Compare commits

..
Author SHA1 Message Date
Niels Lohmann 42e3489abd Prefer strtod_l over libc++'s std::from_chars
libc++'s std::from_chars for float and double is 1.3x to 2.8x slower
per number than Apple's strtod_l, and it was tried before Clinger's
fast path. With Apple clang in C++17 mode, parsing random doubles took
1.75x as long as on develop, short numbers such as 123.45 1.3x, and
mesh.json 1.2x.

Use libc++'s std::from_chars only where the C library has no strtod_l.
On Apple platforms, floats are now converted by Clinger's fast path and
strtod_l, which is 0.90x to 1.01x the time of develop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:12:50 +02:00
Niels Lohmann 436bfb1358 Convert floats independently of the C locale's decimal point
Under a locale whose decimal point is longer than one byte (e.g. U+066B
in fa_IR.UTF-8 or ar_EG.UTF-8), every float that reached the strtod
fallback was truncated at the decimal point: "3.14159265358979323846"
became 3.0, and "1.5e400" became 1.0 instead of throwing. With libc++
and in C++11/14, that fallback was taken for most floats.

The lexer now converts floats in this order:

1. std::from_chars, now also for float and double with libc++ 20 or
   later, which does not define __cpp_lib_to_chars (on Apple platforms
   only if the deployment target provides it);
2. Clinger's fast path (double only);
3. strtof_l/strtod_l/strtold_l with a "C" locale created once, on
   glibc, Apple platforms, and MSVC;
4. strtof/strtod/strtold with the decimal point of the current locale,
   which now puts a multi-byte decimal point into a copy of the token.

If std::from_chars reports a value out of range, the result is derived
from the token (+-infinity or +-0) instead of calling strtod, because
implementations disagree on the stored value (P4168). Values that may
be subnormal are left to the next step, because libstdc++ before
GCC 13 reports some of them as out of range.

The conversion helpers moved from the lexer to number_parse.hpp, so the
last-resort path can be tested directly.

Fixes #5660.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:51:50 +02:00
Niels Lohmann 633de8e44b Fix CI: clang-tidy and GCC -Wnoexcept in the locale test (#5613)
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:

- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
  calls whose result was discarded. Check the result, like the other
  resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
  callbacks, which cannot throw but were not declared noexcept.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:20:43 +02:00
Niels Lohmann fc03b9912e Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction

The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.

token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.

Fixes #5198

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the strtod retry loop when the decimal point is unchanged

convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.

Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Weffc++ errors in the #5198 locale test

GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:56:11 +02:00
Niels Lohmann 9e1a09eec0 Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys

CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:

  syntax error while parsing MessagePack object key: only string keys
  are supported, but found nil; last byte: 0xC0

The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.

Refs #2766, #3381

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point the MessagePack key note to the spec's profile section

The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:51:07 +02:00
20 changed files with 1733 additions and 303 deletions
+2 -2
View File
@@ -80,8 +80,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key,
but not found
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity
@@ -73,8 +73,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key,
but not found
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity
@@ -128,13 +128,10 @@ Strong exception safety: if an exception occurs, the original value stays intact
When the JSON pointer traverses intermediate levels that don't exist at all yet (not just a missing
leaf), each missing level is created as an array or an object depending on whether the corresponding
pointer token is a valid array index: the token `0`, a sequence of digits that does not begin with `0`,
or the token `-` creates an array, and every other token creates an object. For example, on an
initially `#!json null` value, `/foo/0/0/0` creates nested arrays, while `/foo/one/one/one` creates
nested objects. Tokens such as `01` or the empty token cannot be array indices (cf. RFC 6901, Sect. 4)
and therefore create objects, just as they would if the level already existed as an object. This is not
specified by the JSON Pointer RFC; it is this library's own, intentional disambiguation rule. See also
[JSON Pointer](../../features/json_pointer.md).
pointer token parses as a non-negative integer: a numeric token creates an array, a non-numeric token
creates an object. For example, on an initially `#!json null` value, `/foo/0/0/0` creates nested arrays,
while `/foo/one/one/one` creates nested objects. This is not specified by the JSON Pointer RFC; it is
this library's own, intentional disambiguation rule. See also [JSON Pointer](../../features/json_pointer.md).
## Examples
+2
View File
@@ -254,6 +254,8 @@ outside of a string, invalid) byte; see the [FAQ entry](../../home/faq.md#nul-by
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating
it as end of input; planned to become the default in version 4.0.0.
- The result of converting floating-point numbers no longer depends on the C locale in version 3.13.0; before, a
locale whose decimal point is longer than one byte (e.g., `fa_IR.UTF-8`) truncated them at the decimal point.
!!! warning "Deprecation"
@@ -174,7 +174,20 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Object keys"
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps with keys other than UTF-8 strings are rejected.
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps
with keys other than text strings (major type 3) are rejected with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set
to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. This is a deliberate restriction of the library's JSON value model, not an oversight: formats built on CBOR
maps with integer keys, such as COSE ([RFC 9052](https://www.rfc-editor.org/rfc/rfc9052.html)) or CWT
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings"
@@ -138,6 +138,21 @@ The library maps MessagePack types to JSON value types as follows:
Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`.
!!! warning "Object keys"
MessagePack allows map keys of any type, whereas JSON only allows strings as keys in object values. Like the
JSON-compatible [profile](https://github.com/msgpack/msgpack/blob/master/spec.md#profile) sketched in the
MessagePack specification, this library restricts map keys to `str` values. Maps with keys of any other type are
rejected with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
@@ -75,6 +75,13 @@ otherwise, it uses unsigned integer storage.
[`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul),
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof), respectively.
- The result of converting floating-point numbers does not depend on the C locale (`LC_NUMERIC`). They are
converted with [`std::from_chars`](https://en.cppreference.com/w/cpp/utility/from_chars) where the standard
library implements it for the number type (with libc++ 20 or later, only for `#!c float` and `#!c double`, and
only where `strtod_l` is unavailable, because that is faster), otherwise with `strtod_l` and the "C" locale where
the C library provides it (glibc, macOS, MSVC), and otherwise with `std::strtod` and the decimal point of the
current locale. Before version 3.13.0, the last way was used much more often, and a locale whose decimal point
is longer than one byte (e.g., `fa_IR.UTF-8`) truncated numbers at the decimal point.
!!! example "Examples"
@@ -85,10 +92,11 @@ otherwise, it uses unsigned integer storage.
### Number limits
- Any 64-bit signed or unsigned integer can be stored without loss of precision.
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion via
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) are not satisfying
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion are not satisfying
[`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing.
- Numbers too close to zero to be represented as `#!c double`, not even as subnormal number (such as `#!c 1E-400`), are
stored as `#!c 0.0`, or as `#!c -0.0` if they are negative.
- Floating-point numbers are rounded to the next number representable as `double`. For instance
`#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18).
This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`.
+9 -2
View File
@@ -343,13 +343,20 @@ A string could not be read from a [binary format](../features/binary_formats/ind
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
[CBOR](../features/binary_formats/cbor.md) and [MessagePack](../features/binary_formats/messagepack.md).
!!! failure "Example messages"
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82
+168 -2
View File
@@ -1324,6 +1324,80 @@ class binary_reader
}
}
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a definite-length CBOR byte array
@@ -1568,7 +1642,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
{
return false;
}
@@ -2069,6 +2143,98 @@ class binary_reader
}
}
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a MessagePack byte array
@@ -2231,7 +2397,7 @@ class binary_reader
{
get();
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
{
return false;
}
+28 -71
View File
@@ -9,10 +9,8 @@
#pragma once
#include <array> // array
#include <clocale> // localeconv
#include <cstddef> // size_t
#include <cstdio> // snprintf
#include <cstdlib> // strtof, strtod, strtold, strtoll, strtoull
#include <initializer_list> // initializer_list
#include <string> // char_traits, string
#include <utility> // move
@@ -206,7 +204,6 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_)
{}
@@ -218,19 +215,6 @@ class lexer : public lexer_base<BasicJsonType>
~lexer() = default;
private:
/////////////////////
// locales
/////////////////////
/// return the locale-dependent decimal point
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr) ? '.' : *(loc->decimal_point);
}
/////////////////////
// scan functions
/////////////////////
@@ -1038,24 +1022,6 @@ class lexer : public lexer_base<BasicJsonType>
}
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(float& f, const char* str, char** endptr) noexcept
{
f = std::strtof(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(double& f, const char* str, char** endptr) noexcept
{
f = std::strtod(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(long double& f, const char* str, char** endptr) noexcept
{
f = std::strtold(str, endptr);
}
/*!
@brief scan a number literal
@@ -1092,9 +1058,10 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise
@note The scanner is independent of the current locale. Internally, the
locale's decimal point is used instead of `.` to work with the
locale-dependent converters.
@note The scanner is independent of the current locale: token_buffer
always holds `.`. Only the last-resort std::strtod fallback of
convert_number() depends on the locale, and it looks up the decimal
point right before converting (see parse_float_locale_aware()).
*/
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{
@@ -1183,7 +1150,7 @@ scan_number_zero:
{
case '.':
{
add(decimal_point_char);
add(current);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -1220,7 +1187,7 @@ scan_number_any1:
case '.':
{
add(decimal_point_char);
add(current);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -1462,9 +1429,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note
// token_buffer holds the locale's decimal point, so the fraction is
// located through decimal_point_position rather than by searching '.'.
// answer - so this loop is skipped for all but a few tokens. The
// fraction is located through decimal_point_position rather than by
// searching '.'.
if (lead_zero != 0)
{
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -1482,8 +1449,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
scan_number() state machine or by the contiguous fast path) and holds '.'
as decimal point, independent of the locale. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@@ -1562,8 +1529,10 @@ scan_number_done:
// this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod.
// otherwise the exact Clinger fast path (double only); otherwise
// strtof/strtod/strtold with the "C" locale where the C library offers
// that; and only as a last resort strtof/strtod/strtold with the
// decimal point of the current locale.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
@@ -1572,17 +1541,16 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
&& parse_float_fast(num_begin, num_end, value_float))
{
return token_type::value_float;
}
if (parse_float_c_locale(num_begin, num_end, value_float))
{
return token_type::value_float;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
parse_float_locale_aware(token_buffer, decimal_point_position, value_float);
return token_type::value_float;
}
@@ -1591,7 +1559,7 @@ scan_number_done:
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and
(as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -1707,16 +1675,11 @@ scan_number_done:
}
#endif
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
// materialize the token exactly as scan_number() would. reset() already
// cleared token_buffer, so append() fills it (assign() is avoided
// because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
decimal_point_position = dot_index;
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
@@ -1983,11 +1946,7 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once)
string_t& get_string()
{
// translate decimal points from locale back to '.' (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
// a number token holds '.' regardless of the locale (#4084)
return token_buffer;
}
@@ -2283,9 +2242,7 @@ scan_number_done:
number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0;
/// the decimal point
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
/// the position of the decimal point in token_buffer
std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
+368 -25
View File
@@ -10,27 +10,79 @@
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <clocale> // LC_NUMERIC, LC_NUMERIC_MASK, newlocale, _create_locale
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <cstdlib> // strtof, strtod, strtold, strtof_l, strtod_l, strtold_l, _strtof_l, _strtod_l, _strtold_l
#include <limits> // numeric_limits
#include <string> // string
#include <utility> // move
#include <nlohmann/detail/macro_scope.hpp>
// strtof_l/strtod_l/strtold_l convert with a given locale object instead of the
// global C locale. They are not part of ISO C or C++, so they are only used where
// the C library is known to declare them: Microsoft's UCRT (as _strtod_l etc.),
// Apple's libc (in <xlocale.h>, which must follow <cstdlib>), and glibc (as GNU
// extensions, visible because g++ and clang++ define _GNU_SOURCE for C++).
// Everything else, e.g. MinGW (whose runtime lacks them), musl (which declares
// only some of them), Android, or uClibc, uses parse_float_locale_aware().
#if defined(_MSC_VER) && !defined(__MINGW32__) && _MSC_VER >= 1900
#define JSON_HAS_C_LOCALE_STRTOD 1
#elif defined(__APPLE__)
#include <xlocale.h> // newlocale, strtof_l, strtod_l, strtold_l
#define JSON_HAS_C_LOCALE_STRTOD 1
#elif defined(__GLIBC__) && defined(__USE_GNU) && !defined(__UCLIBC__)
#define JSON_HAS_C_LOCALE_STRTOD 1
#else
#define JSON_HAS_C_LOCALE_STRTOD 0
#endif
// std::from_chars lives in <charconv>, but being in C++17 mode does not
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
// include with __has_include so such toolchains fall back to the scalar path.
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
#if __has_include(<charconv>)
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
#include <charconv> // from_chars
#include <system_error> // errc
// std::from_chars is used for floating-point numbers
// - for float, double, and long double if __cpp_lib_to_chars announces
// complete support (only checked in C++17 or later: some standard
// libraries, e.g. libstdc++ 15, define it even in C++14 mode, where
// <charconv> is not included);
// - for float and double with libc++ 20 or later, which does not define
// __cpp_lib_to_chars because long double is missing, but only where
// the C library offers no strtod_l: libc++'s implementation is slower
// than Apple's strtod_l (by 1.3x to 2.8x per number), and it would be
// tried before Clinger's fast path. On Apple platforms, it is also only
// available when deploying to macOS/iOS 26 or later; for older
// deployment targets, _LIBCPP_AVAILABILITY_HAS_FROM_CHARS_FLOATING_POINT
// is 0.
#if defined(__cpp_lib_to_chars)
#define JSON_HAS_FLOAT_FROM_CHARS 1
#define JSON_HAS_LONG_DOUBLE_FROM_CHARS 1
#elif !JSON_HAS_C_LOCALE_STRTOD && defined(_LIBCPP_VERSION) && defined(_LIBCPP_AVAILABILITY_HAS_FROM_CHARS_FLOATING_POINT)
#if _LIBCPP_VERSION >= 200000 && _LIBCPP_AVAILABILITY_HAS_FROM_CHARS_FLOATING_POINT
#define JSON_HAS_FLOAT_FROM_CHARS 1
#endif
#endif
#endif
#endif
#ifndef JSON_HAS_FLOAT_FROM_CHARS
#define JSON_HAS_FLOAT_FROM_CHARS 0
#endif
#ifndef JSON_HAS_LONG_DOUBLE_FROM_CHARS
#define JSON_HAS_LONG_DOUBLE_FROM_CHARS 0
#endif
// This file contains the value-conversion helpers used by the lexer to turn an
// already-validated number token into a value, without the locale/errno
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
// stays focused on scanning; see lexer::convert_number().
// already-validated number token into a value, where possible without the
// locale/errno overhead of std::strtoull/std::strtod. They are free functions so
// the lexer stays focused on scanning; see lexer::convert_number().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -118,14 +170,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
@@ -136,7 +186,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
@@ -175,7 +224,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (static_cast<DecimalPointType>(c) == decimal_point)
else if (c == '.')
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
@@ -260,35 +309,134 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
{
return false;
}
/*!
@brief derive the value of a number token that is out of range
The token [first, last) is a valid JSON number whose value cannot be
represented by @a FloatType. The result follows from the token alone: a value
of at least 1 can only overflow and becomes ±infinity (which the parser reports
as out_of_range.406), a smaller one can only underflow and becomes ±0. The sign
is taken from a leading '-', and the magnitude from the decimal exponent of the
first nonzero digit.
A value slightly below the smallest normal number may still be representable
as a subnormal number, which some implementations also report as out of range
(libstdc++'s std::from_chars before GCC 13, which relies on the ERANGE of
strtod for long double, and in GCC 11 for all types). Therefore ±0 is only
returned if the value is below half the smallest subnormal number whatever its
digits are.
@param[in] first pointer to the first character of the token
@param[in] last pointer past the last character
@param[out] out ±infinity or ±0 on success
@return true if @a out was set; false if the value may be a subnormal number,
in which case the caller converts the token another way
*/
template<typename FloatType>
bool parse_float_out_of_range(const char* first, const char* last, FloatType& out) noexcept
{
const bool negative = first != last && *first == '-';
const char* p = negative ? first + 1 : first;
// the decimal exponent of the first nonzero digit, from its position
// relative to the decimal point
std::int64_t exponent = 0;
bool nonzero = false;
for (; p != last && *p >= '0' && *p <= '9'; ++p)
{
if (nonzero)
{
++exponent;
}
else
{
nonzero = *p != '0';
}
}
if (p != last && *p == '.')
{
for (++p; p != last && *p >= '0' && *p <= '9'; ++p)
{
if (!nonzero)
{
--exponent;
nonzero = *p != '0';
}
}
}
if (nonzero && p != last && (*p == 'e' || *p == 'E'))
{
++p;
const bool negative_exponent = p != last && *p == '-';
if (p != last && (*p == '-' || *p == '+'))
{
++p;
}
// saturate: a larger exponent is far out of range for every type
constexpr std::int64_t saturation = 100000000000000000; // 10^17
std::int64_t explicit_exponent = 0;
for (; p != last && *p >= '0' && *p <= '9'; ++p)
{
if (explicit_exponent < saturation)
{
explicit_exponent = (explicit_exponent * 10) + (*p - '0');
}
}
exponent += negative_exponent ? -explicit_exponent : explicit_exponent;
}
if (nonzero && exponent >= 0)
{
out = negative ? -std::numeric_limits<FloatType>::infinity() : std::numeric_limits<FloatType>::infinity();
return true;
}
// The value is below 10^(exponent + 1). It rounds to zero if that is at most
// half the smallest subnormal number, 2^(min_exponent - digits - 1). The
// bound rounds log10(2) up to 0.30103 and the product toward zero, and the
// margin of 2 keeps it on the safe side.
constexpr std::int64_t zero_exponent = (static_cast<std::int64_t>(std::numeric_limits<FloatType>::min_exponent - std::numeric_limits<FloatType>::digits - 1) * 30103 / 100000) - 2;
if (!nonzero || exponent <= zero_exponent)
{
out = negative ? -FloatType(0) : FloatType(0);
return true;
}
return false;
}
/*!
@brief parse a float with std::from_chars (Eisel-Lemire) when available
std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
over the whole value range (not just the Clinger subset). It is used only where
the standard library implements it for @a FloatType (see
JSON_HAS_FLOAT_FROM_CHARS) and only when it consumes the entire token
([first, last)).
For an under- or overflow (std::errc::result_out_of_range), implementations
disagree on the value they store: libstdc++ leaves it unchanged, whereas libc++
and the MSVC STL store ±0 or ±infinity (P4168). The result is therefore derived
from the token, see parse_float_out_of_range().
@return true if the value was parsed exactly and fully; false to fall back
*/
template<typename FloatType>
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
{
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
// in C++14 mode, where <charconv> is not included.
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
#if JSON_HAS_FLOAT_FROM_CHARS
const auto result = std::from_chars(first, last, out);
if (JSON_HEDLEY_UNLIKELY(result.ec == std::errc::result_out_of_range && result.ptr == last))
{
return parse_float_out_of_range(first, last, out);
}
return result.ec == std::errc() && result.ptr == last;
#else
static_cast<void>(first);
@@ -298,5 +446,200 @@ bool parse_float_from_chars(const char* first, const char* last, FloatType& out)
#endif
}
#if JSON_HAS_FLOAT_FROM_CHARS && !JSON_HAS_LONG_DOUBLE_FROM_CHARS
/// libc++ implements std::from_chars for float and double, but not for long double
inline bool parse_float_from_chars(const char* /*first*/, const char* /*last*/, long double& /*out*/) noexcept
{
return false;
}
#endif
#if JSON_HAS_C_LOCALE_STRTOD
#if defined(_MSC_VER)
using c_locale_t = _locale_t;
/// the "C" locale for the numeric category, created on first use and never freed
inline c_locale_t c_numeric_locale() noexcept
{
static const c_locale_t c_locale = _create_locale(LC_NUMERIC, "C");
return c_locale;
}
inline void strtof_c_locale(float& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = _strtof_l(str, endptr, loc);
}
inline void strtof_c_locale(double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = _strtod_l(str, endptr, loc);
}
inline void strtof_c_locale(long double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = _strtold_l(str, endptr, loc);
}
#else
using c_locale_t = locale_t;
/// the "C" locale for the numeric category, created on first use and never freed
inline c_locale_t c_numeric_locale() noexcept
{
static const c_locale_t c_locale = newlocale(LC_NUMERIC_MASK, "C", nullptr);
return c_locale;
}
inline void strtof_c_locale(float& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = strtof_l(str, endptr, loc);
}
inline void strtof_c_locale(double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = strtod_l(str, endptr, loc);
}
inline void strtof_c_locale(long double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = strtold_l(str, endptr, loc);
}
#endif
#endif
/*!
@brief parse a float with strtof_l/strtod_l/strtold_l in the "C" locale
These functions round correctly like strtod, but take the "C" locale as an
argument instead of using the global one, so the decimal point is always '.'.
The locale object is created on first use and never freed, so it remains valid
for parsers that run during static destruction.
@param[in] first pointer to the first character of the token, which must be
followed by a NUL character
@param[in] last pointer past the last character
@param[out] out the parsed value (±infinity or ±0 if out of range)
@return true if the value was parsed from the entire token; false if the C
library offers no such functions (see JSON_HAS_C_LOCALE_STRTOD) or the
locale could not be created, in which case the caller falls back to
parse_float_locale_aware()
*/
template<typename FloatType>
bool parse_float_c_locale(const char* first, const char* last, FloatType& out) noexcept
{
#if JSON_HAS_C_LOCALE_STRTOD
const c_locale_t loc = c_numeric_locale();
if (JSON_HEDLEY_UNLIKELY(loc == nullptr))
{
return false;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness)
strtof_c_locale(out, first, &endptr, loc);
return endptr == last;
#else
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(out);
return false;
#endif
}
JSON_HEDLEY_NON_NULL(2)
inline void strtof_global_locale(float& f, const char* str, char** endptr) noexcept
{
f = std::strtof(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
inline void strtof_global_locale(double& f, const char* str, char** endptr) noexcept
{
f = std::strtod(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
inline void strtof_global_locale(long double& f, const char* str, char** endptr) noexcept
{
f = std::strtold(str, endptr);
}
/// return the decimal point of the current locale
inline std::string locale_decimal_point()
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr || *loc->decimal_point == '\0') ? "." : loc->decimal_point;
}
/*!
@brief parse a float with strtof/strtod/strtold in the current locale
This is the last resort for platforms without std::from_chars for @a FloatType
and without parse_float_c_locale(). These functions expect the decimal point
of the *current* locale, so the '.' in the token is replaced by it. It is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX handler,
or another thread) must not truncate the value (#5198). A single-byte decimal
point is substituted in place and restored afterwards, because the token is
also handed to the SAX interface. A longer one (e.g., the two-byte U+066B of
fa_IR.UTF-8 or ar_EG.UTF-8) is put into a copy of the token instead.
The token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the lookup
and the call, and the conversion is repeated with the new decimal point. If it
did not change, the value strtod parsed up to that point is kept.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
@param[in,out] token the token, with '.' as decimal point
@param[in] decimal_point_position the position of the '.' in @a token,
or std::string::npos if it has none
@param[out] out the parsed value
*/
template<typename StringType, typename FloatType>
void parse_float_locale_aware(StringType& token, std::size_t decimal_point_position, FloatType& out)
{
const bool has_dot = decimal_point_position != std::string::npos;
std::string decimal_point = locale_decimal_point();
for (;;)
{
char* endptr = nullptr; // NOLINT(misc-const-correctness)
bool complete = false;
if (!has_dot || decimal_point.size() == 1)
{
const bool substitute = has_dot && decimal_point[0] != '.';
if (substitute)
{
token[decimal_point_position] = static_cast<typename StringType::value_type>(decimal_point[0]);
}
strtof_global_locale(out, token.data(), &endptr);
if (substitute)
{
token[decimal_point_position] = '.';
}
complete = endptr == token.data() + token.size();
}
else
{
std::string copy(token.data(), token.size());
copy.replace(decimal_point_position, 1, decimal_point);
strtof_global_locale(out, copy.c_str(), &endptr);
complete = endptr == copy.c_str() + copy.size();
}
if (JSON_HEDLEY_LIKELY(complete))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
std::string current_decimal_point = locale_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = std::move(current_decimal_point);
}
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+5 -9
View File
@@ -416,19 +416,15 @@ class json_pointer
// convert null values to arrays or objects before continuing
if (ptr->is_null())
{
// check if the reference token is a valid array index, that is
// a nonempty sequence of digits without a leading '0'
// (cf. RFC 6901, Sect. 4); tokens that could never be a valid
// array index (such as "01" or "") are treated as object keys
const bool nums = !reference_token.empty()
&& (reference_token.size() == 1 || reference_token[0] != '0')
&& std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
// check if the reference token is a number
const bool nums =
std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
{
return std::isdigit(x);
});
// change value to an array for array indices or "-" or to object otherwise
// change value to an array for numbers or "-" or to object otherwise
*ptr = (nums || reference_token == "-")
? detail::value_t::array
: detail::value_t::object;
@@ -42,6 +42,9 @@
#undef JSON_HAS_RANGES
#undef JSON_HAS_STD_FORMAT
#undef JSON_HAS_STATIC_RTTI
#undef JSON_HAS_FLOAT_FROM_CHARS
#undef JSON_HAS_LONG_DOUBLE_FROM_CHARS
#undef JSON_HAS_C_LOCALE_STRTOD
#undef JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
#undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_PRECISE_STREAM_POSITION
+572 -107
View File
@@ -8472,10 +8472,8 @@ NLOHMANN_JSON_NAMESPACE_END
#include <array> // array
#include <clocale> // localeconv
#include <cstddef> // size_t
#include <cstdio> // snprintf
#include <cstdlib> // strtof, strtod, strtold, strtoll, strtoull
#include <initializer_list> // initializer_list
#include <string> // char_traits, string
#include <utility> // move
@@ -8496,28 +8494,80 @@ NLOHMANN_JSON_NAMESPACE_END
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <clocale> // LC_NUMERIC, LC_NUMERIC_MASK, newlocale, _create_locale
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <cstdlib> // strtof, strtod, strtold, strtof_l, strtod_l, strtold_l, _strtof_l, _strtod_l, _strtold_l
#include <limits> // numeric_limits
#include <string> // string
#include <utility> // move
// #include <nlohmann/detail/macro_scope.hpp>
// strtof_l/strtod_l/strtold_l convert with a given locale object instead of the
// global C locale. They are not part of ISO C or C++, so they are only used where
// the C library is known to declare them: Microsoft's UCRT (as _strtod_l etc.),
// Apple's libc (in <xlocale.h>, which must follow <cstdlib>), and glibc (as GNU
// extensions, visible because g++ and clang++ define _GNU_SOURCE for C++).
// Everything else, e.g. MinGW (whose runtime lacks them), musl (which declares
// only some of them), Android, or uClibc, uses parse_float_locale_aware().
#if defined(_MSC_VER) && !defined(__MINGW32__) && _MSC_VER >= 1900
#define JSON_HAS_C_LOCALE_STRTOD 1
#elif defined(__APPLE__)
#include <xlocale.h> // newlocale, strtof_l, strtod_l, strtold_l
#define JSON_HAS_C_LOCALE_STRTOD 1
#elif defined(__GLIBC__) && defined(__USE_GNU) && !defined(__UCLIBC__)
#define JSON_HAS_C_LOCALE_STRTOD 1
#else
#define JSON_HAS_C_LOCALE_STRTOD 0
#endif
// std::from_chars lives in <charconv>, but being in C++17 mode does not
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
// include with __has_include so such toolchains fall back to the scalar path.
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
#if __has_include(<charconv>)
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
#include <charconv> // from_chars
#include <system_error> // errc
// std::from_chars is used for floating-point numbers
// - for float, double, and long double if __cpp_lib_to_chars announces
// complete support (only checked in C++17 or later: some standard
// libraries, e.g. libstdc++ 15, define it even in C++14 mode, where
// <charconv> is not included);
// - for float and double with libc++ 20 or later, which does not define
// __cpp_lib_to_chars because long double is missing, but only where
// the C library offers no strtod_l: libc++'s implementation is slower
// than Apple's strtod_l (by 1.3x to 2.8x per number), and it would be
// tried before Clinger's fast path. On Apple platforms, it is also only
// available when deploying to macOS/iOS 26 or later; for older
// deployment targets, _LIBCPP_AVAILABILITY_HAS_FROM_CHARS_FLOATING_POINT
// is 0.
#if defined(__cpp_lib_to_chars)
#define JSON_HAS_FLOAT_FROM_CHARS 1
#define JSON_HAS_LONG_DOUBLE_FROM_CHARS 1
#elif !JSON_HAS_C_LOCALE_STRTOD && defined(_LIBCPP_VERSION) && defined(_LIBCPP_AVAILABILITY_HAS_FROM_CHARS_FLOATING_POINT)
#if _LIBCPP_VERSION >= 200000 && _LIBCPP_AVAILABILITY_HAS_FROM_CHARS_FLOATING_POINT
#define JSON_HAS_FLOAT_FROM_CHARS 1
#endif
#endif
#endif
#endif
#ifndef JSON_HAS_FLOAT_FROM_CHARS
#define JSON_HAS_FLOAT_FROM_CHARS 0
#endif
#ifndef JSON_HAS_LONG_DOUBLE_FROM_CHARS
#define JSON_HAS_LONG_DOUBLE_FROM_CHARS 0
#endif
// This file contains the value-conversion helpers used by the lexer to turn an
// already-validated number token into a value, without the locale/errno
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
// stays focused on scanning; see lexer::convert_number().
// already-validated number token into a value, where possible without the
// locale/errno overhead of std::strtoull/std::strtod. They are free functions so
// the lexer stays focused on scanning; see lexer::convert_number().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -8605,14 +8655,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
@@ -8623,7 +8671,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
@@ -8662,7 +8709,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (static_cast<DecimalPointType>(c) == decimal_point)
else if (c == '.')
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
@@ -8747,35 +8794,134 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
{
return false;
}
/*!
@brief derive the value of a number token that is out of range
The token [first, last) is a valid JSON number whose value cannot be
represented by @a FloatType. The result follows from the token alone: a value
of at least 1 can only overflow and becomes ±infinity (which the parser reports
as out_of_range.406), a smaller one can only underflow and becomes ±0. The sign
is taken from a leading '-', and the magnitude from the decimal exponent of the
first nonzero digit.
A value slightly below the smallest normal number may still be representable
as a subnormal number, which some implementations also report as out of range
(libstdc++'s std::from_chars before GCC 13, which relies on the ERANGE of
strtod for long double, and in GCC 11 for all types). Therefore ±0 is only
returned if the value is below half the smallest subnormal number whatever its
digits are.
@param[in] first pointer to the first character of the token
@param[in] last pointer past the last character
@param[out] out ±infinity or ±0 on success
@return true if @a out was set; false if the value may be a subnormal number,
in which case the caller converts the token another way
*/
template<typename FloatType>
bool parse_float_out_of_range(const char* first, const char* last, FloatType& out) noexcept
{
const bool negative = first != last && *first == '-';
const char* p = negative ? first + 1 : first;
// the decimal exponent of the first nonzero digit, from its position
// relative to the decimal point
std::int64_t exponent = 0;
bool nonzero = false;
for (; p != last && *p >= '0' && *p <= '9'; ++p)
{
if (nonzero)
{
++exponent;
}
else
{
nonzero = *p != '0';
}
}
if (p != last && *p == '.')
{
for (++p; p != last && *p >= '0' && *p <= '9'; ++p)
{
if (!nonzero)
{
--exponent;
nonzero = *p != '0';
}
}
}
if (nonzero && p != last && (*p == 'e' || *p == 'E'))
{
++p;
const bool negative_exponent = p != last && *p == '-';
if (p != last && (*p == '-' || *p == '+'))
{
++p;
}
// saturate: a larger exponent is far out of range for every type
constexpr std::int64_t saturation = 100000000000000000; // 10^17
std::int64_t explicit_exponent = 0;
for (; p != last && *p >= '0' && *p <= '9'; ++p)
{
if (explicit_exponent < saturation)
{
explicit_exponent = (explicit_exponent * 10) + (*p - '0');
}
}
exponent += negative_exponent ? -explicit_exponent : explicit_exponent;
}
if (nonzero && exponent >= 0)
{
out = negative ? -std::numeric_limits<FloatType>::infinity() : std::numeric_limits<FloatType>::infinity();
return true;
}
// The value is below 10^(exponent + 1). It rounds to zero if that is at most
// half the smallest subnormal number, 2^(min_exponent - digits - 1). The
// bound rounds log10(2) up to 0.30103 and the product toward zero, and the
// margin of 2 keeps it on the safe side.
constexpr std::int64_t zero_exponent = (static_cast<std::int64_t>(std::numeric_limits<FloatType>::min_exponent - std::numeric_limits<FloatType>::digits - 1) * 30103 / 100000) - 2;
if (!nonzero || exponent <= zero_exponent)
{
out = negative ? -FloatType(0) : FloatType(0);
return true;
}
return false;
}
/*!
@brief parse a float with std::from_chars (Eisel-Lemire) when available
std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
over the whole value range (not just the Clinger subset). It is used only where
the standard library implements it for @a FloatType (see
JSON_HAS_FLOAT_FROM_CHARS) and only when it consumes the entire token
([first, last)).
For an under- or overflow (std::errc::result_out_of_range), implementations
disagree on the value they store: libstdc++ leaves it unchanged, whereas libc++
and the MSVC STL store ±0 or ±infinity (P4168). The result is therefore derived
from the token, see parse_float_out_of_range().
@return true if the value was parsed exactly and fully; false to fall back
*/
template<typename FloatType>
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
{
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
// in C++14 mode, where <charconv> is not included.
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
#if JSON_HAS_FLOAT_FROM_CHARS
const auto result = std::from_chars(first, last, out);
if (JSON_HEDLEY_UNLIKELY(result.ec == std::errc::result_out_of_range && result.ptr == last))
{
return parse_float_out_of_range(first, last, out);
}
return result.ec == std::errc() && result.ptr == last;
#else
static_cast<void>(first);
@@ -8785,6 +8931,201 @@ bool parse_float_from_chars(const char* first, const char* last, FloatType& out)
#endif
}
#if JSON_HAS_FLOAT_FROM_CHARS && !JSON_HAS_LONG_DOUBLE_FROM_CHARS
/// libc++ implements std::from_chars for float and double, but not for long double
inline bool parse_float_from_chars(const char* /*first*/, const char* /*last*/, long double& /*out*/) noexcept
{
return false;
}
#endif
#if JSON_HAS_C_LOCALE_STRTOD
#if defined(_MSC_VER)
using c_locale_t = _locale_t;
/// the "C" locale for the numeric category, created on first use and never freed
inline c_locale_t c_numeric_locale() noexcept
{
static const c_locale_t c_locale = _create_locale(LC_NUMERIC, "C");
return c_locale;
}
inline void strtof_c_locale(float& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = _strtof_l(str, endptr, loc);
}
inline void strtof_c_locale(double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = _strtod_l(str, endptr, loc);
}
inline void strtof_c_locale(long double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = _strtold_l(str, endptr, loc);
}
#else
using c_locale_t = locale_t;
/// the "C" locale for the numeric category, created on first use and never freed
inline c_locale_t c_numeric_locale() noexcept
{
static const c_locale_t c_locale = newlocale(LC_NUMERIC_MASK, "C", nullptr);
return c_locale;
}
inline void strtof_c_locale(float& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = strtof_l(str, endptr, loc);
}
inline void strtof_c_locale(double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = strtod_l(str, endptr, loc);
}
inline void strtof_c_locale(long double& f, const char* str, char** endptr, c_locale_t loc) noexcept
{
f = strtold_l(str, endptr, loc);
}
#endif
#endif
/*!
@brief parse a float with strtof_l/strtod_l/strtold_l in the "C" locale
These functions round correctly like strtod, but take the "C" locale as an
argument instead of using the global one, so the decimal point is always '.'.
The locale object is created on first use and never freed, so it remains valid
for parsers that run during static destruction.
@param[in] first pointer to the first character of the token, which must be
followed by a NUL character
@param[in] last pointer past the last character
@param[out] out the parsed value (±infinity or ±0 if out of range)
@return true if the value was parsed from the entire token; false if the C
library offers no such functions (see JSON_HAS_C_LOCALE_STRTOD) or the
locale could not be created, in which case the caller falls back to
parse_float_locale_aware()
*/
template<typename FloatType>
bool parse_float_c_locale(const char* first, const char* last, FloatType& out) noexcept
{
#if JSON_HAS_C_LOCALE_STRTOD
const c_locale_t loc = c_numeric_locale();
if (JSON_HEDLEY_UNLIKELY(loc == nullptr))
{
return false;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness)
strtof_c_locale(out, first, &endptr, loc);
return endptr == last;
#else
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(out);
return false;
#endif
}
JSON_HEDLEY_NON_NULL(2)
inline void strtof_global_locale(float& f, const char* str, char** endptr) noexcept
{
f = std::strtof(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
inline void strtof_global_locale(double& f, const char* str, char** endptr) noexcept
{
f = std::strtod(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
inline void strtof_global_locale(long double& f, const char* str, char** endptr) noexcept
{
f = std::strtold(str, endptr);
}
/// return the decimal point of the current locale
inline std::string locale_decimal_point()
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr || *loc->decimal_point == '\0') ? "." : loc->decimal_point;
}
/*!
@brief parse a float with strtof/strtod/strtold in the current locale
This is the last resort for platforms without std::from_chars for @a FloatType
and without parse_float_c_locale(). These functions expect the decimal point
of the *current* locale, so the '.' in the token is replaced by it. It is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX handler,
or another thread) must not truncate the value (#5198). A single-byte decimal
point is substituted in place and restored afterwards, because the token is
also handed to the SAX interface. A longer one (e.g., the two-byte U+066B of
fa_IR.UTF-8 or ar_EG.UTF-8) is put into a copy of the token instead.
The token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the lookup
and the call, and the conversion is repeated with the new decimal point. If it
did not change, the value strtod parsed up to that point is kept.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
@param[in,out] token the token, with '.' as decimal point
@param[in] decimal_point_position the position of the '.' in @a token,
or std::string::npos if it has none
@param[out] out the parsed value
*/
template<typename StringType, typename FloatType>
void parse_float_locale_aware(StringType& token, std::size_t decimal_point_position, FloatType& out)
{
const bool has_dot = decimal_point_position != std::string::npos;
std::string decimal_point = locale_decimal_point();
for (;;)
{
char* endptr = nullptr; // NOLINT(misc-const-correctness)
bool complete = false;
if (!has_dot || decimal_point.size() == 1)
{
const bool substitute = has_dot && decimal_point[0] != '.';
if (substitute)
{
token[decimal_point_position] = static_cast<typename StringType::value_type>(decimal_point[0]);
}
strtof_global_locale(out, token.data(), &endptr);
if (substitute)
{
token[decimal_point_position] = '.';
}
complete = endptr == token.data() + token.size();
}
else
{
std::string copy(token.data(), token.size());
copy.replace(decimal_point_position, 1, decimal_point);
strtof_global_locale(out, copy.c_str(), &endptr);
complete = endptr == copy.c_str() + copy.size();
}
if (JSON_HEDLEY_LIKELY(complete))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
std::string current_decimal_point = locale_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = std::move(current_decimal_point);
}
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -9303,7 +9644,6 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_)
{}
@@ -9315,19 +9655,6 @@ class lexer : public lexer_base<BasicJsonType>
~lexer() = default;
private:
/////////////////////
// locales
/////////////////////
/// return the locale-dependent decimal point
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr) ? '.' : *(loc->decimal_point);
}
/////////////////////
// scan functions
/////////////////////
@@ -10135,24 +10462,6 @@ class lexer : public lexer_base<BasicJsonType>
}
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(float& f, const char* str, char** endptr) noexcept
{
f = std::strtof(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(double& f, const char* str, char** endptr) noexcept
{
f = std::strtod(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(long double& f, const char* str, char** endptr) noexcept
{
f = std::strtold(str, endptr);
}
/*!
@brief scan a number literal
@@ -10189,9 +10498,10 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise
@note The scanner is independent of the current locale. Internally, the
locale's decimal point is used instead of `.` to work with the
locale-dependent converters.
@note The scanner is independent of the current locale: token_buffer
always holds `.`. Only the last-resort std::strtod fallback of
convert_number() depends on the locale, and it looks up the decimal
point right before converting (see parse_float_locale_aware()).
*/
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{
@@ -10280,7 +10590,7 @@ scan_number_zero:
{
case '.':
{
add(decimal_point_char);
add(current);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -10317,7 +10627,7 @@ scan_number_any1:
case '.':
{
add(decimal_point_char);
add(current);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -10559,9 +10869,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note
// token_buffer holds the locale's decimal point, so the fraction is
// located through decimal_point_position rather than by searching '.'.
// answer - so this loop is skipped for all but a few tokens. The
// fraction is located through decimal_point_position rather than by
// searching '.'.
if (lead_zero != 0)
{
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -10579,8 +10889,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
scan_number() state machine or by the contiguous fast path) and holds '.'
as decimal point, independent of the locale. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@@ -10659,8 +10969,10 @@ scan_number_done:
// this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod.
// otherwise the exact Clinger fast path (double only); otherwise
// strtof/strtod/strtold with the "C" locale where the C library offers
// that; and only as a last resort strtof/strtod/strtold with the
// decimal point of the current locale.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
@@ -10669,17 +10981,16 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
&& parse_float_fast(num_begin, num_end, value_float))
{
return token_type::value_float;
}
if (parse_float_c_locale(num_begin, num_end, value_float))
{
return token_type::value_float;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
parse_float_locale_aware(token_buffer, decimal_point_position, value_float);
return token_type::value_float;
}
@@ -10688,7 +10999,7 @@ scan_number_done:
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and
(as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -10804,16 +11115,11 @@ scan_number_done:
}
#endif
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
// materialize the token exactly as scan_number() would. reset() already
// cleared token_buffer, so append() fills it (assign() is avoided
// because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
decimal_point_position = dot_index;
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
@@ -11080,11 +11386,7 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once)
string_t& get_string()
{
// translate decimal points from locale back to '.' (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
// a number token holds '.' regardless of the locale (#4084)
return token_buffer;
}
@@ -11380,9 +11682,7 @@ scan_number_done:
number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0;
/// the decimal point
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
/// the position of the decimal point in token_buffer
std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
@@ -14059,6 +14359,80 @@ class binary_reader
}
}
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a definite-length CBOR byte array
@@ -14303,7 +14677,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
{
return false;
}
@@ -14804,6 +15178,98 @@ class binary_reader
}
}
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a MessagePack byte array
@@ -14966,7 +15432,7 @@ class binary_reader
{
get();
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
{
return false;
}
@@ -19060,19 +19526,15 @@ class json_pointer
// convert null values to arrays or objects before continuing
if (ptr->is_null())
{
// check if the reference token is a valid array index, that is
// a nonempty sequence of digits without a leading '0'
// (cf. RFC 6901, Sect. 4); tokens that could never be a valid
// array index (such as "01" or "") are treated as object keys
const bool nums = !reference_token.empty()
&& (reference_token.size() == 1 || reference_token[0] != '0')
&& std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
// check if the reference token is a number
const bool nums =
std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
{
return std::isdigit(x);
});
// change value to an array for array indices or "-" or to object otherwise
// change value to an array for numbers or "-" or to object otherwise
*ptr = (nums || reference_token == "-")
? detail::value_t::array
: detail::value_t::object;
@@ -32549,6 +33011,9 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#undef JSON_HAS_RANGES
#undef JSON_HAS_STD_FORMAT
#undef JSON_HAS_STATIC_RTTI
#undef JSON_HAS_FLOAT_FROM_CHARS
#undef JSON_HAS_LONG_DOUBLE_FROM_CHARS
#undef JSON_HAS_C_LOCALE_STRTOD
#undef JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
#undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_PRECISE_STREAM_POSITION
+43 -2
View File
@@ -1830,10 +1830,51 @@ TEST_CASE("CBOR")
SECTION("invalid string in map")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01}), true, false).is_discarded());
}
SECTION("non-string key (see #2766 and #3381)")
{
// only text strings map to JSON object keys; any other key is
// rejected with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0xA1, 0x01, 0x01}, "an unsigned integer; last byte: 0x01"},
{{0xA1, 0x20, 0x01}, "a negative integer; last byte: 0x20"},
{{0xA1, 0x41, 0x61, 0x01}, "a byte string; last byte: 0x41"},
{{0xA1, 0x80, 0x01}, "an array; last byte: 0x80"},
{{0xA1, 0xA0, 0x01}, "a map; last byte: 0xA0"},
{{0xA1, 0xC0, 0x61, 0x61, 0x01}, "a tag; last byte: 0xC0"},
{{0xA1, 0xF4, 0x01}, "a boolean; last byte: 0xF4"},
{{0xA1, 0xF5, 0x01}, "a boolean; last byte: 0xF5"},
{{0xA1, 0xF6, 0x01}, "null; last byte: 0xF6"},
{{0xA1, 0xF7, 0x01}, "undefined; last byte: 0xF7"},
{{0xA1, 0xF9, 0x3C, 0x00, 0x01}, "a floating-point number; last byte: 0xF9"},
{{0xA1, 0xFA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFA"},
{{0xA1, 0xFB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFB"},
{{0xA1, 0xE0, 0x01}, "a simple value; last byte: 0xE0"},
{{0xA1, 0xF8, 0x20, 0x01}, "a simple value; last byte: 0xF8"},
// indefinite-length map
{{0xBF, 0x01, 0x01, 0xFF}, "an unsigned integer; last byte: 0x01"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_cbor(c.first, true, false).is_discarded());
}
// a key of major type 3 with a reserved length is still reported as
// a malformed string, and a missing key as the end of input
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
{
// a two-character text string (major type 3) whose bytes are not
@@ -2284,7 +2325,7 @@ TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
SECTION("a break marker outside an indefinite-length string is not a string")
{
// 0xFF only closes a string that was opened; on its own it is not one
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
}
}
+141 -1
View File
@@ -13,7 +13,10 @@
using nlohmann::json;
#include <cfloat> // FLT_EVAL_METHOD
#include <cmath> // signbit
#include <cstdlib> // strtod
#include <limits> // numeric_limits
#include <map> // map
#include <sstream> // stringstream
#include <string> // string
#include <vector> // vector
@@ -666,7 +669,7 @@ TEST_CASE("parse_float_fast declines what it cannot convert exactly")
// always safe: the caller then falls back to a slower, exact conversion.
const auto fast = [](const std::string & s, double & out)
{
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), '.', out);
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out);
};
double out = 0;
@@ -700,3 +703,140 @@ TEST_CASE("parse_float_fast declines what it cannot convert exactly")
CHECK_FALSE(fast("1e23", out));
CHECK_FALSE(fast("1e-23", out));
}
namespace
{
template<typename FloatType>
bool out_of_range_value(const std::string& s, FloatType& out)
{
return nlohmann::detail::parse_float_out_of_range(s.data(), s.data() + s.size(), out);
}
} // namespace
TEST_CASE("parse_float_out_of_range derives the value from the token")
{
// std::from_chars reports numbers out of range without a portable value
// (P4168), so the value is derived from the token
const double inf = std::numeric_limits<double>::infinity();
double out = 1.0;
SECTION("overflow")
{
CHECK(out_of_range_value("1e400", out));
CHECK(out == inf);
CHECK(out_of_range_value("-1E+400", out));
CHECK(out == -inf);
CHECK(out_of_range_value("123.456e306", out));
CHECK(out == inf);
CHECK(out_of_range_value("0.001e99999999999999999999", out));
CHECK(out == inf);
CHECK(out_of_range_value("-1" + std::string(400, '0'), out));
CHECK(out == -inf);
}
SECTION("underflow")
{
CHECK(out_of_range_value("1e-400", out));
CHECK(out == 0.0);
CHECK(!std::signbit(out));
CHECK(out_of_range_value("-1e-400", out));
CHECK(out == 0.0);
CHECK(std::signbit(out));
CHECK(out_of_range_value("0.00012e-321", out));
CHECK(out == 0.0);
CHECK(out_of_range_value("-1234e-99999999999999999999", out));
CHECK(std::signbit(out));
CHECK(out_of_range_value("-0.0", out));
CHECK(out == 0.0);
CHECK(std::signbit(out));
}
SECTION("possibly subnormal")
{
// some implementations report subnormal numbers as out of range; the
// caller then converts them another way
CHECK_FALSE(out_of_range_value("0.0012e-321", out));
CHECK_FALSE(out_of_range_value("2.5e-320", out));
CHECK_FALSE(out_of_range_value("-1e-310", out));
}
SECTION("float")
{
float f = 1.0f;
CHECK(out_of_range_value("-1e39", f));
CHECK(f == -std::numeric_limits<float>::infinity());
CHECK(out_of_range_value("1e-47", f));
CHECK(f == 0.0f);
CHECK_FALSE(out_of_range_value("1e-46", f));
CHECK_FALSE(out_of_range_value("1e-40", f));
}
SECTION("long double")
{
long double ld = 1.0L;
CHECK(out_of_range_value("1e5000", ld));
CHECK(ld == std::numeric_limits<long double>::infinity());
CHECK(out_of_range_value("-1e-5000", ld));
CHECK(ld == 0.0L);
CHECK(std::signbit(ld));
}
}
TEST_CASE("floating-point numbers out of range")
{
// Whichever conversion the platform uses, an overflow throws, and an
// underflow yields a zero with the sign of the number.
using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
SECTION("double")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::parse("1.5e400"), "[json.exception.out_of_range.406] number overflow parsing '1.5e400'", json::out_of_range&);
CHECK_THROWS_WITH_AS(_ = json::parse("-1.5e400"), "[json.exception.out_of_range.406] number overflow parsing '-1.5e400'", json::out_of_range&);
CHECK_THROWS_WITH_AS(_ = json::parse("1e99999999999999999999"), "[json.exception.out_of_range.406] number overflow parsing '1e99999999999999999999'", json::out_of_range&);
CHECK_THROWS_AS(_ = json::parse("1" + std::string(400, '0')), json::out_of_range&);
const json zero = json::parse("1.5e-400");
CHECK(zero == 0.0);
CHECK(!std::signbit(zero.get<double>()));
const json negative_zero = json::parse("-1.5e-400");
CHECK(negative_zero == 0.0);
CHECK(std::signbit(negative_zero.get<double>()));
CHECK(std::signbit(json::parse("-0.0000000001e-99999999999999999999").get<double>()));
// around the smallest subnormal number
CHECK(json::parse("1e-324") == 0.0);
CHECK(json::parse("3e-324") == std::numeric_limits<double>::denorm_min());
CHECK(json::parse("-2.5e-320") == -2.5e-320);
}
SECTION("float")
{
float_json _;
CHECK_THROWS_WITH_AS(_ = float_json::parse("1e39"), "[json.exception.out_of_range.406] number overflow parsing '1e39'", json::out_of_range&);
CHECK_THROWS_WITH_AS(_ = float_json::parse("-1e39"), "[json.exception.out_of_range.406] number overflow parsing '-1e39'", json::out_of_range&);
const float_json zero = float_json::parse("1e-50");
CHECK(zero == 0.0f);
CHECK(!std::signbit(zero.get<float>()));
const float_json negative_zero = float_json::parse("-1e-50");
CHECK(negative_zero == 0.0f);
CHECK(std::signbit(negative_zero.get<float>()));
CHECK(float_json::parse("1e-45") == std::numeric_limits<float>::denorm_min());
}
SECTION("long double")
{
long_double_json _;
CHECK_THROWS_WITH_AS(_ = long_double_json::parse("1e5000"), "[json.exception.out_of_range.406] number overflow parsing '1e5000'", json::out_of_range&);
CHECK_THROWS_WITH_AS(_ = long_double_json::parse("-1e5000"), "[json.exception.out_of_range.406] number overflow parsing '-1e5000'", json::out_of_range&);
const long_double_json zero = long_double_json::parse("1e-5000");
CHECK(zero == 0.0L);
CHECK(!std::signbit(zero.get<long double>()));
const long_double_json negative_zero = long_double_json::parse("-1e-5000");
CHECK(negative_zero == 0.0L);
CHECK(std::signbit(negative_zero.get<long double>()));
}
}
-66
View File
@@ -438,72 +438,6 @@ TEST_CASE("JSON pointers")
}
}
SECTION("creating intermediate levels")
{
SECTION("tokens that are valid array indices create arrays")
{
json j;
j["/0"_json_pointer] = 1;
CHECK(j == json({1}));
json j2;
j2["/2"_json_pointer] = 1;
CHECK(j2 == json({nullptr, nullptr, 1}));
json j3;
j3["/-"_json_pointer] = 1;
CHECK(j3 == json({1}));
json j4;
j4["/foo/0/0"_json_pointer] = 1;
CHECK(j4 == json({{"foo", {{1}}}}));
}
SECTION("tokens that are no valid array indices create objects")
{
json j;
j["/one"_json_pointer] = 1;
CHECK(j == json({{"one", 1}}));
// leading '0' can never be a valid array index (RFC 6901, Sect. 4)
json j2;
j2["/01"_json_pointer] = 1;
CHECK(j2 == json({{"01", 1}}));
// the empty token is a valid object key, but no valid array index
json j3;
j3["/"_json_pointer] = 1;
CHECK(j3 == json({{"", 1}}));
}
SECTION("creating a level yields the same result as reusing it (#5357)")
{
json j;
j["/a/b/01/d"_json_pointer] = "value";
json j_init = json::object();
j_init["/a/b"_json_pointer] = json::object();
j_init["/a/b/01/d"_json_pointer] = "value";
const json expected = json::parse(R"({"a":{"b":{"01":{"d":"value"}}}})");
CHECK(j == expected);
CHECK(j_init == expected);
// unflatten uses the same key
const json flat = {{"/a/b/01/d", "value"}};
CHECK(flat.unflatten() == expected);
}
SECTION("existing arrays still reject invalid indices")
{
json j = {1, 2, 3};
CHECK_THROWS_WITH_AS(j["/01"_json_pointer],
"[json.exception.parse_error.106] parse error: array index '01' must not begin with '0'", json::parse_error&);
CHECK_THROWS_WITH_AS(j.at("/01"_json_pointer),
"[json.exception.parse_error.106] parse error: array index '01' must not begin with '0'", json::parse_error&);
}
}
SECTION("flatten")
{
json j =
+284
View File
@@ -12,7 +12,14 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <array>
#include <clocale>
#include <cstring>
#include <limits>
#include <map>
#include <string>
#include <utility>
#include <vector>
struct ParserImpl final: public nlohmann::json_sax<json>
{
@@ -175,3 +182,280 @@ TEST_CASE("locale-dependent test (LC_NUMERIC=de_DE)")
MESSAGE("locale de_DE is not usable");
}
}
namespace
{
// records the numbers of a flat array and switches LC_NUMERIC to the given
// locale once the array opens - after the lexer was constructed, but before
// any number in the array is lexed
struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
{
explicit LocaleSwitchingSax(const char* switch_to)
: locale_after_open(switch_to)
{}
bool null() override
{
return true;
}
bool boolean(bool /*val*/) override
{
return true;
}
bool number_integer(json::number_integer_t /*val*/) override
{
return true;
}
bool number_unsigned(json::number_unsigned_t /*val*/) override
{
return true;
}
bool number_float(json::number_float_t val, const json::string_t& s) override
{
values.push_back(val);
strings.push_back(s);
return true;
}
bool string(json::string_t& /*val*/) override
{
return true;
}
bool binary(json::binary_t& /*val*/) override
{
return true;
}
bool start_object(std::size_t /*val*/) override
{
return true;
}
bool key(json::string_t& /*val*/) override
{
return true;
}
bool end_object() override
{
return true;
}
bool start_array(std::size_t /*val*/) override
{
switched = std::setlocale(LC_NUMERIC, locale_after_open.c_str()) != nullptr;
return true;
}
bool end_array() override
{
return true;
}
bool parse_error(std::size_t /*val*/, const std::string& /*val*/, const nlohmann::detail::exception& /*val*/) override
{
return false;
}
std::string locale_after_open;
bool switched = false;
std::vector<json::number_float_t> values {}; // NOLINT(readability-redundant-member-init)
std::vector<json::string_t> strings {}; // NOLINT(readability-redundant-member-init)
};
} // namespace
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
{
// The numbers are chosen so that the conversion takes the slower paths: too
// many significant digits for Clinger's fast path, an underflow, and a plain
// value. Without std::from_chars and strtod_l, this is the strtod fallback,
// which honors the locale that is current at conversion time.
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
std::string text = "[";
for (const auto& n : numbers)
{
text += (text.size() == 1 ? "" : ",") + n;
}
text += "]";
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
// reference values, parsed without a locale switch
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
const json expected = json::parse(text);
const long_double_json expected_ld = long_double_json::parse(text);
const std::array<std::pair<const char*, const char*>, 2> transitions =
{
{
{"C", "de_DE"},
{"de_DE", "C"}
}
};
for (const auto& transition : transitions)
{
CAPTURE(transition.first);
CAPTURE(transition.second);
if (std::setlocale(LC_NUMERIC, transition.first) == nullptr)
{
MESSAGE("locale is not usable");
continue;
}
// SAX parsing
{
LocaleSwitchingSax sax(transition.second);
CHECK(json::sax_parse(text, &sax));
if (sax.switched)
{
CHECK(sax.values == expected.get<std::vector<json::number_float_t>>());
CHECK(sax.strings == numbers);
}
}
// DOM parsing with a callback
{
bool switched = false;
const auto cb = [&](int /*depth*/, json::parse_event_t event, json& /*parsed*/) noexcept
{
if (event == json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const json j = json::parse(text, cb);
if (switched)
{
CHECK(j == expected);
}
}
// a long double goes through std::strtold unless std::from_chars supports it
{
bool switched = false;
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
{
if (event == long_double_json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const long_double_json j = long_double_json::parse(text, cb);
if (switched)
{
CHECK(j == expected_ld);
}
}
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
namespace
{
// sets LC_NUMERIC to the first installed locale whose decimal point is longer
// than one byte, e.g. U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8)
const char* set_multi_byte_decimal_point_locale()
{
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
for (const char* name : names)
{
if (std::setlocale(LC_NUMERIC, name) != nullptr && std::strlen(std::localeconv()->decimal_point) > 1)
{
return name;
}
}
return nullptr;
}
} // namespace
TEST_CASE("locale with a multi-byte decimal point")
{
// Such a decimal point cannot be substituted in place for '.'; before
// #5660, the strtod fallback stopped there and returned the integer part.
const char* name = set_multi_byte_decimal_point_locale();
if (name == nullptr)
{
MESSAGE("no locale with a multi-byte decimal point is usable");
}
else
{
const std::string locale_name = name;
CAPTURE(locale_name);
// too many significant digits for Clinger's fast path
CHECK(json::parse("3.141592653589793238462643383279") == 3.141592653589793);
CHECK(json::parse("1.7976931348623157e308") == (std::numeric_limits<double>::max)());
CHECK(json::accept("3.14159265358979323846"));
// a subnormal number
CHECK(json::parse("-2.5e-320") == -2.5e-320);
// out of range
json _;
CHECK_THROWS_WITH_AS(_ = json::parse("1.5e400"), "[json.exception.out_of_range.406] number overflow parsing '1.5e400'", json::out_of_range&);
CHECK(json::parse("1.5e-400") == 0.0);
// float and long double as number_float_t
using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
CHECK(float_json::parse("1.5") == 1.5f);
CHECK(long_double_json::parse("1.5") == 1.5L);
// a value Clinger's fast path converts
CHECK(json::parse("12.5") == 12.5);
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
TEST_CASE("conversion with the decimal point of the current locale")
{
// parse_float_locale_aware() is the last resort for platforms without
// std::from_chars and strtod_l, so it is called directly here
const auto convert = [](std::string token, double & out)
{
nlohmann::detail::parse_float_locale_aware(token, token.find('.'), out);
// the token is also handed to the SAX interface and must keep its '.'
return token;
};
std::vector<const char*> names = {"C", "de_DE", "de_DE.UTF-8"};
const char* multi_byte = set_multi_byte_decimal_point_locale();
if (multi_byte != nullptr)
{
names.push_back(multi_byte);
}
for (const char* name : names)
{
if (std::setlocale(LC_NUMERIC, name) == nullptr)
{
continue;
}
const std::string locale_name = name;
CAPTURE(locale_name);
double d = 0;
CHECK(convert("3.141592653589793238462643383279", d) == "3.141592653589793238462643383279");
CHECK(d == 3.141592653589793);
CHECK(convert("-2.5e-320", d) == "-2.5e-320");
CHECK(d == -2.5e-320);
CHECK(convert("12345678901234567890", d) == "12345678901234567890");
CHECK(d == 12345678901234567890.0);
float f = 0;
std::string token = "1.5";
nlohmann::detail::parse_float_locale_aware(token, 1, f);
CHECK(f == 1.5f);
long double ld = 0;
nlohmann::detail::parse_float_locale_aware(token, 1, ld);
CHECK(ld == 1.5L);
CHECK(token == "1.5");
// the lexer only passes valid tokens; for others, the conversion stops
// early, and the value parsed up to there is kept
CHECK(convert("1.5x", d) == "1.5x");
CHECK(d == 1.5);
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
+60 -1
View File
@@ -1551,10 +1551,69 @@ TEST_CASE("MessagePack")
SECTION("invalid string in map")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found an integer; last byte: 0xFF", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01}), true, false).is_discarded());
}
SECTION("non-string key (see #3381)")
{
// only strings map to JSON object keys; any other key is rejected
// with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0x81, 0xC0, 0x01}, "nil; last byte: 0xC0"},
{{0x81, 0xC2, 0x01}, "a boolean; last byte: 0xC2"},
{{0x81, 0xC3, 0x01}, "a boolean; last byte: 0xC3"},
{{0x81, 0xCA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a float; last byte: 0xCA"},
{{0x81, 0xCB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a float; last byte: 0xCB"},
{{0x81, 0xC4, 0x00, 0x01}, "a bin; last byte: 0xC4"},
{{0x81, 0xC5, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC5"},
{{0x81, 0xC6, 0x00, 0x00, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC6"},
{{0x81, 0xC7, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC7"},
{{0x81, 0xC8, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC8"},
{{0x81, 0xC9, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC9"},
{{0x81, 0xD4, 0x01, 0x00, 0x01}, "an ext; last byte: 0xD4"},
{{0x81, 0xD5, 0x01, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD5"},
{{0x81, 0xD6, 0x01, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD6"},
{{0x81, 0xD7, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD7"},
{{0x81, 0xD8, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD8"},
{{0x81, 0xCC, 0x01, 0x01}, "an integer; last byte: 0xCC"},
{{0x81, 0xCD, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCD"},
{{0x81, 0xCE, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCE"},
{{0x81, 0xCF, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCF"},
{{0x81, 0xD0, 0x01, 0x01}, "an integer; last byte: 0xD0"},
{{0x81, 0xD1, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD1"},
{{0x81, 0xD2, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD2"},
{{0x81, 0xD3, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD3"},
{{0x81, 0x00, 0x01}, "an integer; last byte: 0x00"},
{{0x81, 0x7F, 0x01}, "an integer; last byte: 0x7F"},
{{0x81, 0xE0, 0x01}, "an integer; last byte: 0xE0"},
{{0x81, 0x80, 0x01}, "a map; last byte: 0x80"},
{{0x81, 0x8F, 0x01}, "a map; last byte: 0x8F"},
{{0x81, 0xDE, 0x00, 0x00, 0x01}, "a map; last byte: 0xDE"},
{{0x81, 0xDF, 0x00, 0x00, 0x00, 0x00, 0x01}, "a map; last byte: 0xDF"},
{{0x81, 0x90, 0x01}, "an array; last byte: 0x90"},
{{0x81, 0x9F, 0x01}, "an array; last byte: 0x9F"},
{{0x81, 0xDC, 0x00, 0x00, 0x01}, "an array; last byte: 0xDC"},
{{0x81, 0xDD, 0x00, 0x00, 0x00, 0x00, 0x01}, "an array; last byte: 0xDD"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_msgpack(c.first, true, false).is_discarded());
}
json _;
// the unused byte 0xC1 is still reported as a malformed string
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xC1, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xC1", json::parse_error&);
// a missing key is still reported as the end of input
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
{
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
+3 -3
View File
@@ -1018,7 +1018,7 @@ TEST_CASE("regression tests 1")
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x98", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an array; last byte: 0x98", json::parse_error&);
// related test case: nonempty UTF-8 string (indefinite length)
std::vector<uint8_t> const vec1 {0x7f, 0x61, 0x61};
@@ -1065,7 +1065,7 @@ TEST_CASE("regression tests 1")
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
// related test case: double-precision
std::vector<uint8_t> const vec2
@@ -1077,7 +1077,7 @@ TEST_CASE("regression tests 1")
0x96, 0x96, 0xb4, 0xb4, 0xfa, 0x94, 0x94, 0x61,
0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0xfb
};
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
}
SECTION("issue #452 - Heap-buffer-overflow (OSS-Fuzz issue 585)")