mirror of
https://github.com/nlohmann/json.git
synced 2026-09-28 02:30:32 +00:00
Compare commits
4
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a3c2b0897e | ||
|
|
cfe7c9e732 | ||
|
|
f37492a6d6 | ||
|
|
b69794bd80 |
@@ -363,7 +363,7 @@ std::cout << j_string << " == " << serialized_string << std::endl;
|
||||
|
||||
[`.dump()`](https://json.nlohmann.me/api/basic_json/dump/) returns the originally stored string value.
|
||||
|
||||
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
|
||||
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
|
||||
|
||||
#### To/from streams (e.g., files, string streams)
|
||||
|
||||
@@ -1915,7 +1915,7 @@ The library supports **Unicode input** as follows:
|
||||
- [Unicode noncharacters](https://www.unicode.org/faq/private_use.html#nonchar1) will not be replaced by the library.
|
||||
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
|
||||
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
|
||||
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
|
||||
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
|
||||
- To store wide strings (e.g., `std::wstring`), you need to convert them to a UTF-8 encoded `std::string` before, see [an example](https://json.nlohmann.me/home/faq/#wide-string-handling).
|
||||
|
||||
### Comments in JSON
|
||||
|
||||
@@ -25,15 +25,10 @@ and `ensure_ascii` parameters.
|
||||
result consists of ASCII characters only.
|
||||
|
||||
`error_handler` (in)
|
||||
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md)):
|
||||
|
||||
- `strict`: throw a [`type_error`](../../home/exceptions.md#type-errors) exception in case a decoding error occurs
|
||||
(default),
|
||||
- `replace`: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER),
|
||||
- `ignore`: ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the output unchanged,
|
||||
and invalid bytes are dropped, and
|
||||
- `keep`: keep invalid UTF-8 sequences during serialization; all bytes are copied to the output unchanged, so the
|
||||
result is not valid UTF-8.
|
||||
: how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
|
||||
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
|
||||
with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
|
||||
output unchanged, and invalid bytes are dropped)).
|
||||
|
||||
## Return value
|
||||
|
||||
@@ -99,4 +94,3 @@ Binary values are serialized as an object containing two keys:
|
||||
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
|
||||
- Error handlers added in version 3.4.0.
|
||||
- Serialization of binary values added in version 3.8.0.
|
||||
- Error handler value `keep` added in version 3.13.0.
|
||||
|
||||
@@ -4,13 +4,12 @@
|
||||
enum class error_handler_t {
|
||||
strict,
|
||||
replace,
|
||||
ignore,
|
||||
keep
|
||||
ignore
|
||||
};
|
||||
```
|
||||
|
||||
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
|
||||
`basic_json` value. Four values are differentiated:
|
||||
`basic_json` value. Three values are differentiated:
|
||||
|
||||
strict
|
||||
: throw a `type_error` exception in case of invalid UTF-8
|
||||
@@ -21,12 +20,6 @@ replace
|
||||
ignore
|
||||
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
|
||||
|
||||
keep
|
||||
: keep invalid UTF-8 sequences; all bytes are copied to the output unchanged. Valid characters are still escaped as
|
||||
usual (e.g., `"`, `\\`, and control characters), so the result has valid JSON syntax, but it is not valid UTF-8.
|
||||
In particular, [`parse`](parse.md) rejects it, and with `ensure_ascii` set to `true`, the invalid bytes are the
|
||||
only non-ASCII bytes of the output.
|
||||
|
||||
## Examples
|
||||
|
||||
??? example
|
||||
@@ -47,4 +40,3 @@ keep
|
||||
## Version history
|
||||
|
||||
- Added in version 3.4.0.
|
||||
- Added value `keep` in version 3.13.0.
|
||||
|
||||
@@ -1,4 +1,3 @@
|
||||
#include <iomanip>
|
||||
#include <iostream>
|
||||
#include <nlohmann/json.hpp>
|
||||
|
||||
@@ -22,12 +21,4 @@ int main()
|
||||
<< "\nstring with ignored invalid characters: "
|
||||
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
|
||||
<< '\n';
|
||||
|
||||
// the invalid byte is kept; print the result byte-wise to make it visible
|
||||
std::cout << "string with kept invalid characters:";
|
||||
for (const unsigned char c : j_invalid.dump(-1, ' ', false, json::error_handler_t::keep))
|
||||
{
|
||||
std::cout << ' ' << std::hex << std::setw(2) << std::setfill('0') << static_cast<int>(c);
|
||||
}
|
||||
std::cout << '\n';
|
||||
}
|
||||
|
||||
@@ -1,4 +1,3 @@
|
||||
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
|
||||
string with replaced invalid characters: "ä�ü"
|
||||
string with ignored invalid characters: "äü"
|
||||
string with kept invalid characters: 22 c3 a4 a9 c3 bc 22
|
||||
|
||||
@@ -64,7 +64,6 @@ serialization fails by default. The fourth argument of `dump` selects an
|
||||
- `strict` (default) — throw a [`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) exception.
|
||||
- `replace` — replace invalid bytes with the Unicode replacement character U+FFFD (`�`).
|
||||
- `ignore` — silently drop invalid bytes.
|
||||
- `keep` — copy invalid bytes to the output unchanged; the result is not valid UTF-8.
|
||||
|
||||
??? example
|
||||
|
||||
|
||||
@@ -755,7 +755,6 @@ The `dump()` function only works with UTF-8 encoded strings; that is, if you ass
|
||||
- Pass an error handler as last parameter to the `dump()` function to avoid this exception:
|
||||
- `json::error_handler_t::replace` will replace invalid bytes sequences with `U+FFFD`
|
||||
- `json::error_handler_t::ignore` will silently ignore invalid byte sequences
|
||||
- `json::error_handler_t::keep` will copy invalid byte sequences to the output unchanged
|
||||
|
||||
### json.exception.type_error.317
|
||||
|
||||
|
||||
@@ -85,7 +85,7 @@ The library supports **Unicode input** as follows:
|
||||
- The library will not replace [Unicode noncharacters](http://www.unicode.org/faq/private_use.html#nonchar1).
|
||||
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
|
||||
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
|
||||
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
|
||||
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
|
||||
|
||||
In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding.
|
||||
|
||||
|
||||
@@ -206,7 +206,6 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
|
||||
: ia(std::move(adapter))
|
||||
, ignore_comments(ignore_comments_)
|
||||
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
|
||||
, discard_number_values(discard_number_values_)
|
||||
{}
|
||||
|
||||
@@ -222,8 +221,7 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
// locales
|
||||
/////////////////////
|
||||
|
||||
/// return the locale-dependent decimal point
|
||||
JSON_HEDLEY_PURE
|
||||
/// return the decimal point of the current locale
|
||||
static char get_decimal_point() noexcept
|
||||
{
|
||||
const auto* loc = localeconv();
|
||||
@@ -1092,9 +1090,10 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
token_type::value_float if number could be successfully scanned,
|
||||
token_type::parse_error otherwise
|
||||
|
||||
@note The scanner is independent of the current locale. Internally, the
|
||||
locale's decimal point is used instead of `.` to work with the
|
||||
locale-dependent converters.
|
||||
@note The scanner is independent of the current locale: token_buffer
|
||||
always holds `.`. Only the std::strtod fallback of convert_number()
|
||||
depends on the locale, and it looks up the decimal point right
|
||||
before converting (see convert_float_locale_aware()).
|
||||
*/
|
||||
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
|
||||
{
|
||||
@@ -1183,7 +1182,7 @@ scan_number_zero:
|
||||
{
|
||||
case '.':
|
||||
{
|
||||
add(decimal_point_char);
|
||||
add(current);
|
||||
decimal_point_position = token_buffer.size() - 1;
|
||||
goto scan_number_decimal1;
|
||||
}
|
||||
@@ -1220,7 +1219,7 @@ scan_number_any1:
|
||||
|
||||
case '.':
|
||||
{
|
||||
add(decimal_point_char);
|
||||
add(current);
|
||||
decimal_point_position = token_buffer.size() - 1;
|
||||
goto scan_number_decimal1;
|
||||
}
|
||||
@@ -1462,9 +1461,9 @@ scan_number_done:
|
||||
|
||||
// Only a number below 1 can carry further insignificant zeros, and only
|
||||
// while the count stays at the limit does removing them change the
|
||||
// answer - so this loop is skipped for all but a few tokens. Note
|
||||
// token_buffer holds the locale's decimal point, so the fraction is
|
||||
// located through decimal_point_position rather than by searching '.'.
|
||||
// answer - so this loop is skipped for all but a few tokens. The
|
||||
// fraction is located through decimal_point_position rather than by
|
||||
// searching '.'.
|
||||
if (lead_zero != 0)
|
||||
{
|
||||
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
|
||||
@@ -1482,8 +1481,8 @@ scan_number_done:
|
||||
@brief convert the number text in token_buffer to its value and token type
|
||||
|
||||
The digit sequence in token_buffer has already been validated (by the
|
||||
scan_number() state machine or by the contiguous fast path) and holds the
|
||||
locale decimal point in place of '.'. Integers are parsed first and fall
|
||||
scan_number() state machine or by the contiguous fast path) and holds '.'
|
||||
as decimal point, independent of the locale. Integers are parsed first and fall
|
||||
back to floating point on overflow. This is shared so both scanners produce
|
||||
identical results.
|
||||
|
||||
@@ -1563,7 +1562,7 @@ scan_number_done:
|
||||
// integer conversion above overflowed. Prefer std::from_chars
|
||||
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
|
||||
// otherwise the exact Clinger fast path (double only); otherwise the
|
||||
// locale-aware strtof/strtod.
|
||||
// locale-aware strtof/strtod/strtold.
|
||||
if (parse_float_from_chars(num_begin, num_end, value_float))
|
||||
{
|
||||
return token_type::value_float;
|
||||
@@ -1572,26 +1571,75 @@ scan_number_done:
|
||||
// extra pass over the token's bytes, which otherwise shows up on
|
||||
// high-precision inputs such as canada.json
|
||||
if (mantissa_fits_clinger(mantissa_end)
|
||||
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
|
||||
&& parse_float_fast(num_begin, num_end, value_float))
|
||||
{
|
||||
return token_type::value_float;
|
||||
}
|
||||
|
||||
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
|
||||
strtof(value_float, token_buffer.data(), &endptr);
|
||||
|
||||
// we checked the number format before
|
||||
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
|
||||
|
||||
convert_float_locale_aware();
|
||||
return token_type::value_float;
|
||||
}
|
||||
|
||||
/*!
|
||||
@brief convert the float in token_buffer with strtof/strtod/strtold
|
||||
|
||||
These functions expect the decimal point of the *current* locale, so it is
|
||||
looked up right before the conversion instead of once when the lexer is
|
||||
constructed: a locale change in between (by a parser callback, a SAX
|
||||
handler, or another thread) must not truncate the value (#5198). The
|
||||
token has been validated before, so if the conversion stops early and the
|
||||
decimal point changed in the meantime, the locale changed between the
|
||||
lookup and the call, and the conversion is repeated with the new decimal
|
||||
point. If the decimal point did not change, a retry cannot succeed: the
|
||||
locale's decimal point is not a single character (e.g., the two-byte
|
||||
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
|
||||
The value strtod parsed up to that point is kept, as before this change.
|
||||
|
||||
Note that changing the locale in another thread *while* strtod runs is
|
||||
undefined behavior of the C library, which this function cannot prevent.
|
||||
*/
|
||||
void convert_float_locale_aware()
|
||||
{
|
||||
const bool has_dot = decimal_point_position != std::string::npos;
|
||||
char decimal_point = get_decimal_point();
|
||||
for (;;)
|
||||
{
|
||||
const bool substitute = has_dot && decimal_point != '.';
|
||||
if (substitute)
|
||||
{
|
||||
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
|
||||
}
|
||||
|
||||
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
|
||||
strtof(value_float, token_buffer.data(), &endptr);
|
||||
|
||||
if (substitute)
|
||||
{
|
||||
// get_string() hands the token to the SAX interface with '.'
|
||||
token_buffer[decimal_point_position] = '.';
|
||||
}
|
||||
|
||||
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
|
||||
{
|
||||
return;
|
||||
}
|
||||
|
||||
// retry only if the locale changed; otherwise, this would loop forever
|
||||
const char current_decimal_point = get_decimal_point();
|
||||
if (current_decimal_point == decimal_point)
|
||||
{
|
||||
return;
|
||||
}
|
||||
decimal_point = current_decimal_point;
|
||||
}
|
||||
}
|
||||
|
||||
/*!
|
||||
@brief contiguous fast path for scanning a number
|
||||
|
||||
Parses the whole number token straight from the input buffer, avoiding the
|
||||
per-character get()/add() of scan_number(). On success it fills token_buffer
|
||||
(with the locale decimal point substituted, as scan_number() does) and
|
||||
(as scan_number() does) and
|
||||
returns the token type. On anything it does not fully recognize as a
|
||||
well-formed number it makes no state change and returns
|
||||
token_type::uninitialized, so the caller falls back to scan_number(), which
|
||||
@@ -1707,16 +1755,11 @@ scan_number_done:
|
||||
}
|
||||
#endif
|
||||
|
||||
// materialize the token exactly as scan_number() would, substituting the
|
||||
// locale decimal point so convert_number()'s strtof fallback stays valid.
|
||||
// reset() already cleared token_buffer, so append() fills it (assign() is
|
||||
// avoided because custom string_t types need not provide it)
|
||||
// materialize the token exactly as scan_number() would. reset() already
|
||||
// cleared token_buffer, so append() fills it (assign() is avoided
|
||||
// because custom string_t types need not provide it)
|
||||
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
|
||||
if (dot_index != std::string::npos)
|
||||
{
|
||||
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
|
||||
decimal_point_position = dot_index;
|
||||
}
|
||||
decimal_point_position = dot_index;
|
||||
|
||||
ia.bulk_skip(len - 1);
|
||||
position.chars_read_total += (len - 1);
|
||||
@@ -1983,11 +2026,7 @@ scan_number_done:
|
||||
/// return current string value (implicitly resets the token; useful only once)
|
||||
string_t& get_string()
|
||||
{
|
||||
// translate decimal points from locale back to '.' (#4084)
|
||||
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
|
||||
{
|
||||
token_buffer[decimal_point_position] = '.';
|
||||
}
|
||||
// a number token holds '.' regardless of the locale (#4084)
|
||||
return token_buffer;
|
||||
}
|
||||
|
||||
@@ -2283,9 +2322,7 @@ scan_number_done:
|
||||
number_unsigned_t value_unsigned = 0;
|
||||
number_float_t value_float = 0;
|
||||
|
||||
/// the decimal point
|
||||
const char_int_type decimal_point_char = '.';
|
||||
/// the position of the decimal point in the input
|
||||
/// the position of the decimal point in token_buffer
|
||||
std::size_t decimal_point_position = std::string::npos;
|
||||
|
||||
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
|
||||
|
||||
@@ -118,14 +118,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
|
||||
long double keep the std::strtof/std::strtold paths (see the templated overload
|
||||
below).
|
||||
|
||||
@param[in] first pointer to the first character of the number
|
||||
@param[in] last pointer past the last character
|
||||
@param[in] decimal_point the (locale-dependent) decimal point character
|
||||
@param[out] out the parsed value on success
|
||||
@param[in] first pointer to the first character of the number
|
||||
@param[in] last pointer past the last character
|
||||
@param[out] out the parsed value on success
|
||||
@return true if the value was parsed exactly; false to fall back to strtod
|
||||
*/
|
||||
template<typename DecimalPointType>
|
||||
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
|
||||
inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
|
||||
{
|
||||
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
|
||||
// Clinger's fast path is only exact when double operations are evaluated in
|
||||
@@ -136,7 +134,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
|
||||
// std::from_chars / std::strtod path.
|
||||
static_cast<void>(first);
|
||||
static_cast<void>(last);
|
||||
static_cast<void>(decimal_point);
|
||||
static_cast<void>(out);
|
||||
return false;
|
||||
#else
|
||||
@@ -175,7 +172,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
|
||||
++num_digits;
|
||||
fractional_digits += static_cast<int>(seen_dot);
|
||||
}
|
||||
else if (static_cast<DecimalPointType>(c) == decimal_point)
|
||||
else if (c == '.')
|
||||
{
|
||||
if (JSON_HEDLEY_UNLIKELY(seen_dot))
|
||||
{
|
||||
@@ -260,8 +257,8 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
|
||||
}
|
||||
|
||||
/// fast float path is only exact for `double`; decline for float/long double
|
||||
template<typename DecimalPointType, typename FloatType>
|
||||
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
|
||||
template<typename FloatType>
|
||||
bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
|
||||
{
|
||||
return false;
|
||||
}
|
||||
@@ -273,9 +270,7 @@ std::from_chars is locale-independent, correctly rounded, and - via the
|
||||
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
|
||||
over the whole value range (not just the Clinger subset). It is used only when
|
||||
__cpp_lib_to_chars indicates full floating-point support and only when it
|
||||
consumes the entire token ([first, last)); a partial parse means the buffer
|
||||
uses a non-'.' locale decimal point, in which case the caller falls back to the
|
||||
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
|
||||
consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
|
||||
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
|
||||
expects (side-stepping the P4168 divergence between implementations).
|
||||
|
||||
|
||||
@@ -48,8 +48,7 @@ enum class error_handler_t
|
||||
{
|
||||
strict, ///< throw a type_error exception in case of invalid UTF-8
|
||||
replace, ///< replace invalid UTF-8 sequences with U+FFFD
|
||||
ignore, ///< ignore invalid UTF-8 sequences
|
||||
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
|
||||
ignore ///< ignore invalid UTF-8 sequences
|
||||
};
|
||||
|
||||
template<typename BasicJsonType>
|
||||
@@ -1020,47 +1019,6 @@ class serializer
|
||||
break;
|
||||
}
|
||||
|
||||
case error_handler_t::keep:
|
||||
{
|
||||
// drop whatever the incomplete sequence left in
|
||||
// the buffer (only copied if !EnsureAscii) and copy
|
||||
// the ill-formed bytes from the input instead
|
||||
bytes = bytes_after_last_accept;
|
||||
|
||||
if (undumped_chars > 0)
|
||||
{
|
||||
// the pending bytes of the incomplete sequence
|
||||
// are ill-formed; the current byte may be OK for
|
||||
// itself, so we would like to read it again
|
||||
for (std::size_t j = i - undumped_chars; j < i; ++j)
|
||||
{
|
||||
string_buffer[bytes++] = s[j];
|
||||
}
|
||||
--i;
|
||||
}
|
||||
else
|
||||
{
|
||||
// the current byte cannot start any sequence
|
||||
string_buffer[bytes++] = s[i];
|
||||
}
|
||||
|
||||
// write buffer and reset index; there must be 13 bytes
|
||||
// left, as this is the maximal number of bytes to be
|
||||
// written ("\uxxxx\uxxxx\0") for one code point
|
||||
if (string_buffer.size() - bytes < 13)
|
||||
{
|
||||
put_buffer(string_buffer, bytes);
|
||||
bytes = 0;
|
||||
}
|
||||
|
||||
bytes_after_last_accept = bytes;
|
||||
undumped_chars = 0;
|
||||
|
||||
// continue processing the string
|
||||
state = UTF8_ACCEPT;
|
||||
break;
|
||||
}
|
||||
|
||||
default: // LCOV_EXCL_LINE
|
||||
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
|
||||
}
|
||||
@@ -1106,15 +1064,6 @@ class serializer
|
||||
break;
|
||||
}
|
||||
|
||||
case error_handler_t::keep:
|
||||
{
|
||||
// write all accepted bytes
|
||||
put_buffer(string_buffer, bytes_after_last_accept);
|
||||
// copy the bytes of the incomplete sequence unchanged
|
||||
put_string(s, s.size() - undumped_chars, s.size());
|
||||
break;
|
||||
}
|
||||
|
||||
case error_handler_t::replace:
|
||||
{
|
||||
// write all accepted bytes
|
||||
|
||||
@@ -8605,14 +8605,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
|
||||
long double keep the std::strtof/std::strtold paths (see the templated overload
|
||||
below).
|
||||
|
||||
@param[in] first pointer to the first character of the number
|
||||
@param[in] last pointer past the last character
|
||||
@param[in] decimal_point the (locale-dependent) decimal point character
|
||||
@param[out] out the parsed value on success
|
||||
@param[in] first pointer to the first character of the number
|
||||
@param[in] last pointer past the last character
|
||||
@param[out] out the parsed value on success
|
||||
@return true if the value was parsed exactly; false to fall back to strtod
|
||||
*/
|
||||
template<typename DecimalPointType>
|
||||
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
|
||||
inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
|
||||
{
|
||||
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
|
||||
// Clinger's fast path is only exact when double operations are evaluated in
|
||||
@@ -8623,7 +8621,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
|
||||
// std::from_chars / std::strtod path.
|
||||
static_cast<void>(first);
|
||||
static_cast<void>(last);
|
||||
static_cast<void>(decimal_point);
|
||||
static_cast<void>(out);
|
||||
return false;
|
||||
#else
|
||||
@@ -8662,7 +8659,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
|
||||
++num_digits;
|
||||
fractional_digits += static_cast<int>(seen_dot);
|
||||
}
|
||||
else if (static_cast<DecimalPointType>(c) == decimal_point)
|
||||
else if (c == '.')
|
||||
{
|
||||
if (JSON_HEDLEY_UNLIKELY(seen_dot))
|
||||
{
|
||||
@@ -8747,8 +8744,8 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
|
||||
}
|
||||
|
||||
/// fast float path is only exact for `double`; decline for float/long double
|
||||
template<typename DecimalPointType, typename FloatType>
|
||||
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
|
||||
template<typename FloatType>
|
||||
bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
|
||||
{
|
||||
return false;
|
||||
}
|
||||
@@ -8760,9 +8757,7 @@ std::from_chars is locale-independent, correctly rounded, and - via the
|
||||
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
|
||||
over the whole value range (not just the Clinger subset). It is used only when
|
||||
__cpp_lib_to_chars indicates full floating-point support and only when it
|
||||
consumes the entire token ([first, last)); a partial parse means the buffer
|
||||
uses a non-'.' locale decimal point, in which case the caller falls back to the
|
||||
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
|
||||
consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
|
||||
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
|
||||
expects (side-stepping the P4168 divergence between implementations).
|
||||
|
||||
@@ -9303,7 +9298,6 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
|
||||
: ia(std::move(adapter))
|
||||
, ignore_comments(ignore_comments_)
|
||||
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
|
||||
, discard_number_values(discard_number_values_)
|
||||
{}
|
||||
|
||||
@@ -9319,8 +9313,7 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
// locales
|
||||
/////////////////////
|
||||
|
||||
/// return the locale-dependent decimal point
|
||||
JSON_HEDLEY_PURE
|
||||
/// return the decimal point of the current locale
|
||||
static char get_decimal_point() noexcept
|
||||
{
|
||||
const auto* loc = localeconv();
|
||||
@@ -10189,9 +10182,10 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
token_type::value_float if number could be successfully scanned,
|
||||
token_type::parse_error otherwise
|
||||
|
||||
@note The scanner is independent of the current locale. Internally, the
|
||||
locale's decimal point is used instead of `.` to work with the
|
||||
locale-dependent converters.
|
||||
@note The scanner is independent of the current locale: token_buffer
|
||||
always holds `.`. Only the std::strtod fallback of convert_number()
|
||||
depends on the locale, and it looks up the decimal point right
|
||||
before converting (see convert_float_locale_aware()).
|
||||
*/
|
||||
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
|
||||
{
|
||||
@@ -10280,7 +10274,7 @@ scan_number_zero:
|
||||
{
|
||||
case '.':
|
||||
{
|
||||
add(decimal_point_char);
|
||||
add(current);
|
||||
decimal_point_position = token_buffer.size() - 1;
|
||||
goto scan_number_decimal1;
|
||||
}
|
||||
@@ -10317,7 +10311,7 @@ scan_number_any1:
|
||||
|
||||
case '.':
|
||||
{
|
||||
add(decimal_point_char);
|
||||
add(current);
|
||||
decimal_point_position = token_buffer.size() - 1;
|
||||
goto scan_number_decimal1;
|
||||
}
|
||||
@@ -10559,9 +10553,9 @@ scan_number_done:
|
||||
|
||||
// Only a number below 1 can carry further insignificant zeros, and only
|
||||
// while the count stays at the limit does removing them change the
|
||||
// answer - so this loop is skipped for all but a few tokens. Note
|
||||
// token_buffer holds the locale's decimal point, so the fraction is
|
||||
// located through decimal_point_position rather than by searching '.'.
|
||||
// answer - so this loop is skipped for all but a few tokens. The
|
||||
// fraction is located through decimal_point_position rather than by
|
||||
// searching '.'.
|
||||
if (lead_zero != 0)
|
||||
{
|
||||
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
|
||||
@@ -10579,8 +10573,8 @@ scan_number_done:
|
||||
@brief convert the number text in token_buffer to its value and token type
|
||||
|
||||
The digit sequence in token_buffer has already been validated (by the
|
||||
scan_number() state machine or by the contiguous fast path) and holds the
|
||||
locale decimal point in place of '.'. Integers are parsed first and fall
|
||||
scan_number() state machine or by the contiguous fast path) and holds '.'
|
||||
as decimal point, independent of the locale. Integers are parsed first and fall
|
||||
back to floating point on overflow. This is shared so both scanners produce
|
||||
identical results.
|
||||
|
||||
@@ -10660,7 +10654,7 @@ scan_number_done:
|
||||
// integer conversion above overflowed. Prefer std::from_chars
|
||||
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
|
||||
// otherwise the exact Clinger fast path (double only); otherwise the
|
||||
// locale-aware strtof/strtod.
|
||||
// locale-aware strtof/strtod/strtold.
|
||||
if (parse_float_from_chars(num_begin, num_end, value_float))
|
||||
{
|
||||
return token_type::value_float;
|
||||
@@ -10669,26 +10663,75 @@ scan_number_done:
|
||||
// extra pass over the token's bytes, which otherwise shows up on
|
||||
// high-precision inputs such as canada.json
|
||||
if (mantissa_fits_clinger(mantissa_end)
|
||||
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
|
||||
&& parse_float_fast(num_begin, num_end, value_float))
|
||||
{
|
||||
return token_type::value_float;
|
||||
}
|
||||
|
||||
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
|
||||
strtof(value_float, token_buffer.data(), &endptr);
|
||||
|
||||
// we checked the number format before
|
||||
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
|
||||
|
||||
convert_float_locale_aware();
|
||||
return token_type::value_float;
|
||||
}
|
||||
|
||||
/*!
|
||||
@brief convert the float in token_buffer with strtof/strtod/strtold
|
||||
|
||||
These functions expect the decimal point of the *current* locale, so it is
|
||||
looked up right before the conversion instead of once when the lexer is
|
||||
constructed: a locale change in between (by a parser callback, a SAX
|
||||
handler, or another thread) must not truncate the value (#5198). The
|
||||
token has been validated before, so if the conversion stops early and the
|
||||
decimal point changed in the meantime, the locale changed between the
|
||||
lookup and the call, and the conversion is repeated with the new decimal
|
||||
point. If the decimal point did not change, a retry cannot succeed: the
|
||||
locale's decimal point is not a single character (e.g., the two-byte
|
||||
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
|
||||
The value strtod parsed up to that point is kept, as before this change.
|
||||
|
||||
Note that changing the locale in another thread *while* strtod runs is
|
||||
undefined behavior of the C library, which this function cannot prevent.
|
||||
*/
|
||||
void convert_float_locale_aware()
|
||||
{
|
||||
const bool has_dot = decimal_point_position != std::string::npos;
|
||||
char decimal_point = get_decimal_point();
|
||||
for (;;)
|
||||
{
|
||||
const bool substitute = has_dot && decimal_point != '.';
|
||||
if (substitute)
|
||||
{
|
||||
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
|
||||
}
|
||||
|
||||
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
|
||||
strtof(value_float, token_buffer.data(), &endptr);
|
||||
|
||||
if (substitute)
|
||||
{
|
||||
// get_string() hands the token to the SAX interface with '.'
|
||||
token_buffer[decimal_point_position] = '.';
|
||||
}
|
||||
|
||||
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
|
||||
{
|
||||
return;
|
||||
}
|
||||
|
||||
// retry only if the locale changed; otherwise, this would loop forever
|
||||
const char current_decimal_point = get_decimal_point();
|
||||
if (current_decimal_point == decimal_point)
|
||||
{
|
||||
return;
|
||||
}
|
||||
decimal_point = current_decimal_point;
|
||||
}
|
||||
}
|
||||
|
||||
/*!
|
||||
@brief contiguous fast path for scanning a number
|
||||
|
||||
Parses the whole number token straight from the input buffer, avoiding the
|
||||
per-character get()/add() of scan_number(). On success it fills token_buffer
|
||||
(with the locale decimal point substituted, as scan_number() does) and
|
||||
(as scan_number() does) and
|
||||
returns the token type. On anything it does not fully recognize as a
|
||||
well-formed number it makes no state change and returns
|
||||
token_type::uninitialized, so the caller falls back to scan_number(), which
|
||||
@@ -10804,16 +10847,11 @@ scan_number_done:
|
||||
}
|
||||
#endif
|
||||
|
||||
// materialize the token exactly as scan_number() would, substituting the
|
||||
// locale decimal point so convert_number()'s strtof fallback stays valid.
|
||||
// reset() already cleared token_buffer, so append() fills it (assign() is
|
||||
// avoided because custom string_t types need not provide it)
|
||||
// materialize the token exactly as scan_number() would. reset() already
|
||||
// cleared token_buffer, so append() fills it (assign() is avoided
|
||||
// because custom string_t types need not provide it)
|
||||
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
|
||||
if (dot_index != std::string::npos)
|
||||
{
|
||||
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
|
||||
decimal_point_position = dot_index;
|
||||
}
|
||||
decimal_point_position = dot_index;
|
||||
|
||||
ia.bulk_skip(len - 1);
|
||||
position.chars_read_total += (len - 1);
|
||||
@@ -11080,11 +11118,7 @@ scan_number_done:
|
||||
/// return current string value (implicitly resets the token; useful only once)
|
||||
string_t& get_string()
|
||||
{
|
||||
// translate decimal points from locale back to '.' (#4084)
|
||||
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
|
||||
{
|
||||
token_buffer[decimal_point_position] = '.';
|
||||
}
|
||||
// a number token holds '.' regardless of the locale (#4084)
|
||||
return token_buffer;
|
||||
}
|
||||
|
||||
@@ -11380,9 +11414,7 @@ scan_number_done:
|
||||
number_unsigned_t value_unsigned = 0;
|
||||
number_float_t value_float = 0;
|
||||
|
||||
/// the decimal point
|
||||
const char_int_type decimal_point_char = '.';
|
||||
/// the position of the decimal point in the input
|
||||
/// the position of the decimal point in token_buffer
|
||||
std::size_t decimal_point_position = std::string::npos;
|
||||
|
||||
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
|
||||
@@ -23931,8 +23963,7 @@ enum class error_handler_t
|
||||
{
|
||||
strict, ///< throw a type_error exception in case of invalid UTF-8
|
||||
replace, ///< replace invalid UTF-8 sequences with U+FFFD
|
||||
ignore, ///< ignore invalid UTF-8 sequences
|
||||
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
|
||||
ignore ///< ignore invalid UTF-8 sequences
|
||||
};
|
||||
|
||||
template<typename BasicJsonType>
|
||||
@@ -24903,47 +24934,6 @@ class serializer
|
||||
break;
|
||||
}
|
||||
|
||||
case error_handler_t::keep:
|
||||
{
|
||||
// drop whatever the incomplete sequence left in
|
||||
// the buffer (only copied if !EnsureAscii) and copy
|
||||
// the ill-formed bytes from the input instead
|
||||
bytes = bytes_after_last_accept;
|
||||
|
||||
if (undumped_chars > 0)
|
||||
{
|
||||
// the pending bytes of the incomplete sequence
|
||||
// are ill-formed; the current byte may be OK for
|
||||
// itself, so we would like to read it again
|
||||
for (std::size_t j = i - undumped_chars; j < i; ++j)
|
||||
{
|
||||
string_buffer[bytes++] = s[j];
|
||||
}
|
||||
--i;
|
||||
}
|
||||
else
|
||||
{
|
||||
// the current byte cannot start any sequence
|
||||
string_buffer[bytes++] = s[i];
|
||||
}
|
||||
|
||||
// write buffer and reset index; there must be 13 bytes
|
||||
// left, as this is the maximal number of bytes to be
|
||||
// written ("\uxxxx\uxxxx\0") for one code point
|
||||
if (string_buffer.size() - bytes < 13)
|
||||
{
|
||||
put_buffer(string_buffer, bytes);
|
||||
bytes = 0;
|
||||
}
|
||||
|
||||
bytes_after_last_accept = bytes;
|
||||
undumped_chars = 0;
|
||||
|
||||
// continue processing the string
|
||||
state = UTF8_ACCEPT;
|
||||
break;
|
||||
}
|
||||
|
||||
default: // LCOV_EXCL_LINE
|
||||
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
|
||||
}
|
||||
@@ -24989,15 +24979,6 @@ class serializer
|
||||
break;
|
||||
}
|
||||
|
||||
case error_handler_t::keep:
|
||||
{
|
||||
// write all accepted bytes
|
||||
put_buffer(string_buffer, bytes_after_last_accept);
|
||||
// copy the bytes of the incomplete sequence unchanged
|
||||
put_string(s, s.size() - undumped_chars, s.size());
|
||||
break;
|
||||
}
|
||||
|
||||
case error_handler_t::replace:
|
||||
{
|
||||
// write all accepted bytes
|
||||
|
||||
@@ -666,7 +666,7 @@ TEST_CASE("parse_float_fast declines what it cannot convert exactly")
|
||||
// always safe: the caller then falls back to a slower, exact conversion.
|
||||
const auto fast = [](const std::string & s, double & out)
|
||||
{
|
||||
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), '.', out);
|
||||
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out);
|
||||
};
|
||||
double out = 0;
|
||||
|
||||
|
||||
@@ -12,7 +12,12 @@
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
|
||||
#include <array>
|
||||
#include <clocale>
|
||||
#include <map>
|
||||
#include <string>
|
||||
#include <utility>
|
||||
#include <vector>
|
||||
|
||||
struct ParserImpl final: public nlohmann::json_sax<json>
|
||||
{
|
||||
@@ -175,3 +180,208 @@ TEST_CASE("locale-dependent test (LC_NUMERIC=de_DE)")
|
||||
MESSAGE("locale de_DE is not usable");
|
||||
}
|
||||
}
|
||||
|
||||
namespace
|
||||
{
|
||||
// records the numbers of a flat array and switches LC_NUMERIC to the given
|
||||
// locale once the array opens - after the lexer was constructed, but before
|
||||
// any number in the array is lexed
|
||||
struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
|
||||
{
|
||||
explicit LocaleSwitchingSax(const char* switch_to)
|
||||
: locale_after_open(switch_to)
|
||||
{}
|
||||
|
||||
bool null() override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool boolean(bool /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool number_integer(json::number_integer_t /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool number_unsigned(json::number_unsigned_t /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool number_float(json::number_float_t val, const json::string_t& s) override
|
||||
{
|
||||
values.push_back(val);
|
||||
strings.push_back(s);
|
||||
return true;
|
||||
}
|
||||
bool string(json::string_t& /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool binary(json::binary_t& /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool start_object(std::size_t /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool key(json::string_t& /*val*/) override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool end_object() override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool start_array(std::size_t /*val*/) override
|
||||
{
|
||||
switched = std::setlocale(LC_NUMERIC, locale_after_open.c_str()) != nullptr;
|
||||
return true;
|
||||
}
|
||||
bool end_array() override
|
||||
{
|
||||
return true;
|
||||
}
|
||||
bool parse_error(std::size_t /*val*/, const std::string& /*val*/, const nlohmann::detail::exception& /*val*/) override
|
||||
{
|
||||
return false;
|
||||
}
|
||||
|
||||
std::string locale_after_open;
|
||||
bool switched = false;
|
||||
std::vector<json::number_float_t> values {}; // NOLINT(readability-redundant-member-init)
|
||||
std::vector<json::string_t> strings {}; // NOLINT(readability-redundant-member-init)
|
||||
};
|
||||
} // namespace
|
||||
|
||||
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
|
||||
{
|
||||
// The numbers are chosen so that the conversion also takes the strtod
|
||||
// fallback, which honors the locale that is current at conversion time:
|
||||
// too many significant digits for Clinger's fast path, an underflow that
|
||||
// std::from_chars rejects, and a plain value.
|
||||
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
|
||||
std::string text = "[";
|
||||
for (const auto& n : numbers)
|
||||
{
|
||||
text += (text.size() == 1 ? "" : ",") + n;
|
||||
}
|
||||
text += "]";
|
||||
|
||||
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
|
||||
|
||||
// reference values, parsed without a locale switch
|
||||
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
|
||||
const json expected = json::parse(text);
|
||||
const long_double_json expected_ld = long_double_json::parse(text);
|
||||
|
||||
const std::array<std::pair<const char*, const char*>, 2> transitions =
|
||||
{
|
||||
{
|
||||
{"C", "de_DE"},
|
||||
{"de_DE", "C"}
|
||||
}
|
||||
};
|
||||
|
||||
for (const auto& transition : transitions)
|
||||
{
|
||||
CAPTURE(transition.first);
|
||||
CAPTURE(transition.second);
|
||||
|
||||
if (std::setlocale(LC_NUMERIC, transition.first) == nullptr)
|
||||
{
|
||||
MESSAGE("locale is not usable");
|
||||
continue;
|
||||
}
|
||||
|
||||
// SAX parsing
|
||||
{
|
||||
LocaleSwitchingSax sax(transition.second);
|
||||
CHECK(json::sax_parse(text, &sax));
|
||||
if (sax.switched)
|
||||
{
|
||||
CHECK(sax.values == expected.get<std::vector<json::number_float_t>>());
|
||||
CHECK(sax.strings == numbers);
|
||||
}
|
||||
}
|
||||
|
||||
// DOM parsing with a callback
|
||||
{
|
||||
bool switched = false;
|
||||
const auto cb = [&](int /*depth*/, json::parse_event_t event, json& /*parsed*/)
|
||||
{
|
||||
if (event == json::parse_event_t::array_start)
|
||||
{
|
||||
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
|
||||
}
|
||||
return true;
|
||||
};
|
||||
const json j = json::parse(text, cb);
|
||||
if (switched)
|
||||
{
|
||||
CHECK(j == expected);
|
||||
}
|
||||
}
|
||||
|
||||
// a long double goes through std::strtold unless std::from_chars supports it
|
||||
{
|
||||
bool switched = false;
|
||||
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/)
|
||||
{
|
||||
if (event == long_double_json::parse_event_t::array_start)
|
||||
{
|
||||
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
|
||||
}
|
||||
return true;
|
||||
};
|
||||
const long_double_json j = long_double_json::parse(text, cb);
|
||||
if (switched)
|
||||
{
|
||||
CHECK(j == expected_ld);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
std::setlocale(LC_NUMERIC, "C");
|
||||
}
|
||||
|
||||
TEST_CASE("locale with a multi-byte decimal point")
|
||||
{
|
||||
// Some locales use a decimal point that is not a single character, e.g.
|
||||
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
|
||||
// substituted in place for '.', so the strtod fallback stops early. The
|
||||
// conversion must still terminate rather than retry forever.
|
||||
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
|
||||
bool tested = false;
|
||||
for (const char* name : names)
|
||||
{
|
||||
if (std::setlocale(LC_NUMERIC, name) == nullptr)
|
||||
{
|
||||
continue;
|
||||
}
|
||||
const std::string decimal_point = std::localeconv()->decimal_point;
|
||||
if (decimal_point.size() < 2)
|
||||
{
|
||||
continue;
|
||||
}
|
||||
CAPTURE(name);
|
||||
tested = true;
|
||||
|
||||
// too many significant digits for Clinger's fast path, and an underflow
|
||||
// that std::from_chars rejects: both reach the strtod fallback
|
||||
json j;
|
||||
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
|
||||
CHECK(j.is_array());
|
||||
CHECK(json::accept("3.14159265358979323846"));
|
||||
|
||||
// a value the locale-independent paths convert is not affected
|
||||
CHECK(json::parse("12.5") == 12.5);
|
||||
}
|
||||
if (!tested)
|
||||
{
|
||||
MESSAGE("no locale with a multi-byte decimal point is usable");
|
||||
}
|
||||
|
||||
std::setlocale(LC_NUMERIC, "C");
|
||||
}
|
||||
|
||||
@@ -766,15 +766,6 @@ TEST_CASE("regression tests 2")
|
||||
CHECK(j == k);
|
||||
}
|
||||
|
||||
SECTION("issue #4552 - UTF-8 invalid characters are not always ignored when dumping with error_handler_t::ignore")
|
||||
{
|
||||
json node;
|
||||
node["test"] = "test\334\005";
|
||||
CHECK(node.dump(-1, ' ', false, json::error_handler_t::ignore) == "{\"test\":\"test\\u0005\"}");
|
||||
CHECK(node.dump(-1, ' ', false, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
|
||||
CHECK(node.dump(-1, ' ', true, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
TEST_CASE("regression test - parser callback must not lose a duplicate key's prior value")
|
||||
|
||||
@@ -92,8 +92,6 @@ TEST_CASE("serialization")
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"äü\"");
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"ä\xEF\xBF\xBDü\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"\\u00e4\\ufffd\\u00fc\"");
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"ä\xA9ü\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\\u00e4\xA9\\u00fc\"");
|
||||
}
|
||||
|
||||
SECTION("invalid character (regression guard for shared UTF-8 decoder, see #5529)")
|
||||
@@ -116,8 +114,6 @@ TEST_CASE("serialization")
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123\"");
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd\"");
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xC2\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xC2\"");
|
||||
}
|
||||
|
||||
SECTION("unexpected character")
|
||||
@@ -130,39 +126,6 @@ TEST_CASE("serialization")
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123456\"");
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\x34\x35\x36\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd456\"");
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
|
||||
}
|
||||
|
||||
SECTION("keep: valid characters are still escaped")
|
||||
{
|
||||
// an invalid byte followed by characters that must be escaped
|
||||
const json j = "\xC2\"\\\n\xFF\x05";
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
|
||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
|
||||
}
|
||||
|
||||
SECTION("keep: truncated multibyte sequences")
|
||||
{
|
||||
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
|
||||
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
|
||||
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
|
||||
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
|
||||
}
|
||||
|
||||
SECTION("keep: long string with many invalid bytes")
|
||||
{
|
||||
// exceeds the internal string buffer several times
|
||||
std::string input;
|
||||
std::string expected = "\"";
|
||||
for (int i = 0; i < 2000; ++i)
|
||||
{
|
||||
input += "\xFF\xE2\x82\n\xC3\xA4";
|
||||
expected += "\xFF\xE2\x82\\n\xC3\xA4";
|
||||
}
|
||||
expected += "\"";
|
||||
const json j = input;
|
||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == expected);
|
||||
}
|
||||
|
||||
SECTION("U+FFFD Substitution of Maximal Subparts")
|
||||
|
||||
@@ -14,7 +14,6 @@
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
|
||||
#include <algorithm>
|
||||
#include <fstream>
|
||||
#include <sstream>
|
||||
#include <iostream>
|
||||
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
static std::string s_replaced2;
|
||||
static std::string s_replaced_ascii;
|
||||
static std::string s_replaced2_ascii;
|
||||
static std::string s_kept;
|
||||
static std::string s_kept2;
|
||||
static std::string s_kept_ascii;
|
||||
|
||||
// dumping with ignore/replace/keep must not throw in any case
|
||||
// dumping with ignore/replace must not throw in any case
|
||||
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
|
||||
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
|
||||
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
|
||||
|
||||
if (success_expected)
|
||||
{
|
||||
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
// all dumps should agree on the string
|
||||
CHECK(s_strict == s_ignored);
|
||||
CHECK(s_strict == s_replaced);
|
||||
CHECK(s_strict == s_kept);
|
||||
}
|
||||
else
|
||||
{
|
||||
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
|
||||
// check that replace string contains a replacement character
|
||||
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
|
||||
|
||||
// ignore drops the invalid bytes, keep copies them
|
||||
CHECK(s_ignored != s_kept);
|
||||
CHECK(s_ignored_ascii != s_kept_ascii);
|
||||
|
||||
// unless a byte needs escaping, keep copies the input unchanged
|
||||
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
|
||||
{
|
||||
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
|
||||
});
|
||||
if (!needs_escaping)
|
||||
{
|
||||
CHECK(s_kept == "\"" + json_string + "\"");
|
||||
}
|
||||
}
|
||||
|
||||
// check that prefix and suffix are preserved
|
||||
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
|
||||
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
|
||||
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
|
||||
CHECK(s_kept2.substr(1, 3) == "abc");
|
||||
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
|
||||
}
|
||||
|
||||
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
|
||||
|
||||
@@ -14,7 +14,6 @@
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
|
||||
#include <algorithm>
|
||||
#include <fstream>
|
||||
#include <sstream>
|
||||
#include <iostream>
|
||||
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
static std::string s_replaced2;
|
||||
static std::string s_replaced_ascii;
|
||||
static std::string s_replaced2_ascii;
|
||||
static std::string s_kept;
|
||||
static std::string s_kept2;
|
||||
static std::string s_kept_ascii;
|
||||
|
||||
// dumping with ignore/replace/keep must not throw in any case
|
||||
// dumping with ignore/replace must not throw in any case
|
||||
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
|
||||
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
|
||||
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
|
||||
|
||||
if (success_expected)
|
||||
{
|
||||
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
// all dumps should agree on the string
|
||||
CHECK(s_strict == s_ignored);
|
||||
CHECK(s_strict == s_replaced);
|
||||
CHECK(s_strict == s_kept);
|
||||
}
|
||||
else
|
||||
{
|
||||
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
|
||||
// check that replace string contains a replacement character
|
||||
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
|
||||
|
||||
// ignore drops the invalid bytes, keep copies them
|
||||
CHECK(s_ignored != s_kept);
|
||||
CHECK(s_ignored_ascii != s_kept_ascii);
|
||||
|
||||
// unless a byte needs escaping, keep copies the input unchanged
|
||||
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
|
||||
{
|
||||
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
|
||||
});
|
||||
if (!needs_escaping)
|
||||
{
|
||||
CHECK(s_kept == "\"" + json_string + "\"");
|
||||
}
|
||||
}
|
||||
|
||||
// check that prefix and suffix are preserved
|
||||
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
|
||||
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
|
||||
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
|
||||
CHECK(s_kept2.substr(1, 3) == "abc");
|
||||
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
|
||||
}
|
||||
|
||||
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
|
||||
|
||||
@@ -14,7 +14,6 @@
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
|
||||
#include <algorithm>
|
||||
#include <fstream>
|
||||
#include <sstream>
|
||||
#include <iostream>
|
||||
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
static std::string s_replaced2;
|
||||
static std::string s_replaced_ascii;
|
||||
static std::string s_replaced2_ascii;
|
||||
static std::string s_kept;
|
||||
static std::string s_kept2;
|
||||
static std::string s_kept_ascii;
|
||||
|
||||
// dumping with ignore/replace/keep must not throw in any case
|
||||
// dumping with ignore/replace must not throw in any case
|
||||
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
|
||||
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
|
||||
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
|
||||
|
||||
if (success_expected)
|
||||
{
|
||||
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
// all dumps should agree on the string
|
||||
CHECK(s_strict == s_ignored);
|
||||
CHECK(s_strict == s_replaced);
|
||||
CHECK(s_strict == s_kept);
|
||||
}
|
||||
else
|
||||
{
|
||||
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
|
||||
// check that replace string contains a replacement character
|
||||
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
|
||||
|
||||
// ignore drops the invalid bytes, keep copies them
|
||||
CHECK(s_ignored != s_kept);
|
||||
CHECK(s_ignored_ascii != s_kept_ascii);
|
||||
|
||||
// unless a byte needs escaping, keep copies the input unchanged
|
||||
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
|
||||
{
|
||||
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
|
||||
});
|
||||
if (!needs_escaping)
|
||||
{
|
||||
CHECK(s_kept == "\"" + json_string + "\"");
|
||||
}
|
||||
}
|
||||
|
||||
// check that prefix and suffix are preserved
|
||||
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
|
||||
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
|
||||
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
|
||||
CHECK(s_kept2.substr(1, 3) == "abc");
|
||||
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
|
||||
}
|
||||
|
||||
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
|
||||
|
||||
@@ -14,7 +14,6 @@
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
|
||||
#include <algorithm>
|
||||
#include <fstream>
|
||||
#include <sstream>
|
||||
#include <iostream>
|
||||
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
static std::string s_replaced2;
|
||||
static std::string s_replaced_ascii;
|
||||
static std::string s_replaced2_ascii;
|
||||
static std::string s_kept;
|
||||
static std::string s_kept2;
|
||||
static std::string s_kept_ascii;
|
||||
|
||||
// dumping with ignore/replace/keep must not throw in any case
|
||||
// dumping with ignore/replace must not throw in any case
|
||||
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
|
||||
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
|
||||
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
|
||||
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
|
||||
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
|
||||
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
|
||||
|
||||
if (success_expected)
|
||||
{
|
||||
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
// all dumps should agree on the string
|
||||
CHECK(s_strict == s_ignored);
|
||||
CHECK(s_strict == s_replaced);
|
||||
CHECK(s_strict == s_kept);
|
||||
}
|
||||
else
|
||||
{
|
||||
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
|
||||
// check that replace string contains a replacement character
|
||||
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
|
||||
|
||||
// ignore drops the invalid bytes, keep copies them
|
||||
CHECK(s_ignored != s_kept);
|
||||
CHECK(s_ignored_ascii != s_kept_ascii);
|
||||
|
||||
// unless a byte needs escaping, keep copies the input unchanged
|
||||
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
|
||||
{
|
||||
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
|
||||
});
|
||||
if (!needs_escaping)
|
||||
{
|
||||
CHECK(s_kept == "\"" + json_string + "\"");
|
||||
}
|
||||
}
|
||||
|
||||
// check that prefix and suffix are preserved
|
||||
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
|
||||
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
|
||||
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
|
||||
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
|
||||
CHECK(s_kept2.substr(1, 3) == "abc");
|
||||
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
|
||||
}
|
||||
|
||||
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
|
||||
|
||||
Reference in New Issue
Block a user