Compare commits

..
Author SHA1 Message Date
Niels Lohmann b2666c334e Merge branch 'develop' into issue4552-ignore
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 20:58:34 +02:00
Niels Lohmann ad39cda092 Mention error_handler_t::keep in the README
The README's two notes on dump() throwing for non-UTF-8 strings listed
only the replace and ignore handlers. The FAQ already mentions keep.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 18:40:41 +02:00
Niels Lohmann 8aabb9981f Add error_handler_t::keep to copy invalid UTF-8 bytes unchanged
error_handler_t::ignore drops invalid bytes although its docs promised to copy them (#4552). Add keep, which copies each ill-formed subsequence byte-for-byte while still escaping valid characters.

Fixes #4552

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-26 07:37:05 +02:00
21 changed files with 295 additions and 159 deletions
+2 -2
View File
@@ -363,7 +363,7 @@ std::cout << j_string << " == " << serialized_string << std::endl;
[`.dump()`](https://json.nlohmann.me/api/basic_json/dump/) returns the originally stored string value.
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
#### To/from streams (e.g., files, string streams)
@@ -1915,7 +1915,7 @@ The library supports **Unicode input** as follows:
- [Unicode noncharacters](https://www.unicode.org/faq/private_use.html#nonchar1) will not be replaced by the library.
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
- To store wide strings (e.g., `std::wstring`), you need to convert them to a UTF-8 encoded `std::string` before, see [an example](https://json.nlohmann.me/home/faq/#wide-string-handling).
### Comments in JSON
+1 -11
View File
@@ -272,15 +272,6 @@ basic_json(basic_json&& other) noexcept;
When used without parentheses around an empty initializer list, `basic_json()` is called instead of this
function, yielding the JSON `#!json null` value.
- Overload 4:
!!! info "Implicit conversion"
The conversion is implicit unless [`JSON_USE_IMPLICIT_CONVERSIONS`](../macros/json_use_implicit_conversions.md)
is defined to `0` and `BasicJsonType::string_t` differs from `string_t`. In that case, the constructor is
`explicit`, so a JSON value with a different string type is no longer silently converted, for example when it is
passed to a function taking `#!cpp const json&`. Write `#!cpp json(other)` or `#!cpp other.get<json>()` instead.
- Overload 7:
!!! info "Preconditions"
@@ -429,8 +420,7 @@ basic_json(basic_json&& other) noexcept;
1. Since version 1.0.0.
2. Since version 1.0.0.
3. Since version 2.1.0.
4. Since version 3.2.0. Explicit for different string types if `JSON_USE_IMPLICIT_CONVERSIONS` is `0` since
version 3.13.0.
4. Since version 3.2.0.
5. Since version 1.0.0.
6. Since version 1.0.0.
7. Since version 1.0.0.
+10 -4
View File
@@ -25,10 +25,15 @@ and `ensure_ascii` parameters.
result consists of ASCII characters only.
`error_handler` (in)
: how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped)).
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md)):
- `strict`: throw a [`type_error`](../../home/exceptions.md#type-errors) exception in case a decoding error occurs
(default),
- `replace`: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER),
- `ignore`: ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the output unchanged,
and invalid bytes are dropped, and
- `keep`: keep invalid UTF-8 sequences during serialization; all bytes are copied to the output unchanged, so the
result is not valid UTF-8.
## Return value
@@ -94,3 +99,4 @@ Binary values are serialized as an object containing two keys:
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
- Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0.
- Error handler value `keep` added in version 3.13.0.
@@ -4,12 +4,13 @@
enum class error_handler_t {
strict,
replace,
ignore
ignore,
keep
};
```
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Three values are differentiated:
`basic_json` value. Four values are differentiated:
strict
: throw a `type_error` exception in case of invalid UTF-8
@@ -20,6 +21,12 @@ replace
ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences; all bytes are copied to the output unchanged. Valid characters are still escaped as
usual (e.g., `"`, `\\`, and control characters), so the result has valid JSON syntax, but it is not valid UTF-8.
In particular, [`parse`](parse.md) rejects it, and with `ensure_ascii` set to `true`, the invalid bytes are the
only non-ASCII bytes of the output.
## Examples
??? example
@@ -40,3 +47,4 @@ ignore
## Version history
- Added in version 3.4.0.
- Added value `keep` in version 3.13.0.
@@ -5,9 +5,7 @@
```
When defined to `0`, implicit conversions are switched off. By default, implicit conversions are switched on. The
value directly affects [`operator ValueType`](../basic_json/operator_ValueType.md) and the
[converting constructor](../basic_json/basic_json.md) from a `basic_json` specialization with a different string
type (overload 4).
value directly affects [`operator ValueType`](../basic_json/operator_ValueType.md).
## Default definition
@@ -59,25 +57,6 @@ By default, implicit conversions are enabled.
auto s = j.get<std::string>();
```
??? example "Conversion between `basic_json` specializations"
A `basic_json` specialization with a different string type is also no longer converted implicitly when
`JSON_USE_IMPLICIT_CONVERSIONS` is defined to `0`:
```cpp
using wjson = nlohmann::basic_json<std::map, std::vector, std::wstring>;
void load(const nlohmann::json& j);
wjson wj = /* ... */;
load(wj); // error: no implicit conversion
load(nlohmann::json(wj)); // OK: explicit conversion
load(wj.get<nlohmann::json>()); // OK: explicit conversion
```
Specializations that share the same string type, such as `json` and `ordered_json`, remain implicitly
convertible.
## See also
- [**operator ValueType**](../basic_json/operator_ValueType.md) - get a value (implicit)
@@ -87,4 +66,3 @@ By default, implicit conversions are enabled.
## Version history
- Added in version 3.9.0.
- Also affects the conversion between `basic_json` specializations with different string types since version 3.13.0.
@@ -1,3 +1,4 @@
#include <iomanip>
#include <iostream>
#include <nlohmann/json.hpp>
@@ -21,4 +22,12 @@ int main()
<< "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< '\n';
// the invalid byte is kept; print the result byte-wise to make it visible
std::cout << "string with kept invalid characters:";
for (const unsigned char c : j_invalid.dump(-1, ' ', false, json::error_handler_t::keep))
{
std::cout << ' ' << std::hex << std::setw(2) << std::setfill('0') << static_cast<int>(c);
}
std::cout << '\n';
}
@@ -1,3 +1,4 @@
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü"
string with kept invalid characters: 22 c3 a4 a9 c3 bc 22
@@ -64,6 +64,7 @@ serialization fails by default. The fourth argument of `dump` selects an
- `strict` (default) — throw a [`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) exception.
- `replace` — replace invalid bytes with the Unicode replacement character U+FFFD (`�`).
- `ignore` — silently drop invalid bytes.
- `keep` — copy invalid bytes to the output unchanged; the result is not valid UTF-8.
??? example
+1
View File
@@ -755,6 +755,7 @@ The `dump()` function only works with UTF-8 encoded strings; that is, if you ass
- Pass an error handler as last parameter to the `dump()` function to avoid this exception:
- `json::error_handler_t::replace` will replace invalid bytes sequences with `U+FFFD`
- `json::error_handler_t::ignore` will silently ignore invalid byte sequences
- `json::error_handler_t::keep` will copy invalid byte sequences to the output unchanged
### json.exception.type_error.317
+1 -1
View File
@@ -85,7 +85,7 @@ The library supports **Unicode input** as follows:
- The library will not replace [Unicode noncharacters](http://www.unicode.org/faq/private_use.html#nonchar1).
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding.
@@ -291,9 +291,7 @@ void to_json(BasicJsonType& j, const std::optional<T>& opt) noexcept
{
if (opt.has_value())
{
// explicit construction, as the conversion from a basic_json with a different
// string type is explicit if JSON_USE_IMPLICIT_CONVERSIONS is 0 (#2649)
j = BasicJsonType(*opt);
j = *opt;
}
else
{
+52 -1
View File
@@ -48,7 +48,8 @@ enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore ///< ignore invalid UTF-8 sequences
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
};
template<typename BasicJsonType>
@@ -1019,6 +1020,47 @@ class serializer
break;
}
case error_handler_t::keep:
{
// drop whatever the incomplete sequence left in
// the buffer (only copied if !EnsureAscii) and copy
// the ill-formed bytes from the input instead
bytes = bytes_after_last_accept;
if (undumped_chars > 0)
{
// the pending bytes of the incomplete sequence
// are ill-formed; the current byte may be OK for
// itself, so we would like to read it again
for (std::size_t j = i - undumped_chars; j < i; ++j)
{
string_buffer[bytes++] = s[j];
}
--i;
}
else
{
// the current byte cannot start any sequence
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -1064,6 +1106,15 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write all accepted bytes
put_buffer(string_buffer, bytes_after_last_accept);
// copy the bytes of the incomplete sequence unchanged
put_string(s, s.size() - undumped_chars, s.size());
break;
}
case error_handler_t::replace:
{
// write all accepted bytes
+3 -34
View File
@@ -1605,42 +1605,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
assert_invariant();
}
private:
/// whether a basic_json specialization can be converted implicitly into this one;
/// with JSON_USE_IMPLICIT_CONVERSIONS set to 0, this is only the case if both share
/// the same string type (see https://github.com/nlohmann/json/issues/2649)
template<typename BasicJsonType>
using is_implicitly_convertible_basic_json = std::integral_constant < bool,
(JSON_USE_IMPLICIT_CONVERSIONS != 0)
|| std::is_same<typename BasicJsonType::string_t, string_t>::value >;
/// tag to select the constructor that performs the conversion from another basic_json specialization
struct convert_basic_json_tag {};
public:
/// @brief create a JSON value from an existing one
/// @sa https://json.nlohmann.me/api/basic_json/basic_json/
template < typename BasicJsonType,
detail::enable_if_t <
detail::is_basic_json<BasicJsonType>::value&& !std::is_same<basic_json, BasicJsonType>::value
&& is_implicitly_convertible_basic_json<BasicJsonType>::value, int > = 0 >
detail::is_basic_json<BasicJsonType>::value&& !std::is_same<basic_json, BasicJsonType>::value, int > = 0 >
basic_json(const BasicJsonType& val)
: basic_json(val, convert_basic_json_tag{})
{}
/// @brief create a JSON value from an existing one
/// @sa https://json.nlohmann.me/api/basic_json/basic_json/
template < typename BasicJsonType,
detail::enable_if_t <
detail::is_basic_json<BasicJsonType>::value&& !std::is_same<basic_json, BasicJsonType>::value
&& !is_implicitly_convertible_basic_json<BasicJsonType>::value, int > = 0 >
explicit basic_json(const BasicJsonType& val)
: basic_json(val, convert_basic_json_tag{})
{}
private:
template<typename BasicJsonType>
basic_json(const BasicJsonType& val, convert_basic_json_tag /*unused*/)
#if JSON_DIAGNOSTIC_POSITIONS
: start_position(val.start_pos()),
end_position(val.end_pos())
@@ -1696,7 +1666,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
assert_invariant();
}
public:
/// @brief create a container (array or object) from an initializer list
/// @sa https://json.nlohmann.me/api/basic_json/basic_json/
basic_json(initializer_list_t init,
@@ -2463,7 +2432,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
int > = 0 >
BasicJsonType get_impl(detail::priority_tag<2> /*unused*/) const
{
return BasicJsonType(*this);
return *this;
}
/*!
@@ -2602,7 +2571,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
int> = 0>
ValueType & get_to(ValueType& v) const
{
v = ValueType(*this);
v = *this;
return v;
}
+56 -38
View File
@@ -6820,9 +6820,7 @@ void to_json(BasicJsonType& j, const std::optional<T>& opt) noexcept
{
if (opt.has_value())
{
// explicit construction, as the conversion from a basic_json with a different
// string type is explicit if JSON_USE_IMPLICIT_CONVERSIONS is 0 (#2649)
j = BasicJsonType(*opt);
j = *opt;
}
else
{
@@ -23933,7 +23931,8 @@ enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore ///< ignore invalid UTF-8 sequences
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
};
template<typename BasicJsonType>
@@ -24904,6 +24903,47 @@ class serializer
break;
}
case error_handler_t::keep:
{
// drop whatever the incomplete sequence left in
// the buffer (only copied if !EnsureAscii) and copy
// the ill-formed bytes from the input instead
bytes = bytes_after_last_accept;
if (undumped_chars > 0)
{
// the pending bytes of the incomplete sequence
// are ill-formed; the current byte may be OK for
// itself, so we would like to read it again
for (std::size_t j = i - undumped_chars; j < i; ++j)
{
string_buffer[bytes++] = s[j];
}
--i;
}
else
{
// the current byte cannot start any sequence
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -24949,6 +24989,15 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write all accepted bytes
put_buffer(string_buffer, bytes_after_last_accept);
// copy the bytes of the incomplete sequence unchanged
put_string(s, s.size() - undumped_chars, s.size());
break;
}
case error_handler_t::replace:
{
// write all accepted bytes
@@ -27490,42 +27539,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
assert_invariant();
}
private:
/// whether a basic_json specialization can be converted implicitly into this one;
/// with JSON_USE_IMPLICIT_CONVERSIONS set to 0, this is only the case if both share
/// the same string type (see https://github.com/nlohmann/json/issues/2649)
template<typename BasicJsonType>
using is_implicitly_convertible_basic_json = std::integral_constant < bool,
(JSON_USE_IMPLICIT_CONVERSIONS != 0)
|| std::is_same<typename BasicJsonType::string_t, string_t>::value >;
/// tag to select the constructor that performs the conversion from another basic_json specialization
struct convert_basic_json_tag {};
public:
/// @brief create a JSON value from an existing one
/// @sa https://json.nlohmann.me/api/basic_json/basic_json/
template < typename BasicJsonType,
detail::enable_if_t <
detail::is_basic_json<BasicJsonType>::value&& !std::is_same<basic_json, BasicJsonType>::value
&& is_implicitly_convertible_basic_json<BasicJsonType>::value, int > = 0 >
detail::is_basic_json<BasicJsonType>::value&& !std::is_same<basic_json, BasicJsonType>::value, int > = 0 >
basic_json(const BasicJsonType& val)
: basic_json(val, convert_basic_json_tag{})
{}
/// @brief create a JSON value from an existing one
/// @sa https://json.nlohmann.me/api/basic_json/basic_json/
template < typename BasicJsonType,
detail::enable_if_t <
detail::is_basic_json<BasicJsonType>::value&& !std::is_same<basic_json, BasicJsonType>::value
&& !is_implicitly_convertible_basic_json<BasicJsonType>::value, int > = 0 >
explicit basic_json(const BasicJsonType& val)
: basic_json(val, convert_basic_json_tag{})
{}
private:
template<typename BasicJsonType>
basic_json(const BasicJsonType& val, convert_basic_json_tag /*unused*/)
#if JSON_DIAGNOSTIC_POSITIONS
: start_position(val.start_pos()),
end_position(val.end_pos())
@@ -27581,7 +27600,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
assert_invariant();
}
public:
/// @brief create a container (array or object) from an initializer list
/// @sa https://json.nlohmann.me/api/basic_json/basic_json/
basic_json(initializer_list_t init,
@@ -28348,7 +28366,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
int > = 0 >
BasicJsonType get_impl(detail::priority_tag<2> /*unused*/) const
{
return BasicJsonType(*this);
return *this;
}
/*!
@@ -28487,7 +28505,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
int> = 0>
ValueType & get_to(ValueType& v) const
{
v = ValueType(*this);
v = *this;
return v;
}
-36
View File
@@ -13,7 +13,6 @@
#include <cstdint>
#include <string>
#include <type_traits>
#include <utility>
#include <vector>
@@ -375,39 +374,4 @@ TEST_CASE("alternative string type")
const auto j2 = j.flatten();
CHECK(j2.dump() == R"({"/foo/0":"bar","/foo/1":"baz"})");
}
SECTION("conversion between basic_json specializations (#2649)")
{
// explicit conversions are always possible
CHECK(std::is_constructible<nlohmann::json, alt_json>::value);
CHECK(std::is_constructible<alt_json, nlohmann::json>::value);
CHECK(std::is_constructible<nlohmann::json, nlohmann::ordered_json>::value);
CHECK(std::is_constructible<nlohmann::ordered_json, nlohmann::json>::value);
// specializations with the same string type are implicitly convertible
CHECK(std::is_convertible<nlohmann::ordered_json, nlohmann::json>::value);
CHECK(std::is_convertible<nlohmann::json, nlohmann::ordered_json>::value);
// specializations with different string types are only implicitly convertible
// if implicit conversions are enabled
#if JSON_USE_IMPLICIT_CONVERSIONS
CHECK(std::is_convertible<alt_json, nlohmann::json>::value);
CHECK(std::is_convertible<nlohmann::json, alt_json>::value);
#else
CHECK_FALSE(std::is_convertible<alt_json, nlohmann::json>::value);
CHECK_FALSE(std::is_convertible<nlohmann::json, alt_json>::value);
#endif
// get<BasicJsonType>() works in either case
const nlohmann::json j = {{"foo", 1}, {"bar", true}};
CHECK(j.get<nlohmann::ordered_json>() == nlohmann::ordered_json(j));
// (only a number is converted here, as objects and strings are affected by #3425)
CHECK(nlohmann::json(42).get<alt_json>() == 42);
CHECK(alt_json(nlohmann::json(42)) == 42);
// get_to() also works in either case
alt_json a;
nlohmann::json(42).get_to(a);
CHECK(a == 42);
}
}
+9
View File
@@ -766,6 +766,15 @@ TEST_CASE("regression tests 2")
CHECK(j == k);
}
SECTION("issue #4552 - UTF-8 invalid characters are not always ignored when dumping with error_handler_t::ignore")
{
json node;
node["test"] = "test\334\005";
CHECK(node.dump(-1, ' ', false, json::error_handler_t::ignore) == "{\"test\":\"test\\u0005\"}");
CHECK(node.dump(-1, ' ', false, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
CHECK(node.dump(-1, ' ', true, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
}
}
TEST_CASE("regression test - parser callback must not lose a duplicate key's prior value")
+37
View File
@@ -92,6 +92,8 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"äü\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"ä\xEF\xBF\xBDü\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"\\u00e4\\ufffd\\u00fc\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"ä\xA9ü\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\\u00e4\xA9\\u00fc\"");
}
SECTION("invalid character (regression guard for shared UTF-8 decoder, see #5529)")
@@ -114,6 +116,8 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xC2\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xC2\"");
}
SECTION("unexpected character")
@@ -126,6 +130,39 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123456\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\x34\x35\x36\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd456\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
}
SECTION("keep: valid characters are still escaped")
{
// an invalid byte followed by characters that must be escaped
const json j = "\xC2\"\\\n\xFF\x05";
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
}
SECTION("keep: truncated multibyte sequences")
{
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
}
SECTION("keep: long string with many invalid bytes")
{
// exceeds the internal string buffer several times
std::string input;
std::string expected = "\"";
for (int i = 0; i < 2000; ++i)
{
input += "\xFF\xE2\x82\n\xC3\xA4";
expected += "\xFF\xE2\x82\\n\xC3\xA4";
}
expected += "\"";
const json j = input;
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == expected);
}
SECTION("U+FFFD Substitution of Maximal Subparts")
+25 -1
View File
@@ -14,6 +14,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -75,8 +76,11 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace must not throw in any case
// dumping with ignore/replace/keep must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -85,6 +89,9 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -94,6 +101,7 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -105,6 +113,20 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -116,6 +138,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+25 -1
View File
@@ -14,6 +14,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -75,8 +76,11 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace must not throw in any case
// dumping with ignore/replace/keep must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -85,6 +89,9 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -94,6 +101,7 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -105,6 +113,20 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -116,6 +138,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+25 -1
View File
@@ -14,6 +14,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -75,8 +76,11 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace must not throw in any case
// dumping with ignore/replace/keep must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -85,6 +89,9 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -94,6 +101,7 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -105,6 +113,20 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -116,6 +138,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+25 -1
View File
@@ -14,6 +14,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -75,8 +76,11 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace must not throw in any case
// dumping with ignore/replace/keep must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -85,6 +89,9 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -94,6 +101,7 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -105,6 +113,20 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -116,6 +138,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);