nlohmann::basic_json::error_handler_t¶
enum class error_handler_t {
strict,
replace,
ignore,
keep
};
This enumeration is used to choose how to treat ill-formed UTF-8 in a string value or object key:
dumpuses it while serializing abasic_jsonvalue to text.to_cbor,to_msgpack,to_ubjson,to_bjdata, andto_bsonuse it while serializing abasic_jsonvalue to that binary format. Their default iskeep, as no binary writer checked before this parameter was added. CBOR, UBJSON, BJData, and BSON require valid UTF-8, so for these four the default isstrictifJSON_STRICT_BINARY_UTF8is enabled; MessagePack's specification explicitly allows a string to contain ill-formed UTF-8, soto_msgpackstays atkeep.to_bon8does not take this parameter: BON8 always validates, since UTF-8 lead bytes are structural to that format.from_cbor,from_msgpack,from_ubjson,from_bjdata, andfrom_bsonuse it while parsing that binary format, to decide whether to check a string value or object key for well-formed UTF-8 at all; by default (keep) they do not, as no binary reader did before this parameter was added.from_bon8does not take this parameter, for the same reasonto_bon8does not.
Four values are differentiated:
- strict
- throw a
type_error/parse_errorexception in case of invalid UTF-8 - replace
- replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER)
- ignore
- ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
- keep
- keep invalid UTF-8 sequences unchanged; only meaningful for the binary formats mentioned above, since [
dump] (dump.md) itself must produce text, andkeepthere writes the ill-formed bytes to the output as is, so the result is then not valid UTF-8 (but still equals the input bytes exactly, including around any well-formed characters, which are still escaped as usual)
Examples¶
Example
The example below shows how the different values of the error_handler_t influence the behavior of dump when reading serializing an invalid UTF-8 sequence.
#include <iostream>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
// create JSON value with invalid UTF-8 byte sequence
json j_invalid = "ä\xA9ü";
try
{
std::cout << j_invalid.dump() << std::endl;
}
catch (const json::type_error& e)
{
std::cout << e.what() << std::endl;
}
std::cout << "string with replaced invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::replace)
<< "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< "\nstring with the invalid byte kept as is (" << j_invalid.dump(-1, ' ', false, json::error_handler_t::keep).size()
<< " bytes, not valid UTF-8 itself)\n";
}
Output:
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü"
string with the invalid byte kept as is (7 bytes, not valid UTF-8 itself)
See also¶
- dump serializes a JSON value, with an
error_handler_tparameter to configure invalid UTF-8 handling - Handling invalid UTF-8 - the article on handling invalid UTF-8
Version history¶
- Added in version 3.4.0.
- Added
keep, and made this enumeration apply to the binary readers and writers in addition todump, in version 3.13.0.