Files
json/include/nlohmann/detail/output/error_handler.hpp
T
Niels Lohmann fff221c58b Add an error_handler parameter for UTF-8 to the binary readers and writers
Adds error_handler_t::keep (invalid UTF-8 sequences are left unchanged) as
a fourth error_handler_t value, and threads an error_handler parameter
through the binary writers and readers:

- to_cbor/to_ubjson/to_bjdata/to_bson gain a trailing error_handler
  parameter (default strict, matching their existing type_error.316
  behavior); keep writes a string value or object key's bytes as is
  instead of throwing, and replace/ignore sanitize it exactly like
  dump() would, including for the BSON length prefix. to_msgpack and
  to_bon8 are unchanged.
- from_cbor/from_msgpack/from_ubjson/from_bjdata/from_bson gain a
  trailing error_handler parameter (default keep, i.e. the lenient
  behavior every binary reader had in 3.12.0 and still has after
  #5741); strict checks every string value and object key and raises
  parse_error.113 for ill-formed UTF-8, honoring allow_exceptions;
  replace/ignore sanitize it like dump() would. from_bon8 is
  unchanged, since UTF-8 lead bytes are structural there.

dump()'s own keep support writes ill-formed bytes as is, even with
ensure_ascii, while still \u-escaping well-formed characters around
them as usual.

The UTF-8 validity check (is_valid_utf8) and the replace/ignore
sanitizing logic (sanitize_utf8) now live in string_utils.hpp, shared
by the serializer and the binary reader/writer; error_handler_t itself
moved to its own header (detail/output/error_handler.hpp) so that
string_utils.hpp does not need to depend on serializer.hpp.

See #5529 and #5741.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:43:15 +02:00

39 lines
1.4 KiB
C++

// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/// how to treat decoding errors
///
/// @ref basic_json::dump uses this to decide what to do with ill-formed
/// UTF-8 while escaping a string, and the binary writers (@ref
/// basic_json::to_cbor, @ref basic_json::to_ubjson, @ref
/// basic_json::to_bjdata, @ref basic_json::to_bson) use it the same way for
/// string values and object keys. The binary readers (@ref
/// basic_json::from_cbor, @ref basic_json::from_msgpack, @ref
/// basic_json::from_ubjson, @ref basic_json::from_bjdata, @ref
/// basic_json::from_bson) use it to decide whether to check text strings
/// and object keys for well-formed UTF-8 at all, since none of those
/// formats requires a decoder to do so.
enum class error_handler_t
{
strict, ///< throw a type_error/parse_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences unchanged
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END