Files
json/tests/src/unit-binary_utf8_strict.cpp
T
Niels Lohmann d8d47be4a5 Fix CI on develop after merging the ready-to-merge PRs (#5754)
* Keep the serializer conversion for objects whose keys cannot be converted

#5591 added a test converting nlohmann::json into a basic_json whose
string type cannot be constructed from std::string. That instantiates
convert_iteratively(), whose members.emplace_back(next.key(), ...) needs
exactly that key conversion, and broke the build of unit-alt-string.
Dispatch on the key's constructibility and leave such conversions to the
serializers, as the levels above the nesting bound already do (#3425).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the remaining CI failures on develop

- unit-wstring: with a 16-bit wchar_t (Windows), a lone surrogate is
  reported as the ill-formed byte 0xFF since #5704; the std::wstring
  expectations still had the previous <U+0000>.
- ci_single_binaries: json_literals.hpp (#5610) and json.hpp include each
  other on purpose, and IWYU, not following the cycle, asks to replace
  json.hpp with json_fwd.hpp. Report its findings without failing the
  build, as already done for json.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the library warnings and noexcept specifications from the merged PRs

- binary_reader: rename the error_handler constructor parameter, which
  shadowed the member (-Wshadow, -Wshadow-field-in-constructor; #5746)
- basic_json(copy_construct_tag, ...): declare it noexcept when copying
  the base class is (GCC 16 -Wnoexcept; #5690)
- the scalar-on-left legacy comparison operators: noexcept only when
  converting the scalar is, like their member counterparts (#5682, #5751)
- compare_leaves: use std::is_eq/is_lt/is_gt instead of comparing a
  std::partial_ordering with 0 (-Wzero-as-null-pointer-constant; #5686)
- serializer: silence MSVC C4127 for the EnsureAscii template parameter
  (#5741, #5746)
- clang-tidy: return the sanitized reference in binary_writer, take the
  key of ordered_map::find_impl by const reference (#5727), and mark the
  switches over parse_array_index (#5728)
- ordered_map: keep <memory> for std::allocator (IWYU)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Split unit-conversions.cpp so MinGW can link it

clang 18 with the MinGW linker failed to link test-conversions_cpp17
("relocation truncated to fit: IMAGE_REL_AMD64_REL32"). As windows.yml
recommends, keep the objects small by splitting the test file.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the tests added by the merged PRs for all CI configurations

- discard the results of dump() and from_*() in CHECK_THROWS with
  utils::ignore_return_value (GCC -Werror=unused-result)
- give unit-bson's huge_string_t a default constructor (MSVC C2512,
  GCC 5, clang 3.5)
- unit-disabled_exceptions: use the literals namespace when the global
  UDLs are off (ci_test_noglobaludls; #5700)
- unit-binary_utf8_strict: expect the JSON pointer prefix with
  JSON_DIAGNOSTICS (#5741)
- skip the tests that rely on exceptions under JSON_NOEXCEPTION
  (#5678, #5732)
- clang-tidy and clang -Werror: static test data, CAPTURE(...);,
  const-correctness, use-after-move alias, unused conversion operator,
  a missing <iterator> include

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Title the macro examples and add JSON_STRICT_BINARY_UTF8 to the docset

The documentation style check requires "Example: ..." titles on pages with several examples (#5741, #5591) and a docset entry for every macro page.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Regenerate BUILD.bazel and nlohmann_json.natvis

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5746 added detail/output/error_handler.hpp and #5741 the json_abi_sbu8 ABI tag.

* Install libidn11 for the CMake 3.5.0 binary in ci_cmake_flags

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5733 moved ci_cmake_options from ubuntu:focal to ubuntu:24.04, which no longer ships libidn.so.11; the CMake 3.5.0 release binary links against it, so every ci_cmake_flags run has failed since. Install focal's libidn11 package for that matrix entry only.

* Suppress Infer's false STACK_VARIABLE_ADDRESS_ESCAPE in get_impl

get_impl() returns its local by value. A test added by the merged PRs instantiates it with a type Infer misreads, so ci_infer reported the 2021 code for the first time.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:18:19 +02:00

147 lines
8.6 KiB
C++

// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// The binary writers check strings and object keys for valid UTF-8 only if
// JSON_STRICT_BINARY_UTF8 is enabled (planned to be the default in 4.0.0).
// Without it, they write the bytes unchanged, as before version 3.13.0; the
// tests for that are next to the other tests of each format.
#ifdef JSON_STRICT_BINARY_UTF8
#undef JSON_STRICT_BINARY_UTF8
#endif
#define JSON_STRICT_BINARY_UTF8 1
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <cstdint>
#include <vector>
TEST_CASE("JSON_STRICT_BINARY_UTF8 (see #5529, #5651)")
{
SECTION("CBOR")
{
// a string value with ill-formed UTF-8 is rejected
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_cbor(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
// a value read back from CBOR with ill-formed bytes cannot be written
// back either (the reader is lenient regardless of the macro)
const json j = json::from_cbor(std::vector<std::uint8_t>({0x62, 0xc0, 0xae}));
CHECK_THROWS_WITH_AS(json::to_cbor(j), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
}
SECTION("UBJSON")
{
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_ubjson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("BJData")
{
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_bjdata(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("BSON")
{
// to_bson() rejects the same kind of ill-formed string value, before
// any bytes reach the output adapter (the BSON document length
// prefix must be known up front, so nothing is written incrementally)
std::vector<std::uint8_t> out{0x42}; // a sentinel byte the writer must not touch
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] (/s) invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
#endif
CHECK(out == std::vector<std::uint8_t> {0x42});
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xFF"}}), "[json.exception.type_error.316] (/s) invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xFF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
#endif
// a truncated multi-byte sequence
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xC3"}}), "[json.exception.type_error.316] (/s) invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xC3"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
#endif
// an encoded surrogate half (U+D800)
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] (/s) invalid UTF-8 byte at index 0: 0xED", json::type_error&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
#endif
// an overlong encoding of '.'
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] (/s) invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
#endif
// an object key with ill-formed UTF-8 is rejected as well; unlike
// the reader (which never validates element names), the writer
// checks both string values and object keys
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"\xFF", 1}}), "[json.exception.type_error.316] (/\xFF) invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
#endif
}
SECTION("an explicit error_handler overrides the default")
{
// the macro only changes the default of the error_handler parameter
CHECK(json::to_cbor(json("\xFF"), json::error_handler_t::keep) == std::vector<std::uint8_t>({0x61, 0xff}));
CHECK(json::to_ubjson(json("\xFF"), false, false, json::error_handler_t::keep) == std::vector<std::uint8_t>({'S', 'i', 1, 0xff}));
CHECK(json::to_bjdata(json("\xFF"), false, false, json::bjdata_version_t::draft2, json::error_handler_t::keep) == std::vector<std::uint8_t>({'S', 'i', 1, 0xff}));
CHECK(json::from_bson(json::to_bson(json{{"s", "\xFF"}}, json::error_handler_t::keep)) == json{{"s", "\xFF"}});
CHECK(json::to_cbor(json("\xFF"), json::error_handler_t::replace) == std::vector<std::uint8_t>({0x63, 0xef, 0xbf, 0xbd}));
}
SECTION("MessagePack and BON8 are unaffected")
{
// MessagePack allows any bytes in a str, so to_msgpack() still
// defaults to keep (strict only if passed explicitly); BON8 always
// checks, because the lead bytes mark where strings end
CHECK(json::to_msgpack(json("\xFF")) == std::vector<std::uint8_t>({0xa1, 0xff}));
CHECK_THROWS_AS(json::to_msgpack(json("\xFF"), json::error_handler_t::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bon8(json("\xFF")), json::type_error&);
}
}