Trim the compiler-appended NUL from wide/UTF string literals too (#5702)

With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16,
UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]"))
failed with parse_error.101 at the terminating NUL of the literal,
and accept() returned false. The array overload of input_adapter()
only dropped the compiler-added trailing '\0' for arrays of char,
so for wchar_t, char16_t, char32_t, and char8_t arrays that
terminator was passed to the parser as data, which the macro then
rejected.

Broaden the trimming to every character type that a string literal
can use (char, wchar_t, char16_t, char32_t, and, since C++20,
char8_t). Arrays of any other element type (unsigned char,
std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a
trailing zero byte there is still read as genuine data.

Fixes #5658.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-30 20:07:37 +02:00
committed by GitHub
parent bfe0f32d71
commit 4bb1b14b06
4 changed files with 79 additions and 28 deletions
@@ -16,7 +16,8 @@ byte is still not rejected:
[`from_msgpack`](../basic_json/from_msgpack.md), [`from_ubjson`](../basic_json/from_ubjson.md)) are never affected: there, `0x00` is ordinary data.
- A bare `const char*` pointer has no length of its own, so its length is still determined with `strlen()`. The first
NUL byte therefore still marks the end of the input, and nothing after it is read.
- One trailing `'\0'` at the end of a `char` array (e.g., a string literal) is trimmed; see the warning below.
- One trailing `'\0'` at the end of a `char`, `wchar_t`, `char16_t`, `char32_t`, or (C++20) `char8_t` array (e.g., a
string literal) is trimmed; see the warning below.
## Default definition
@@ -57,13 +58,15 @@ The default value is `0` (disabled — existing behavior is preserved).
This macro must be defined **before** including `<nlohmann/json.hpp>`. Defining it after the include has no
effect.
Enabling it also changes how a `char` array (including a string literal, e.g. `json::parse("123")`) is read: such
an array normally carries a trailing `'\0'` contributed by the compiler, not by the source text. With this macro
enabled, that one trailing byte is trimmed if present so that parsing a string literal keeps working; every other
byte in the array - including any `'\0'` that is not the very last element - is read as real data and rejected
like any other unexpected byte. Arrays of any other element type (`unsigned char`, `std::uint8_t`, ...), as used
for CBOR or MessagePack, are never affected by this trimming; their full extent - including a genuine trailing
`0x00` - is always preserved, in both states of this macro.
Enabling it also changes how an array of a text-literal element type (`char`, `wchar_t`, `char16_t`, `char32_t`,
or, since C++20, `char8_t` - including a string literal, e.g. `json::parse("123")` or `json::parse(L"123")`) is
read: such an array normally carries a trailing `'\0'` contributed by the compiler, not by the source text. With
this macro enabled, that one trailing element is trimmed if present so that parsing a string literal keeps
working, for any of these character types; every other element in the array - including any `'\0'` that is not
the very last element - is read as real data and rejected like any other unexpected byte. Arrays of any other
element type (`unsigned char`, `std::uint8_t`, ...), as used for CBOR or MessagePack, are never affected by this
trimming; their full extent - including a genuine trailing `0x00` - is always preserved, in both states of this
macro.
!!! note "ABI compatibility"