mirror of
https://github.com/nlohmann/json.git
synced 2026-09-29 19:20:30 +00:00
fix: treat a NUL byte in the input as an ordinary byte, not EOF
The lexer's token dispatch had `case '\0':` fall through to the same
`end_of_input` handling as the real end-of-file sentinel, with a comment
claiming the NUL case was "needed when parsing from string literals".
That rationale no longer holds: input_adapter(const char*) already uses
strlen() to compute its range, so it never hands the lexer a trailing
NUL, and the const-char* overload is the only "string literal" path the
comment could be referring to. In practice, `case '\0':` only ever fired
on a genuine embedded or trailing NUL byte in real input data (e.g. a
std::string with '\0' appended), which was then silently swallowed as if
it were EOF instead of producing the parse_error.101 any other
unexpected byte gets. Two more spots in the comment-skipping logic had
the same NUL-as-EOF idiom, stopping a `//` or `/* */` comment scan early
at an embedded NUL instead of continuing to the real terminator.
Removing all three still left one real regression: input_adapter's
T(&array)[N] overload (used for a string literal like
json::parse("123"), as opposed to a decayed const char* pointer) passes
the array's full extent through unchanged, trailing '\0' included. That
path was relying on the lexer's old NUL-as-EOF behavior to make ordinary
literal parsing work at all. It now gets its own strlen()-like handling
for char arrays specifically: a single trailing NUL terminator is
excluded, mirroring the pointer overload, while non-char arrays (e.g.
uint8_t buffers for binary formats) are left untouched since a trailing
zero byte there may be data.
Also updates a few existing tests that (mostly incidentally) depended on
a trailing NUL being swallowed - std::array<uint8_t, 5>{"true"} left the
5th element zero-initialized - and adds an FAQ entry.
Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
e5f84e1ebf
commit
633eef8494
@@ -90,6 +90,21 @@ The library supports **Unicode input** as follows:
|
||||
In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding.
|
||||
|
||||
|
||||
### NUL bytes in the input
|
||||
|
||||
!!! question
|
||||
|
||||
Why does parsing fail with "invalid literal" or "unexpected additional data" when my input contains a `'\0'` (NUL) byte?
|
||||
|
||||
A `'\0'` byte that occurs inside or at the end of the input is **not** treated as end-of-input; it is treated as an ordinary, invalid byte, exactly like any other unexpected byte in that position. [RFC 8259](https://tools.ietf.org/html/rfc8259.html) does not give the NUL byte any special end-of-text meaning, so a JSON text that is embedded in a larger byte sequence (for instance, a `std::string` with a trailing `'\0'` appended, or a buffer that happens to be zero-padded) will yield a `parse_error.101`, the same error you would get for any other unexpected trailing or misplaced byte:
|
||||
|
||||
- If the NUL byte follows a complete value, parsing fails with the usual "expected end of input" message the library also gives for any other unexpected trailing byte (e.g., `json::parse(std::string("123") + '\0')` fails the same way `json::parse("123x")` does).
|
||||
- If the NUL byte occurs where a value is expected (for instance, at the very start of the input, or right after a `:` or `,`), the library reports "invalid literal".
|
||||
- A NUL byte inside a quoted string still needs to be escaped as `\u0000`, as required by [RFC 8259](https://tools.ietf.org/html/rfc8259.html#section-7); an unescaped NUL there is reported separately as a control character that must be escaped.
|
||||
|
||||
Only the length actually passed to the parser matters here: parsing a `const char*` (for example, a string literal) uses `strlen()`-like semantics and therefore never sees the terminating NUL, so `json::parse("[1,2,3]")` is unaffected. What is affected is input that explicitly includes a NUL byte as data, such as a `std::string` with `'\0'` appended, or an iterator range/container whose end includes it.
|
||||
|
||||
|
||||
### Wide string handling
|
||||
|
||||
!!! question
|
||||
|
||||
Reference in New Issue
Block a user