Compare commits

..
Author SHA1 Message Date
Niels Lohmann af25bf67f3 Report BON8 input that ends after a UTF-8 lead byte as truncated
A lead byte (0xC2..0xF7) inside a string begins either another character
(if a continuation byte follows) or an integer (otherwise). When the input
ended right after the lead byte, the reader took the missing byte as "not
a continuation byte", ended the string before the lead byte, and treated
the lead byte as the start of the next value. With strict=false, a message
cut off there was therefore read as a shorter value: the 11 bytes of
"😀😀é" cut after 9 bytes gave "😀😀", and ["aé"] cut after 3 of its 5
bytes gave ["a"]. With strict=true, the input was rejected with a
misleading message ("expected end of input"), or, for a key, with
parse_error.112 instead of 110.

Either reading of the lead byte leaves the message incomplete: a string at
the end of a message must be terminated by 0xFF, so the lead byte cannot
belong to a following message. Report parse_error.110 (unexpected end of
input) for strings and keys, as the comment on get_bon8_string() already
requires and as the reference decoder (HikoGUI) does.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 20:25:03 +02:00
19 changed files with 92 additions and 53 deletions
+1 -11
View File
@@ -7,16 +7,6 @@ on:
- develop
paths:
- docs/mkdocs/**
# the site also embeds these files via pymdownx.snippets
# (mkdocs.yml sets restrict_base_path: false for this)
- .clang-tidy
- .github/CODE_OF_CONDUCT.md
- .github/CONTRIBUTING.md
- .github/SECURITY.md
- cmake/clang_flags.cmake
- cmake/gcc_flags.cmake
- tests/fmt_formatter/project/main.cpp
- tools/astyle/.astylerc
workflow_dispatch:
# we don't want to have concurrent jobs, and we don't want to cancel running jobs to avoid broken publications
@@ -33,7 +23,7 @@ jobs:
contents: write
if: github.repository == 'nlohmann/json'
runs-on: ubuntu-latest
runs-on: ubuntu-22.04
steps:
- name: Harden Runner
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
-4
View File
@@ -346,8 +346,6 @@ jobs:
container: intel/oneapi-hpckit:latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Get latest CMake and ninja
uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2
- name: Run CMake
@@ -360,8 +358,6 @@ jobs:
container: nvcr.io/nvidia/nvhpc:25.5-devel-cuda12.9-ubuntu22.04
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Get latest CMake and ninja
uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2
- name: Run CMake
-4
View File
@@ -87,8 +87,6 @@ jobs:
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Get latest CMake and ninja
uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2
- name: Set extra CXX_FLAGS for latest std_version
@@ -125,8 +123,6 @@ jobs:
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Run CMake (Release)
run: cmake -S . -B build -G "Visual Studio 18 2026" -A ARM64 -DJSON_BuildTests=On -DCMAKE_CXX_FLAGS="/W4 /WX"
if: matrix.build_type == 'Release'
+5 -16
View File
@@ -11,31 +11,20 @@ EXAMPLES = $(wildcard mkdocs/docs/examples/*.cpp)
cxx_standard = $(lastword c++11 $(filter c++%, $(subst ., ,$1)))
# common compile flags for the stand-alone example files
EXAMPLE_CPPFLAGS = -I $(SRCDIR) -DJSON_USE_GLOBAL_UDLS=0
EXAMPLE_WARNFLAGS = -Werror=deprecated-declarations
# examples that document deprecated API and are allowed to use it
DEPRECATED_EXAMPLES = $(addprefix mkdocs/docs/examples/, \
json_pointer__operator__equal_stringtype \
json_pointer__operator__notequal_stringtype \
json_pointer__operator_string_t)
$(DEPRECATED_EXAMPLES:=.output) $(DEPRECATED_EXAMPLES:=.test): EXAMPLE_WARNFLAGS = -Wno-deprecated-declarations
# create output from a stand-alone example file
%.output: %.cpp
@echo "standard $(call cxx_standard,$(<:.cpp=))"
@echo "standard $(call cxx_standard $(<:.cpp=))"
$(MAKE) $(<:.cpp=) \
CPPFLAGS="$(EXAMPLE_CPPFLAGS)" \
CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) $(EXAMPLE_WARNFLAGS)"
CPPFLAGS="-I $(SRCDIR) -DJSON_USE_GLOBAL_UDLS=0" \
CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) -Wno-deprecated-declarations"
./$(<:.cpp=) > $@
rm $(<:.cpp=)
# compare created output with current output of the example files
%.test: %.cpp
$(MAKE) $(<:.cpp=) \
CPPFLAGS="$(EXAMPLE_CPPFLAGS)" \
CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) $(EXAMPLE_WARNFLAGS)"
CPPFLAGS="-I $(SRCDIR) -DJSON_USE_GLOBAL_UDLS=0" \
CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) -Wno-deprecated-declarations"
./$(<:.cpp=) > $@
diff $@ $(<:.cpp=.output)
rm $(<:.cpp=) $@
+1 -1
View File
@@ -13,7 +13,7 @@
<key>dashIndexFilePath</key>
<string>index.html</string>
<key>DashDocSetFallbackURL</key>
<string>https://json.nlohmann.me/</string>
<string>https://nlohmann.github.io/json/</string>
<key>isJavaScriptEnabled</key>
<true/>
</dict>
+2 -3
View File
@@ -7,11 +7,10 @@ documentation browsers like [Dash](https://kapeli.com/dash), [Velocity](https://
The docset can be created with
```sh
make JSON_for_Modern_C++.docset
make nlohmann_json.docset
```
The generated folder `JSON_for_Modern_C++.docset` can then be opened in the documentation browser. `make all` builds a
`JSON_for_Modern_C++.tgz` archive instead, and `make install_docset_zeal` installs the docset for Zeal directly.
The generated folder `nlohmann_json.docset` can then be opened in the documentation browser.
A recent version is also part of the [Dash user contributions](https://github.com/Kapeli/Dash-User-Contributions/tree/master/docsets/JSON_for_Modern_C%2B%2B).
@@ -51,7 +51,7 @@ range will yield over/underflow when used in a constructor. During deserializati
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
> interoperable in the sense that implementations will agree exactly on their numeric values.
As this range is a subrange of the exactly supported range [INT64_MIN, INT64_MAX], this class's integer type is
@@ -52,7 +52,7 @@ when used in a constructor. During deserialization, too large or small integer n
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
> interoperable in the sense that implementations will agree exactly on their numeric values.
As this range is a subrange (when considered in conjunction with the `number_integer_t` type) of the exactly supported
@@ -54,7 +54,7 @@ classDiagram
## Notes
For an input with <i>n</i> bytes, 1 is the index of the first character and <i>n</i>+1 is the index of the terminating null byte
For an input with $n$ bytes, 1 is the index of the first character and $n+1$ is the index of the terminating null byte
or the end of file. This also holds true when reading a byte vector for binary formats.
## Examples
@@ -0,0 +1 @@
<a target="_blank" href="https://wandbox.org/permlink/hUJYo1HWmfTBLMGn"><b>online</b></a>
@@ -0,0 +1 @@
<a target="_blank" href="https://wandbox.org/permlink/AWbpa8e1xRV3y4MM"><b>online</b></a>
@@ -61,7 +61,7 @@ The library uses the following mapping from JSON values types to BJData types ac
The following values can **not** be converted to a BJData value:
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
- strings with more than 18446744073709551615 bytes, i.e., $2^{64}-1$ bytes (theoretical)
!!! info "Unused BJData markers"
+1 -1
View File
@@ -273,7 +273,7 @@ When the default type is used, the maximal unsigned integer number that can be s
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are interoperable in the sense that implementations will agree exactly on their numeric values.
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are interoperable in the sense that implementations will agree exactly on their numeric values.
As this range is a subrange of the exactly supported range [`INT64_MIN`, `INT64_MAX`], this class's integer type is interoperable.
@@ -48,7 +48,7 @@ On number interoperability, the following remarks are made:
for numeric magnitude and precision than is widely available.
Note that when such software is used, numbers that are integers and
are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are interoperable in the
are in the range $[-2^{53}+1, 2^{53}-1]$ are interoperable in the
sense that implementations will agree exactly on their numeric
values.
@@ -95,9 +95,9 @@ This is the same behavior as the code `#!c double x = 3.141592653589793238462643
!!! success "Interoperability"
- The library is interoperable with respect to the specification, because its supported range [-2<sup>63</sup>, 2<sup>64</sup>-1] is
larger than the described range [-2<sup>53</sup>+1, 2<sup>53</sup>-1].
- All integers outside the range [-2<sup>63</sup>, 2<sup>64</sup>-1], as well as floating-point numbers are stored as `double`.
- The library is interoperable with respect to the specification, because its supported range $[-2^{63}, 2^{64}-1]$ is
larger than the described range $[-2^{53}+1, 2^{53}-1]$.
- All integers outside the range $[-2^{63}, 2^{64}-1]$, as well as floating-point numbers are stored as `double`.
This also concurs with the specification above.
### Zeros
+4
View File
@@ -349,6 +349,7 @@ markdown_extensions:
- toc:
permalink: true
- md_in_html
- pymdownx.arithmatex
- pymdownx.betterem:
smart_enable: all
- pymdownx.caret
@@ -440,3 +441,6 @@ plugins:
extra_css:
- css/custom.css
extra_javascript:
- https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.0/MathJax.js?config=TeX-MML-AM_CHTML
+5 -5
View File
@@ -79,9 +79,9 @@ def check_structure() -> None:
report("whitespace/line_length", f"{file}:{lineno+1} ({current_section})", f"line is too long ({len(line)} vs. 160 chars)")
# sections in `<!-- NOLINT -->` comments are treated as present
nolint_match = re.match(r"<!--\s*NOLINT\s+(.*?)\s*-->", line)
if nolint_match:
current_section = nolint_match.group(1)
if line.startswith("<!-- NOLINT"):
current_section = line.strip("<!-- NOLINT")
current_section = current_section.strip(" -->")
existing_sections.append(current_section)
# check if sections are correct
@@ -97,7 +97,7 @@ def check_structure() -> None:
if len(unexpected):
report("style/numbering", f"{file}:{lineno} ({current_section})", f'unexpected overloads: {", ".join([f"({x})" for x in unexpected])}')
current_section = line[3:]
current_section = line.strip("## ")
existing_sections.append(current_section)
if current_section in expected_sections:
@@ -141,7 +141,7 @@ def check_structure() -> None:
# check that non-example admonitions have titles
untitled_admonition = re.match(r"^(\?\?\?|!!!) ([^ ]+)$", line)
if untitled_admonition and untitled_admonition.group(2) != "example":
report("style/admonition_title", f"{file}:{lineno+1} ({current_section})", f'"{untitled_admonition.group(2)}" admonitions should have a title')
report("style/admonition_title", f"{file}:{lineno} ({current_section})", f'"{untitled_admonition.group(2)}" admonitions should have a title')
previous_line = line
@@ -3821,6 +3821,11 @@ class binary_reader
if (0xC2 <= byte && byte <= 0xF7)
{
const auto second = get_bon8();
if (second == char_traits<char_type>::eof())
{
// the input ends inside a character or an integer
return unexpect_eof(input_format_t::bon8, "key");
}
unget_bon8(second);
if (is_bon8_continuation(second))
{
@@ -3919,6 +3924,12 @@ class binary_reader
// a lead byte ends the string if no continuation byte follows: it
// is then the first byte of an integer
const auto second = get_bon8();
if (second == char_traits<char_type>::eof())
{
// the input ends inside a character or an integer: either
// way, the message is incomplete
return unexpect_eof(input_format_t::bon8, "string");
}
if (!is_bon8_continuation(second))
{
unget_bon8(second);
+11
View File
@@ -16588,6 +16588,11 @@ class binary_reader
if (0xC2 <= byte && byte <= 0xF7)
{
const auto second = get_bon8();
if (second == char_traits<char_type>::eof())
{
// the input ends inside a character or an integer
return unexpect_eof(input_format_t::bon8, "key");
}
unget_bon8(second);
if (is_bon8_continuation(second))
{
@@ -16686,6 +16691,12 @@ class binary_reader
// a lead byte ends the string if no continuation byte follows: it
// is then the first byte of an integer
const auto second = get_bon8();
if (second == char_traits<char_type>::eof())
{
// the input ends inside a character or an integer: either
// way, the message is incomplete
return unexpect_eof(input_format_t::bon8, "string");
}
if (!is_bon8_continuation(second))
{
unget_bon8(second);
+41
View File
@@ -531,6 +531,47 @@ TEST_CASE("BON8")
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x87, 'a'}), "[json.exception.parse_error.110] parse error at byte 3: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
}
SECTION("input that ends after a UTF-8 lead byte")
{
// the lead byte begins either a character or an integer; both are
// incomplete, so the lead byte must not end the string before it
for (const bool strict :
{
true, false
})
{
CAPTURE(strict)
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{'a', 0xC3}, strict), "[json.exception.parse_error.110] parse error at byte 3: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x81, 'a', 0xC3}, strict), "[json.exception.parse_error.110] parse error at byte 4: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x81, 'a', 0xE2}, strict), "[json.exception.parse_error.110] parse error at byte 4: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x81, 'a', 0xF0}, strict), "[json.exception.parse_error.110] parse error at byte 4: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x81, 0xC3, 0xA9, 0xC3}, strict), "[json.exception.parse_error.110] parse error at byte 5: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x87, 'a', 0xC3}, strict), "[json.exception.parse_error.110] parse error at byte 4: syntax error while parsing BON8 string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x87, 0xC3}, strict), "[json.exception.parse_error.110] parse error at byte 3: syntax error while parsing BON8 key: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bon8(bytes{0x88, 'a', 0x91, 0xE2}, strict), "[json.exception.parse_error.110] parse error at byte 5: syntax error while parsing BON8 key: unexpected end of input", json::parse_error&);
}
}
SECTION("a message that is cut off is not read as a shorter value")
{
const json values = {"\xC3\xA9", "a\xE2\x82\xAC", "\xF0\x9F\x98\x80\xC3\xA9", {"a\xC3\xA9"}, {{"\xC3\xA9", "\xE2\x82\xAC"}}, {{"a", {"b\xC3\xA9", 1}}}};
for (const auto& j : values)
{
const bytes message = json::to_bon8(j);
for (std::size_t length = 0; length < message.size(); ++length)
{
CAPTURE(j)
CAPTURE(length)
bytes prefix = message;
prefix.resize(length);
CHECK(json::from_bon8(prefix, false, false).is_discarded());
// a stream is read byte by byte rather than in bulk
std::istringstream stream(str(prefix));
CHECK(json::from_bon8(stream, false, false).is_discarded());
}
}
}
SECTION("invalid UTF-8")
{
// overlong