Compare commits

..
Author SHA1 Message Date
Claude cf14cbdd02 Deduplicate UBJSON/BJData length-type error message
The "expected length type specification" parse error was built with an
identical if/else block in both get_ubjson_string() and
get_ubjson_size_value(), each branching on the input format to pick the
marker list. Extract a single unexpected_length_type_message() helper that
selects the markers via a ternary and appends the caller-supplied infix,
removing four near-duplicate message strings and centralizing the marker
list. Error messages are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tv6SnfYbcx8y3y4PrECcto
2026-07-22 21:03:03 +00:00
8 changed files with 45 additions and 231 deletions
-8
View File
@@ -15,14 +15,6 @@ guidance.
For vulnerabilities in third-party dependencies or modules, please report them directly to the respective maintainers.
## Unofficial packages
This project does not publish an official npm package. The npm package
[`nlohmann-json`](https://www.npmjs.com/package/nlohmann-json) (or similarly named packages) is not maintained or
endorsed by this project. See the
[package managers documentation](https://json.nlohmann.me/integration/package_managers/#npm) for supported
integration options.
## Additional Resources
- Explore security-related topics and contribute to tools and projects through
+1 -9
View File
@@ -163,15 +163,7 @@ jobs:
export DEBIAN_FRONTEND=noninteractive
apt-get update
apt-get install -y --no-install-recommends software-properties-common ca-certificates gnupg make git
# add-apt-repository resolves the PPA through the Launchpad API,
# which intermittently times out or fails the team lookup (the plain
# "deb ..." sources below never hit Launchpad and never flake).
# Retry with backoff so a transient Launchpad blip does not fail CI.
for attempt in 1 2 3 4 5; do
add-apt-repository -y ppa:ubuntu-toolchain-r/test && break
echo "::warning::add-apt-repository ppa:ubuntu-toolchain-r/test failed (attempt ${attempt}/5); retrying"
sleep $((attempt * 10))
done
add-apt-repository -y ppa:ubuntu-toolchain-r/test
apt-add-repository -y "deb http://archive.ubuntu.com/ubuntu/ bionic main"
apt-add-repository -y "deb http://archive.ubuntu.com/ubuntu/ bionic universe"
apt-add-repository -y "deb http://archive.ubuntu.com/ubuntu/ xenial main"
-11
View File
@@ -43,17 +43,6 @@ Strong guarantee: if an exception is thrown, there are no changes to any JSON va
Throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) if a string stored inside the JSON value
is not UTF-8 encoded and `error_handler` is set to `strict`
!!! warning "Serializing untrusted input"
When serializing values that may contain invalid or untrusted UTF-8 (e.g., bytes taken directly from network
input), `dump()` throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) in the default
`strict` mode. To serialize such data without throwing, pass
[`error_handler_t::replace`](error_handler_t.md) (substitutes U+FFFD) or
[`error_handler_t::ignore`](error_handler_t.md). Callers that serialize untrusted input on a crash-sensitive path
should either choose a non-strict error handler or wrap `dump()` in a `#!cpp try`/`#!cpp catch`.
See the [FAQ](../../home/faq.md#serializing-untrusted-or-invalid-utf-8) for details.
## Complexity
Linear.
-21
View File
@@ -194,27 +194,6 @@ The library uses `std::numeric_limits<number_float_t>::digits10` (15 for IEEE `d
See [this section](../features/types/number_handling.md#number-serialization) on the library's number handling for more information.
### Serializing untrusted or invalid UTF-8
!!! question "Questions"
- Why does `dump()` throw when I serialize data that came from the network?
- Is CVE-2024-34363 a vulnerability in this library?
Crashes reported against this library that stem from an uncaught
[`type_error.316`](exceptions.md#jsonexceptiontype_error316) while serializing unvalidated input (e.g.,
CVE-2024-34363) are a usage issue, not a library vulnerability:
[`dump()`](../api/basic_json/dump.md) throws in its default `strict` mode because
[RFC 8259](https://datatracker.ietf.org/doc/html/rfc8259#section-8.1) requires JSON text to be valid UTF-8.
The recommended pattern is to pass a non-strict [`error_handler`](../api/basic_json/error_handler_t.md) or to handle the
exception:
```cpp
// replace invalid sequences with U+FFFD instead of throwing
const auto s = j.dump(-1, ' ', false, json::error_handler_t::replace);
```
### Using JSON values with `std::format` or `fmt`
!!! question
@@ -930,12 +930,6 @@ If you are using [CocoaPods](https://cocoapods.org), you can use the library by
to your podfile (see [an example](https://bitbucket.org/benman/nlohmann_json-cocoapod/src/master/)). Please file issues
[here](https://bitbucket.org/benman/nlohmann_json-cocoapod/issues?status=new&status=open).
## npm
This project does not publish an official [npm](https://www.npmjs.com) package. The npm package
[`nlohmann-json`](https://www.npmjs.com/package/nlohmann-json) (or similarly named packages) is not maintained or
endorsed by this project. Use one of the package managers listed above, or integrate the single header directly.
## ESP-IDF and PlatformIO
There is no official package published to the [ESP-IDF Component Registry](https://components.espressif.com) or the
+22 -60
View File
@@ -163,39 +163,12 @@ class binary_reader
// BSON //
//////////
/*!
@brief Validate a BSON document's declared size against the bytes read.
A BSON document starts with an int32 that counts its own total length in
bytes, including that prefix and the trailing 0x00. The reader is driven
by the terminator rather than the declared length, so without this check a
nested document could declare a length that disagrees with where its
terminator actually falls and quietly hand the bytes in between to the
enclosing document. A well-formed document is at least 5 bytes (the prefix
plus the terminator); the equality also rejects those impossible sizes,
since at least 5 bytes are always consumed.
@param[in] document_start value of chars_read before the size prefix
@param[in] document_size the declared document size
@return whether the declared size matches the number of bytes read
*/
bool check_bson_document_size(const std::size_t document_start, const std::int32_t document_size)
{
if (JSON_HEDLEY_UNLIKELY(document_size < 0 || static_cast<std::size_t>(document_size) != chars_read - document_start))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format_t::bson, concat("document size ", std::to_string(document_size), " does not match the number of bytes read (", std::to_string(chars_read - document_start), ")"), "document"), nullptr));
}
return true;
}
/*!
@brief Reads in a BSON-object and passes it to the SAX-parser.
@return whether a valid BSON-value was passed to the SAX parser
*/
bool parse_bson_internal()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -209,11 +182,6 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_object();
}
@@ -429,7 +397,6 @@ class binary_reader
*/
bool parse_bson_array()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -443,11 +410,6 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_array();
}
@@ -1907,6 +1869,24 @@ class binary_reader
return true;
}
/*!
@brief create the error message for a missing/invalid length-type marker
Used both when reading a UBJSON/BJData string and when reading an optimized
container size. The accepted markers depend on the input format, and the
size variant appends " after '#'" via @a infix.
@param[in] last_token the offending byte as returned by get_token_string()
@param[in] infix extra context inserted after the marker list
@return the formatted error message
*/
std::string unexpected_length_type_message(const std::string& last_token, const char* infix) const
{
return concat("expected length type specification (",
input_format == input_format_t::bjdata ? "U, i, u, I, m, l, M, L" : "U, i, I, l, L",
")", infix, "; last byte: 0x", last_token);
}
/*!
@brief reads a UBJSON string
@@ -1999,17 +1979,8 @@ class binary_reader
break;
}
auto last_token = get_token_string();
std::string message;
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L); last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "string"), nullptr));
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, ""), "string"), nullptr));
}
/*!
@@ -2288,17 +2259,8 @@ class binary_reader
break;
}
auto last_token = get_token_string();
std::string message;
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L) after '#'; last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "size"), nullptr));
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, " after '#'"), "size"), nullptr));
}
/*!
+22 -60
View File
@@ -10712,39 +10712,12 @@ class binary_reader
// BSON //
//////////
/*!
@brief Validate a BSON document's declared size against the bytes read.
A BSON document starts with an int32 that counts its own total length in
bytes, including that prefix and the trailing 0x00. The reader is driven
by the terminator rather than the declared length, so without this check a
nested document could declare a length that disagrees with where its
terminator actually falls and quietly hand the bytes in between to the
enclosing document. A well-formed document is at least 5 bytes (the prefix
plus the terminator); the equality also rejects those impossible sizes,
since at least 5 bytes are always consumed.
@param[in] document_start value of chars_read before the size prefix
@param[in] document_size the declared document size
@return whether the declared size matches the number of bytes read
*/
bool check_bson_document_size(const std::size_t document_start, const std::int32_t document_size)
{
if (JSON_HEDLEY_UNLIKELY(document_size < 0 || static_cast<std::size_t>(document_size) != chars_read - document_start))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format_t::bson, concat("document size ", std::to_string(document_size), " does not match the number of bytes read (", std::to_string(chars_read - document_start), ")"), "document"), nullptr));
}
return true;
}
/*!
@brief Reads in a BSON-object and passes it to the SAX-parser.
@return whether a valid BSON-value was passed to the SAX parser
*/
bool parse_bson_internal()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -10758,11 +10731,6 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_object();
}
@@ -10978,7 +10946,6 @@ class binary_reader
*/
bool parse_bson_array()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -10992,11 +10959,6 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_array();
}
@@ -12456,6 +12418,24 @@ class binary_reader
return true;
}
/*!
@brief create the error message for a missing/invalid length-type marker
Used both when reading a UBJSON/BJData string and when reading an optimized
container size. The accepted markers depend on the input format, and the
size variant appends " after '#'" via @a infix.
@param[in] last_token the offending byte as returned by get_token_string()
@param[in] infix extra context inserted after the marker list
@return the formatted error message
*/
std::string unexpected_length_type_message(const std::string& last_token, const char* infix) const
{
return concat("expected length type specification (",
input_format == input_format_t::bjdata ? "U, i, u, I, m, l, M, L" : "U, i, I, l, L",
")", infix, "; last byte: 0x", last_token);
}
/*!
@brief reads a UBJSON string
@@ -12548,17 +12528,8 @@ class binary_reader
break;
}
auto last_token = get_token_string();
std::string message;
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L); last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "string"), nullptr));
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, ""), "string"), nullptr));
}
/*!
@@ -12837,17 +12808,8 @@ class binary_reader
break;
}
auto last_token = get_token_string();
std::string message;
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L) after '#'; last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "size"), nullptr));
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, " after '#'"), "size"), nullptr));
}
/*!
-56
View File
@@ -854,62 +854,6 @@ TEST_CASE("Unsupported BSON input")
CHECK(!json::sax_parse(bson, &scp, json::input_format_t::bson));
}
TEST_CASE("BSON document size mismatch")
{
json _;
SECTION("top-level document declaring more bytes than it contains")
{
// empty object, but the length prefix claims 6 bytes instead of 5
std::vector<std::uint8_t> const input = {0x06, 0x00, 0x00, 0x00, 0x00};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("top-level document with a negative size")
{
std::vector<std::uint8_t> const input = {0xFF, 0xFF, 0xFF, 0xFF, 0x00};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size -1 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("embedded document whose size disagrees with its terminator")
{
// the embedded document "d" declares 0x7FFFFFFF bytes but its 0x00
// terminator falls right after {"a":null}; the length prefix would
// otherwise let the following "h" element be read as a member of the
// enclosing document instead of "d"
std::vector<std::uint8_t> const input =
{
0x00, 0x00, 0x00, 0x00, // outer size
0x03, 'd', 0x00, // entry: embedded document "d"
0xFF, 0xFF, 0xFF, 0x7F, // embedded size 0x7FFFFFFF
0x0A, 'a', 0x00, // entry: null "a"
0x00, // embedded end marker
0x08, 'h', 0x00, 0x01, // entry: bool "h" = true
0x00 // outer end marker
};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON document: document size 2147483647 does not match the number of bytes read (8)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("embedded array whose size disagrees with its terminator")
{
// array [42] is 12 bytes, but the length prefix claims 13
std::vector<std::uint8_t> const input =
{
0x00, 0x00, 0x00, 0x00, // outer size
0x04, 'a', 0x00, // entry: array "a"
0x0D, 0x00, 0x00, 0x00, // array size 13 (real is 12)
0x10, '0', 0x00, 0x2A, 0x00, 0x00, 0x00, // entry: int32 "0" = 42
0x00, // array end marker
0x00 // outer end marker
};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 19: syntax error while parsing BSON document: document size 13 does not match the number of bytes read (12)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
}
TEST_CASE("BSON numerical data")
{
SECTION("number")