Compare commits

...
Author SHA1 Message Date
Claude cf14cbdd02 Deduplicate UBJSON/BJData length-type error message
The "expected length type specification" parse error was built with an
identical if/else block in both get_ubjson_string() and
get_ubjson_size_value(), each branching on the input format to pick the
marker list. Extract a single unexpected_length_type_message() helper that
selects the markers via a ternary and appends the caller-supplied infix,
removing four near-duplicate message strings and centralizing the marker
list. Error messages are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tv6SnfYbcx8y3y4PrECcto
2026-07-22 21:03:03 +00:00
Angadi56 3565f40229 reject negative UBJSON/BJData string length (#5284) 2026-07-20 20:32:38 +00:00
Angadi56 1c5a953de5 reject unpaired UTF-16 surrogates in wide-string input (#5276) 2026-07-19 18:33:42 +02:00
dependabot[bot] a03e65420c ⬆️ Bump the codeql-action group with 4 updates (#5275) 2026-07-18 13:59:20 +02:00
dependabot[bot] d6ede37088 ⬆️ Bump actions/stale from 10.3.0 to 10.4.0 (#5277) 2026-07-18 13:58:45 +02:00
Niels Lohmann 722c03495f Extend std::optional null regression coverage (assignment, get_to) (#5269) 2026-07-12 09:16:25 +02:00
Niels Lohmann c197feff81 Extend memcpy fast path to sized sentinels (e.g. std::counted_iterator) (#5268) 2026-07-12 09:16:06 +02:00
Niels Lohmann b2b47c69b1 📝 Document std::pair/std::tuple conversion and C++20 range-view construction (#5271)
Both had zero documentation anywhere in docs/mkdocs/. The tuple/pair
gap was first spotted in the very first git-log audit pass but never
turned into an actionable todo, so it persisted uncaught across four
subsequent passes.

- Document basic positional std::pair/std::tuple <-> json array
  conversion, plus #5016's reference-extraction capability
  (get<std::tuple<T&, T&>>() returning references into the stored
  array elements).
- Document #5205's new json-from-C++20-range-view constructor
  (e.g. nums | std::views::filter(...)).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-11 23:58:49 +02:00
Niels LohmannandClaude Sonnet 5 6a406ee141 Document that JSON_Diagnostics CMake option doesn't apply to pre-installed packages (#5270)
Closes #3106. set(JSON_Diagnostics ON) before find_package() has no
effect on a package built and installed elsewhere (Homebrew, vcpkg, a
system package, etc.) -- the compile definition is baked into the
exported nlohmann_jsonTargets.cmake at install time and the generated
config script never re-reads that variable. Verified empirically
against the real Homebrew-installed 3.12.0 package: the exported
target carries a fixed $<$<BOOL:OFF>:JSON_DIAGNOSTICS=1>, and the
suggested set(JSON_Diagnostics ON) snippet produces no change in
exception output.

Documents the actual working fix (overriding the imported target's
INTERFACE_COMPILE_DEFINITIONS property after find_package()) and the
multi-target "JSON_DIAGNOSTICS redefined" pitfall reported earlier in
the issue thread.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-11 23:58:18 +02:00
24 changed files with 377 additions and 98 deletions
+3 -3
View File
@@ -38,14 +38,14 @@ jobs:
# Initializes the CodeQL tools for scanning. # Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL - name: Initialize CodeQL
uses: github/codeql-action/init@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3 uses: github/codeql-action/init@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with: with:
languages: c-cpp languages: c-cpp
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java). # Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below) # If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild - name: Autobuild
uses: github/codeql-action/autobuild@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3 uses: github/codeql-action/autobuild@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
- name: Perform CodeQL Analysis - name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3 uses: github/codeql-action/analyze@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
+1 -1
View File
@@ -43,6 +43,6 @@ jobs:
output: 'flawfinder_results.sarif' output: 'flawfinder_results.sarif'
- name: Upload analysis results to GitHub Security tab - name: Upload analysis results to GitHub Security tab
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3 uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with: with:
sarif_file: ${{github.workspace}}/flawfinder_results.sarif sarif_file: ${{github.workspace}}/flawfinder_results.sarif
+1 -1
View File
@@ -76,6 +76,6 @@ jobs:
# Upload the results to GitHub's code scanning dashboard. # Upload the results to GitHub's code scanning dashboard.
- name: "Upload to code-scanning" - name: "Upload to code-scanning"
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3 uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with: with:
sarif_file: results.sarif sarif_file: results.sarif
+1 -1
View File
@@ -61,7 +61,7 @@ jobs:
# Upload SARIF file generated in previous step # Upload SARIF file generated in previous step
- name: Upload SARIF file - name: Upload SARIF file
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3 uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with: with:
sarif_file: semgrep.sarif sarif_file: semgrep.sarif
if: always() if: always()
+1 -1
View File
@@ -20,7 +20,7 @@ jobs:
with: with:
egress-policy: audit egress-policy: audit
- uses: actions/stale@eb5cf3af3ac0a1aa4c9c45633dd1ae542a27a899 # v10.3.0 - uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v10.4.0
with: with:
stale-issue-label: 'state: stale' stale-issue-label: 'state: stale'
stale-pr-label: 'state: stale' stale-pr-label: 'state: stale'
+1 -1
View File
@@ -49,7 +49,7 @@ Unlike the [`parse()`](parse.md) function, this function neither throws an excep
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance. : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
- a custom sentinel type for C++20 ranges - a custom sentinel type for C++20 ranges
- `std::counted_iterator` with a different sentinel type - `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
@@ -36,8 +36,10 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: a compatible iterator type : a compatible iterator type
`SentinelType` `SentinelType`
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance a : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
custom sentinel type for C++20 ranges
- a custom sentinel type for C++20 ranges
- `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
+4 -2
View File
@@ -36,8 +36,10 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: a compatible iterator type : a compatible iterator type
`SentinelType` `SentinelType`
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance a : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
custom sentinel type for C++20 ranges
- a custom sentinel type for C++20 ranges
- `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
+4 -2
View File
@@ -39,8 +39,10 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: a compatible iterator type : a compatible iterator type
`SentinelType` `SentinelType`
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance a : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
custom sentinel type for C++20 ranges
- a custom sentinel type for C++20 ranges
- `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
@@ -36,8 +36,10 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: a compatible iterator type : a compatible iterator type
`SentinelType` `SentinelType`
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance a : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
custom sentinel type for C++20 ranges
- a custom sentinel type for C++20 ranges
- `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
@@ -36,8 +36,10 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: a compatible iterator type : a compatible iterator type
`SentinelType` `SentinelType`
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance a : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
custom sentinel type for C++20 ranges
- a custom sentinel type for C++20 ranges
- `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
+1 -1
View File
@@ -48,7 +48,7 @@ static basic_json parse(IteratorType first, SentinelType last,
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance. : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for instance.
- a custom sentinel type for C++20 ranges - a custom sentinel type for C++20 ranges
- `std::counted_iterator` with a different sentinel type - `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
## Parameters ## Parameters
+4 -1
View File
@@ -48,7 +48,10 @@ The SAX event lister must follow the interface of [`json_sax`](../json_sax/index
with a size of 1, 2, or 4 bytes (interpreted respectively as UTF-8, UTF-16, and UTF-32) with a size of 1, 2, or 4 bytes (interpreted respectively as UTF-8, UTF-16, and UTF-32)
`SentinelType` `SentinelType`
: defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for overload (2) : defaults to `IteratorType`; may be a different type comparable to `IteratorType` via `operator!=`, for overload (2), for instance.
- a custom sentinel type for C++20 ranges
- `std::default_sentinel_t`, when `IteratorType` is `std::counted_iterator`
`SAX` `SAX`
: a class fulfilling the SAX event listener interface; see [`json_sax`](../json_sax/index.md) : a class fulfilling the SAX event listener interface; see [`json_sax`](../json_sax/index.md)
@@ -38,7 +38,8 @@ When the macro is not defined, the library will define it to its default value.
Diagnostic messages can also be controlled with the CMake option Diagnostic messages can also be controlled with the CMake option
[`JSON_Diagnostics`](../../integration/cmake.md#json_diagnostics) (`OFF` by default) [`JSON_Diagnostics`](../../integration/cmake.md#json_diagnostics) (`OFF` by default)
which defines `JSON_DIAGNOSTICS` accordingly. which defines `JSON_DIAGNOSTICS` accordingly. Note this only applies when building the
library from source — see the pre-installed-package caveat on that page.
## Examples ## Examples
+36
View File
@@ -47,6 +47,28 @@ json j = {{"one", 1}, {"two", 2}};
auto m = j.get<std::map<std::string, int>>(); // {{"one", 1}, {"two", 2}} auto m = j.get<std::map<std::string, int>>(); // {{"one", 1}, {"two", 2}}
``` ```
`#!cpp std::pair` and `#!cpp std::tuple` are also supported, converting positionally to and from a JSON array:
```cpp
json j = {1.0, "hello", 42};
auto t = j.get<std::tuple<double, std::string, int>>(); // {1.0, "hello", 42}
```
!!! info "Extracting references into a tuple"
A tuple type may also hold references (e.g. `#!cpp std::tuple<double&, std::string&>`) to avoid copying: `get`
then returns a tuple of references pointing directly at the elements stored inside the `basic_json` array,
rather than a tuple of copies:
```cpp
json j = {1.0, "hello"};
auto refs = j.get<std::tuple<double&, std::string&>>();
std::get<1>(refs) = "world"; // modifies j[1] in place
```
A referenced type must be one the library actually stores (or an arithmetic type it can convert to/from);
otherwise this is a compile error.
## Implicit conversions ## Implicit conversions
By default, a JSON value implicitly converts to a compatible C++ type, so the explicit `get` call can often be omitted: By default, a JSON value implicitly converts to a compatible C++ type, so the explicit `get` call can often be omitted:
@@ -136,6 +158,20 @@ std::vector<int> numbers = {1, 2, 3};
json j = numbers; // [1,2,3] json j = numbers; // [1,2,3]
``` ```
!!! info "Constructing from a C++20 range view"
A `json` array can also be constructed directly from a C++20 range view (`std::ranges::view`), such as the result
of `std::views::filter` or `std::views::transform` -- no intermediate container is needed:
```cpp
std::vector<int> nums{1, 2, 37, 42, 21};
auto filtered = nums | std::views::filter([](int i) { return i > 10; });
json j(filtered); // [37,42,21]
```
This requires [`JSON_HAS_RANGES`](../api/macros/json_has_ranges.md) to be enabled and is unavailable on MinGW due
to incomplete C++20 ranges support there.
## Your own types ## Your own types
The conversions above are built in for standard types. To make the same syntax work for **your own** types, provide The conversions above are built in for standard types. To make the same syntax work for **your own** types, provide
+25
View File
@@ -135,6 +135,31 @@ Enable CI build targets. The exact targets are used during the several CI steps
Enable [extended diagnostic messages](../home/exceptions.md#extended-diagnostic-messages) by defining macro [`JSON_DIAGNOSTICS`](../api/macros/json_diagnostics.md). This option is `OFF` by default. Enable [extended diagnostic messages](../home/exceptions.md#extended-diagnostic-messages) by defining macro [`JSON_DIAGNOSTICS`](../api/macros/json_diagnostics.md). This option is `OFF` by default.
!!! warning "Does not apply to a pre-installed package"
This option only takes effect when building nlohmann/json from source as part of your own
CMake project (e.g. via [`FetchContent`](#fetchcontent) or [`add_subdirectory`](#external)).
It has **no effect** on a package that was already built and installed elsewhere (Homebrew,
vcpkg, a system package, etc.) — the resulting compile definition is baked into the exported
`nlohmann_jsonTargets.cmake` at install time, and `set(JSON_Diagnostics ON)` before
`find_package()` does not change it (verified against the Homebrew-installed package: the
exported target still carries a fixed `$<$<BOOL:OFF>:JSON_DIAGNOSTICS=1>`, regardless of any
variable set in the consuming project).
To enable extended diagnostics for a pre-installed package, override the imported target's
property directly after `find_package()`:
```cmake
find_package(nlohmann_json REQUIRED)
set_target_properties(nlohmann_json::nlohmann_json PROPERTIES
INTERFACE_COMPILE_DEFINITIONS "JSON_DIAGNOSTICS=1")
```
This only works cleanly when your project is the sole consumer of that imported target. If
nlohmann_json is pulled in from more than one place in your dependency graph with different
`JSON_DIAGNOSTICS` values, you may see a `"JSON_DIAGNOSTICS" redefined` compiler error, since
conflicting `-D` flags can end up on the same compile command line.
### `JSON_Diagnostic_Positions` ### `JSON_Diagnostic_Positions`
Enable position diagnostics by defining macro [`JSON_DIAGNOSTIC_POSITIONS`](../api/macros/json_diagnostic_positions.md). This option is `OFF` by default. Enable position diagnostics by defining macro [`JSON_DIAGNOSTIC_POSITIONS`](../api/macros/json_diagnostic_positions.md). This option is `OFF` by default.
+49 -26
View File
@@ -1846,6 +1846,47 @@ class binary_reader
return get_ubjson_value(get_char ? get_ignore_noop() : current); return get_ubjson_value(get_char ? get_ignore_noop() : current);
} }
/*!
@brief reject a negative UBJSON/BJData string length
String and key lengths are written with signed integer markers (i, I, l,
L). A negative value is malformed; without this check get_string() would
silently treat it as an empty string and leave the following bytes to be
misread as the next value. This mirrors the non-negative check the
optimized-container count path already performs in get_ubjson_size_value.
@param[in] len the string length read from the input
@return whether the length is valid (non-negative)
*/
template<typename NumberType>
bool check_ubjson_string_length(const NumberType len)
{
if (JSON_HEDLEY_UNLIKELY(len < 0))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(113, chars_read,
exception_message(input_format, "string length must not be negative", "string"), nullptr));
}
return true;
}
/*!
@brief create the error message for a missing/invalid length-type marker
Used both when reading a UBJSON/BJData string and when reading an optimized
container size. The accepted markers depend on the input format, and the
size variant appends " after '#'" via @a infix.
@param[in] last_token the offending byte as returned by get_token_string()
@param[in] infix extra context inserted after the marker list
@return the formatted error message
*/
std::string unexpected_length_type_message(const std::string& last_token, const char* infix) const
{
return concat("expected length type specification (",
input_format == input_format_t::bjdata ? "U, i, u, I, m, l, M, L" : "U, i, I, l, L",
")", infix, "; last byte: 0x", last_token);
}
/*! /*!
@brief reads a UBJSON string @brief reads a UBJSON string
@@ -1883,25 +1924,25 @@ class binary_reader
case 'i': case 'i':
{ {
std::int8_t len{}; std::int8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'I': case 'I':
{ {
std::int16_t len{}; std::int16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'l': case 'l':
{ {
std::int32_t len{}; std::int32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'L': case 'L':
{ {
std::int64_t len{}; std::int64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'u': case 'u':
@@ -1938,17 +1979,8 @@ class binary_reader
break; break;
} }
auto last_token = get_token_string(); auto last_token = get_token_string();
std::string message; return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, ""), "string"), nullptr));
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L); last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "string"), nullptr));
} }
/*! /*!
@@ -2227,17 +2259,8 @@ class binary_reader
break; break;
} }
auto last_token = get_token_string(); auto last_token = get_token_string();
std::string message; return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, " after '#'"), "size"), nullptr));
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L) after '#'; last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "size"), nullptr));
} }
/*! /*!
@@ -220,20 +220,29 @@ class iterator_input_adapter
// whether IteratorType refers to a contiguous range and therefore supports // whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect // a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string). // library iterators such as those of std::vector and std::string).
// The fast path also requires SentinelType == IteratorType so std::distance works. // Computing the available element count needs either same-type iterators
// (plain std::distance) or, in C++20, a sized sentinel (std::ranges::distance),
// e.g. std::counted_iterator paired with std::default_sentinel_t.
static constexpr bool iterator_is_contiguous = static constexpr bool iterator_is_contiguous =
std::is_same<IteratorType, SentinelType>::value && (
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20) #if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::contiguous_iterator<IteratorType> || (std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>)
&& (std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
#endif #endif
std::is_pointer<IteratorType>::value);
// contiguous fast path: bulk copy the remaining range with std::memcpy // contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T> template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/) std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{ {
const std::size_t wanted = count * sizeof(T); const std::size_t wanted = count * sizeof(T);
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
const std::size_t available = static_cast<std::size_t>(std::ranges::distance(current, end)) * sizeof(char_type);
#else
const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type); const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type);
#endif
const std::size_t copied = (std::min)(wanted, available); const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0)) if (JSON_HEDLEY_LIKELY(copied != 0))
{ {
@@ -386,17 +395,30 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
} }
else else
{ {
if (JSON_HEDLEY_UNLIKELY(!input.empty())) // A supplementary code point is a high surrogate (0xD800..0xDBFF)
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
// surrogate, a high surrogate at the end of the input, or a high
// surrogate followed by any other unit is malformed UTF-16. In
// that case the offending unit is passed through unchanged so the
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
// are handled in the lexer.
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{ {
const auto wc2 = static_cast<unsigned int>(input.get_character()); const auto wc2 = static_cast<unsigned int>(input.get_character());
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu)); if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u)); {
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu)); const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu)); utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu)); utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes_filled = 4; utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
valid_pair = true;
}
} }
else
if (!valid_pair)
{ {
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc); utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1; utf8_bytes_filled = 1;
+83 -38
View File
@@ -7207,20 +7207,29 @@ class iterator_input_adapter
// whether IteratorType refers to a contiguous range and therefore supports // whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect // a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string). // library iterators such as those of std::vector and std::string).
// The fast path also requires SentinelType == IteratorType so std::distance works. // Computing the available element count needs either same-type iterators
// (plain std::distance) or, in C++20, a sized sentinel (std::ranges::distance),
// e.g. std::counted_iterator paired with std::default_sentinel_t.
static constexpr bool iterator_is_contiguous = static constexpr bool iterator_is_contiguous =
std::is_same<IteratorType, SentinelType>::value && (
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20) #if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::contiguous_iterator<IteratorType> || (std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>)
&& (std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
#endif #endif
std::is_pointer<IteratorType>::value);
// contiguous fast path: bulk copy the remaining range with std::memcpy // contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T> template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/) std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{ {
const std::size_t wanted = count * sizeof(T); const std::size_t wanted = count * sizeof(T);
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
const std::size_t available = static_cast<std::size_t>(std::ranges::distance(current, end)) * sizeof(char_type);
#else
const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type); const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type);
#endif
const std::size_t copied = (std::min)(wanted, available); const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0)) if (JSON_HEDLEY_LIKELY(copied != 0))
{ {
@@ -7373,17 +7382,30 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
} }
else else
{ {
if (JSON_HEDLEY_UNLIKELY(!input.empty())) // A supplementary code point is a high surrogate (0xD800..0xDBFF)
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
// surrogate, a high surrogate at the end of the input, or a high
// surrogate followed by any other unit is malformed UTF-16. In
// that case the offending unit is passed through unchanged so the
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
// are handled in the lexer.
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{ {
const auto wc2 = static_cast<unsigned int>(input.get_character()); const auto wc2 = static_cast<unsigned int>(input.get_character());
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu)); if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u)); {
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu)); const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu)); utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu)); utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes_filled = 4; utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
valid_pair = true;
}
} }
else
if (!valid_pair)
{ {
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc); utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1; utf8_bytes_filled = 1;
@@ -12373,6 +12395,47 @@ class binary_reader
return get_ubjson_value(get_char ? get_ignore_noop() : current); return get_ubjson_value(get_char ? get_ignore_noop() : current);
} }
/*!
@brief reject a negative UBJSON/BJData string length
String and key lengths are written with signed integer markers (i, I, l,
L). A negative value is malformed; without this check get_string() would
silently treat it as an empty string and leave the following bytes to be
misread as the next value. This mirrors the non-negative check the
optimized-container count path already performs in get_ubjson_size_value.
@param[in] len the string length read from the input
@return whether the length is valid (non-negative)
*/
template<typename NumberType>
bool check_ubjson_string_length(const NumberType len)
{
if (JSON_HEDLEY_UNLIKELY(len < 0))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(113, chars_read,
exception_message(input_format, "string length must not be negative", "string"), nullptr));
}
return true;
}
/*!
@brief create the error message for a missing/invalid length-type marker
Used both when reading a UBJSON/BJData string and when reading an optimized
container size. The accepted markers depend on the input format, and the
size variant appends " after '#'" via @a infix.
@param[in] last_token the offending byte as returned by get_token_string()
@param[in] infix extra context inserted after the marker list
@return the formatted error message
*/
std::string unexpected_length_type_message(const std::string& last_token, const char* infix) const
{
return concat("expected length type specification (",
input_format == input_format_t::bjdata ? "U, i, u, I, m, l, M, L" : "U, i, I, l, L",
")", infix, "; last byte: 0x", last_token);
}
/*! /*!
@brief reads a UBJSON string @brief reads a UBJSON string
@@ -12410,25 +12473,25 @@ class binary_reader
case 'i': case 'i':
{ {
std::int8_t len{}; std::int8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'I': case 'I':
{ {
std::int16_t len{}; std::int16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'l': case 'l':
{ {
std::int32_t len{}; std::int32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'L': case 'L':
{ {
std::int64_t len{}; std::int64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result); return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
} }
case 'u': case 'u':
@@ -12465,17 +12528,8 @@ class binary_reader
break; break;
} }
auto last_token = get_token_string(); auto last_token = get_token_string();
std::string message; return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, ""), "string"), nullptr));
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L); last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "string"), nullptr));
} }
/*! /*!
@@ -12754,17 +12808,8 @@ class binary_reader
break; break;
} }
auto last_token = get_token_string(); auto last_token = get_token_string();
std::string message; return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, unexpected_length_type_message(last_token, " after '#'"), "size"), nullptr));
if (input_format != input_format_t::bjdata)
{
message = "expected length type specification (U, i, I, l, L) after '#'; last byte: 0x" + last_token;
}
else
{
message = "expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x" + last_token;
}
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read, exception_message(input_format, message, "size"), nullptr));
} }
/*! /*!
+13
View File
@@ -2721,6 +2721,19 @@ TEST_CASE("BJData")
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing BJData string: expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x31", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing BJData string: expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x31", json::parse_error&);
} }
SECTION("negative length")
{
json _;
std::vector<uint8_t> const vi = {'S', 'i', 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vi), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative", json::parse_error&);
CHECK(json::from_bjdata(vi, true, false).is_discarded());
std::vector<uint8_t> const vl = {'S', 'l', 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vl), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing BJData string: string length must not be negative", json::parse_error&);
CHECK(json::from_bjdata(vl, true, false).is_discarded());
}
SECTION("parse bjdata markers in ubjson") SECTION("parse bjdata markers in ubjson")
{ {
// create a single-character string for all number types // create a single-character string for all number types
+15
View File
@@ -1782,6 +1782,21 @@ TEST_CASE("std::optional")
"[json.exception.type_error.302] type must be string, but is null", json::type_error&); "[json.exception.type_error.302] type must be string, but is null", json::type_error&);
CHECK_THROWS_WITH_AS(std::optional<int>(j_null), CHECK_THROWS_WITH_AS(std::optional<int>(j_null),
"[json.exception.type_error.302] type must be number, but is null", json::type_error&); "[json.exception.type_error.302] type must be number, but is null", json::type_error&);
// Assignment goes through the same overload resolution as direct
// construction, so it throws for the same reason. This relies on
// basic_json's implicit conversion operator, so it only applies
// when JSON_USE_IMPLICIT_CONVERSIONS is enabled (the default).
#if JSON_USE_IMPLICIT_CONVERSIONS
std::optional<std::string> opt_assign;
CHECK_THROWS_WITH_AS(opt_assign = j_null,
"[json.exception.type_error.302] type must be string, but is null", json::type_error&);
#endif
// get_to() is the correct way to obtain std::nullopt from a JSON null.
std::optional<std::string> opt_get_to = "placeholder";
j_null.get_to(opt_get_to);
CHECK(opt_get_to == std::nullopt);
} }
SECTION("string") SECTION("string")
+25
View File
@@ -1862,6 +1862,31 @@ TEST_CASE("UBJSON")
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON string: expected length type specification (U, i, I, l, L); last byte: 0x31", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON string: expected length type specification (U, i, I, l, L); last byte: 0x31", json::parse_error&);
} }
SECTION("negative length")
{
json _;
std::vector<uint8_t> const vi = {'S', 'i', 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vi), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vi, true, false).is_discarded());
std::vector<uint8_t> const vI = {'S', 'I', 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vI), "[json.exception.parse_error.113] parse error at byte 4: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vI, true, false).is_discarded());
std::vector<uint8_t> const vl = {'S', 'l', 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vl), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vl, true, false).is_discarded());
std::vector<uint8_t> const vL = {'S', 'L', 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vL), "[json.exception.parse_error.113] parse error at byte 10: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vL, true, false).is_discarded());
// a length of zero remains valid and yields an empty string
std::vector<uint8_t> const v0 = {'S', 'i', 0};
CHECK(json::from_ubjson(v0) == json(""));
}
} }
SECTION("array") SECTION("array")
+29
View File
@@ -6,6 +6,13 @@
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me> // SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT // SPDX-License-Identifier: MIT
// cmake/test.cmake selects the C++ standard versions with which to build a
// unit test based on the presence of JSON_HAS_CPP_<VERSION> macros.
// When using macros that are only defined for particular versions of the standard
// (e.g., JSON_HAS_FILESYSTEM for C++17 and up), please mention the corresponding
// version macro in a comment close by, like this:
// JSON_HAS_CPP_<VERSION> (do not remove; see note at top of file)
#include "doctest_compatibility.h" #include "doctest_compatibility.h"
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
@@ -13,6 +20,10 @@ using nlohmann::json;
#include <list> #include <list>
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
#include <iterator>
#endif
namespace namespace
{ {
TEST_CASE("Use arbitrary stdlib container") TEST_CASE("Use arbitrary stdlib container")
@@ -201,4 +212,22 @@ TEST_CASE("Parse with heterogeneous iterator and sentinel types")
CHECK(j2.at(0) == 1); CHECK(j2.at(0) == 1);
} }
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// JSON_HAS_CPP_20 (do not remove; see note at top of file)
TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
{
using iterator_type = std::string::const_iterator;
const std::string json_str = R"({"key":"value","array":[1,2,3]})";
const auto len = static_cast<std::iter_difference_t<iterator_type>>(json_str.size());
const std::counted_iterator<iterator_type> first(json_str.begin(), len);
const json j = json::parse(first, std::default_sentinel);
CHECK(j["key"] == "value");
CHECK(j["array"].size() == 3);
const std::counted_iterator<iterator_type> first2(json_str.begin(), len);
CHECK(json::accept(first2, std::default_sentinel));
}
#endif
} // namespace } // namespace
+33 -1
View File
@@ -53,6 +53,27 @@ TEST_CASE("wide strings")
std::wstring const w = L"\"\xDBFF"; std::wstring const w = L"\"\xDBFF";
json _; json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&); CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// the exact message depends on the width of wchar_t: a 16-bit
// wchar_t passes the lone surrogate to the UTF-8 decoder unchanged
// (rejected as a single ill-formed byte at column 2), while a
// 32-bit wchar_t first encodes it as an ill-formed three-byte
// sequence (rejected one byte later, at column 3)
const char* const error_low_surrogate = sizeof(wchar_t) == 2
? "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'"
: "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xED\xB0'";
const char* const error_high_surrogate = sizeof(wchar_t) == 2
? "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'"
: "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xED\xA0'";
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xDC00), L'"'}), error_low_surrogate, json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xD800), L'a', L'"'}), error_high_surrogate, json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xDC00), L'a', L'"'}), error_low_surrogate, json::parse_error&);
} }
} }
@@ -68,11 +89,22 @@ TEST_CASE("wide strings")
SECTION("invalid std::u16string") SECTION("invalid std::u16string")
{ {
if (wstring_is_utf16()) if (u16string_is_utf16())
{ {
std::u16string const w = u"\"\xDBFF"; std::u16string const w = u"\"\xDBFF";
json _; json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&); CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a valid surrogate pair is still decoded (U+1F600)
CHECK(json::parse(std::u16string{u'"', 0xD83D, 0xDE00, u'"'}).get<std::string>() == "\xF0\x9F\x98\x80");
} }
} }