Compare commits

..
Author SHA1 Message Date
Niels Lohmann 6d426f4e91 Keep converted object keys alive while writing UBJSON and BJData
Since #5746, write_ubjson and write_ubjson_iterative pass each object key
to sanitize_utf8_for_write and keep the returned reference. When
object_t::key_type is not string_t but converts to it, the argument is a
temporary that is destroyed at the end of the statement, and the
function returns a reference to it in every case but a sanitized copy,
so the key bytes are read from a dead object (AddressSanitizer:
stack-use-after-scope). Default json and ordered_json are unaffected.

Bind the key to a named object_key_string_t first: a reference when
key_type is string_t, so no copy is added there, and a converted copy
otherwise. A deleted overload of sanitize_utf8_for_write for anything
other than string_t turns a recurrence into a compile error.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 09:01:48 +02:00
Niels Lohmann d33068da73 Address review comments on the API stability docs and a test comment (#5784)
* Address review comments on #5775 and #5779

Allow new defaulted parameters and new default arguments in the API
stability rules, mention the macro opt-in, and drop the redundant
recompile advice. Describe test-diagnostics-optimized as the regression
test for the fixed #5742.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document what counts as a breaking change in the API stability rules

Spell out the 3.x compatibility rules in the roadmap: new defaulted
parameters, new default arguments, noexcept/constexpr, template
parameters, parse and dump results, accepted input, key iteration order,
iterator invalidation, implicit conversions, to_json/from_json lookup,
json_sax, value_t enumerators, and documented macros, CMake options and
headers. Also list std::hash values as not part of the public API, and
link the macro overview from the section.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 17:44:34 +02:00
Niels Lohmann d8dfc0d0f9 Fix unused-result warnings in the contains and parse_error examples (#5788)
contains(json_pointer) is marked JSON_HEDLEY_WARN_UNUSED_RESULT since #5477,
and parse() has been for longer. The contains example ignored the result in
two try blocks waiting for a parse_error that contains() never throws, so
they printed nothing; print the result for those pointers instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:21:59 +02:00
11 changed files with 241 additions and 127 deletions

No files matched your search

@@ -137,13 +137,10 @@ Strong exception safety: if an exception occurs, the original value stays intact
When the JSON pointer traverses intermediate levels that don't exist at all yet (not just a missing
leaf), each missing level is created as an array or an object depending on whether the corresponding
pointer token is a valid array index: the token `0`, a sequence of digits that does not begin with `0`,
or the token `-` creates an array, and every other token creates an object. For example, on an
initially `#!json null` value, `/foo/0/0/0` creates nested arrays, while `/foo/one/one/one` creates
nested objects. Tokens such as `01` or the empty token cannot be array indices (cf. RFC 6901, Sect. 4)
and therefore create objects, just as they would if the level already existed as an object. This is not
specified by the JSON Pointer RFC; it is this library's own, intentional disambiguation rule. See also
[JSON Pointer](../../features/json_pointer.md).
pointer token parses as a non-negative integer: a numeric token creates an array, a non-numeric token
creates an object. For example, on an initially `#!json null` value, `/foo/0/0/0` creates nested arrays,
while `/foo/one/one/one` creates nested objects. This is not specified by the JSON Pointer RFC; it is
this library's own, intentional disambiguation rule. See also [JSON Pointer](../../features/json_pointer.md).
!!! warning "Deprecation"
+23 -9
View File
@@ -36,13 +36,26 @@ work items are tracked in the [GitHub milestones](https://github.com/nlohmann/js
## API stability
Releases follow [semantic versioning](https://semver.org): a minor or patch release of version 3.x does not break code
that uses the public API. In particular, a 3.x release does not:
that uses the public API, unless that code opts in to a change with a macro as described [below](#version-40). In
particular, a 3.x release does not:
- change the signature of a function (its parameter types, return type, number of parameters, or the const-ness of a
member function);
- remove or rename a function or class;
- make breaking changes to the signature of a function: the types or order of its existing parameters, its return type,
its `noexcept` or `constexpr` specifier, or the const-ness of a member function. New parameters may be added if they
have a default value;
- remove or rename a function or class, or change the template parameters of a public class template;
- change which exceptions a function throws, or the [exception ids](../home/exceptions.md);
- change access specifiers or default arguments.
- change access specifiers, or change or remove existing default arguments. New default arguments may be added;
- change the JSON type that a valid input parses to, or the text that `dump()` produces for a valid value;
- accept input that was rejected before, or reject input that was accepted before;
- change the order in which the keys of an object are iterated. The default type sorts keys, and
[`ordered_json`](../api/ordered_json.md) keeps insertion order;
- change when iterators, pointers, or references are invalidated, or the state of a moved-from `basic_json`;
- add or remove implicit conversions from `basic_json`;
- change how `to_json` and `from_json` functions are found, or the behavior of
[`adl_serializer`](../api/adl_serializer/index.md);
- add pure virtual functions to the [`json_sax`](../api/json_sax/index.md) interface;
- remove, rename, renumber, or add enumerators of `value_t`;
- remove or rename a documented macro, CMake option, CMake target, or header, or change what a documented macro does.
Exceptions to these rules, for instance when fixing a bug requires changing the exception a function throws, are
documented in the [release notes](../home/releases.md).
@@ -51,13 +64,14 @@ The following are **not** part of the public API and may change in any release,
- The text of exception messages returned by `what()`. Use the [exception id](../home/exceptions.md) to tell errors
apart.
- The ABI, including `sizeof(basic_json)` and the memory layout of its values. Recompile your code when you upgrade the
library. The [versioned inline namespace](../features/namespace.md) turns mixing versions into a link error.
- The ABI, including `sizeof(basic_json)` and the memory layout of its values. The
[versioned inline namespace](../features/namespace.md) turns mixing versions into a link error.
- The hash values returned by `std::hash` for `basic_json`. Numbers that compare equal still hash equally.
- Everything in namespace `nlohmann::detail`, and macros and type traits that are not documented in the
[API reference](../api/basic_json/index.md).
Changes that would break the public API are only added behind a macro whose default keeps the 3.x behavior, see
[Version 4.0](#version-40).
Breaking changes are only added behind a macro whose default keeps the 3.x behavior. See [Version 4.0](#version-40) and
the [macro overview](../features/macros.md).
## Version 4.0
@@ -19,25 +19,9 @@ int main()
<< j.contains("/array/1"_json_pointer) << '\n'
<< j.contains("/array/-"_json_pointer) << '\n'
<< j.contains("/array/4"_json_pointer) << '\n'
<< j.contains("/baz"_json_pointer) << std::endl;
try
{
// try to use an array index with leading '0'
j.contains("/array/01"_json_pointer);
}
catch (const json::parse_error& e)
{
std::cout << e.what() << '\n';
}
try
{
// try to use an array index that is not a number
j.contains("/array/one"_json_pointer);
}
catch (const json::parse_error& e)
{
std::cout << e.what() << '\n';
}
<< j.contains("/baz"_json_pointer) << '\n'
// an array index with a leading '0' is not found
<< j.contains("/array/01"_json_pointer) << '\n'
// an array index that is not a number is not found
<< j.contains("/array/one"_json_pointer) << std::endl;
}
@@ -5,3 +5,5 @@ true
false
false
false
false
false
+1 -1
View File
@@ -8,7 +8,7 @@ int main()
try
{
// parsing input with a syntax error
json::parse("[1,2,3,]");
json j = json::parse("[1,2,3,]");
}
catch (const json::parse_error& e)
{
+5 -9
View File
@@ -473,19 +473,15 @@ class json_pointer
// convert null values to arrays or objects before continuing
if (ptr->is_null())
{
// check if the reference token is a valid array index, that is
// a nonempty sequence of digits without a leading '0'
// (cf. RFC 6901, Sect. 4); tokens that could never be a valid
// array index (such as "01" or "") are treated as object keys
const bool nums = !reference_token.empty()
&& (reference_token.size() == 1 || reference_token[0] != '0')
&& std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
// check if the reference token is a number
const bool nums =
std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
{
return std::isdigit(x);
});
// change value to an array for array indices or "-" or to object otherwise
// change value to an array for numbers or "-" or to object otherwise
*ptr = (nums || reference_token == "-")
? detail::value_t::array
: detail::value_t::object;
@@ -85,6 +85,12 @@ template<typename BasicJsonType, typename CharType, typename OutputSinkType = ou
class binary_writer
{
using string_t = typename BasicJsonType::string_t;
/// an object key as string_t: a reference when object_t::key_type already is
/// string_t, otherwise a converted copy that outlives sanitize_utf8_for_write's result
using object_key_string_t = typename std::conditional <
std::is_same<typename BasicJsonType::object_t::key_type, string_t>::value,
const string_t&, string_t >::type;
using binary_t = typename BasicJsonType::binary_t;
using number_float_t = typename BasicJsonType::number_float_t;
@@ -819,8 +825,10 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
// a converted key must outlive the reference returned by sanitize_utf8_for_write
const object_key_string_t key_string = el.first;
string_t storage;
const string_t& key = sanitize_utf8_for_write(el.first, j, storage);
const string_t& key = sanitize_utf8_for_write(key_string, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
@@ -1366,8 +1374,10 @@ class binary_writer
continue;
}
// a converted key must outlive the reference returned by sanitize_utf8_for_write
const object_key_string_t key_string = current.object_it->first;
string_t storage;
const string_t& key = sanitize_utf8_for_write(current.object_it->first, j, storage);
const string_t& key = sanitize_utf8_for_write(key_string, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
@@ -2730,6 +2740,11 @@ class binary_writer
itself in every case but a sanitized `replace`/`ignore` one, so @a
storage must outlive the returned reference only then.
@a s must be an lvalue that outlives the returned reference. An object key
whose `key_type` is not @ref string_t must therefore first be converted
into a named string_t (see @ref object_key_string_t); the deleted overload
below enforces this at compile time.
@param[in] s the string (value or object key) to write
@param[in] context the value @a s belongs to (for diagnostics)
@param[out] storage backing storage for a sanitized copy
@@ -2759,6 +2774,10 @@ class binary_writer
}
}
/// deleted: anything but a string_t would bind a temporary that dies before the returned reference is used
template < typename T, enable_if_t < !std::is_same<T, string_t>::value, int > = 0 >
const string_t& sanitize_utf8_for_write(const T& /*s*/, const BasicJsonType& /*context*/, string_t& /*storage*/) const = delete;
/*!
@brief write an integer in the shortest encoding
+26 -11
View File
@@ -20391,19 +20391,15 @@ class json_pointer
// convert null values to arrays or objects before continuing
if (ptr->is_null())
{
// check if the reference token is a valid array index, that is
// a nonempty sequence of digits without a leading '0'
// (cf. RFC 6901, Sect. 4); tokens that could never be a valid
// array index (such as "01" or "") are treated as object keys
const bool nums = !reference_token.empty()
&& (reference_token.size() == 1 || reference_token[0] != '0')
&& std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
// check if the reference token is a number
const bool nums =
std::all_of(reference_token.begin(), reference_token.end(),
[](const unsigned char x)
{
return std::isdigit(x);
});
// change value to an array for array indices or "-" or to object otherwise
// change value to an array for numbers or "-" or to object otherwise
*ptr = (nums || reference_token == "-")
? detail::value_t::array
: detail::value_t::object;
@@ -21594,6 +21590,12 @@ template<typename BasicJsonType, typename CharType, typename OutputSinkType = ou
class binary_writer
{
using string_t = typename BasicJsonType::string_t;
/// an object key as string_t: a reference when object_t::key_type already is
/// string_t, otherwise a converted copy that outlives sanitize_utf8_for_write's result
using object_key_string_t = typename std::conditional <
std::is_same<typename BasicJsonType::object_t::key_type, string_t>::value,
const string_t&, string_t >::type;
using binary_t = typename BasicJsonType::binary_t;
using number_float_t = typename BasicJsonType::number_float_t;
@@ -22328,8 +22330,10 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
// a converted key must outlive the reference returned by sanitize_utf8_for_write
const object_key_string_t key_string = el.first;
string_t storage;
const string_t& key = sanitize_utf8_for_write(el.first, j, storage);
const string_t& key = sanitize_utf8_for_write(key_string, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
@@ -22875,8 +22879,10 @@ class binary_writer
continue;
}
// a converted key must outlive the reference returned by sanitize_utf8_for_write
const object_key_string_t key_string = current.object_it->first;
string_t storage;
const string_t& key = sanitize_utf8_for_write(current.object_it->first, j, storage);
const string_t& key = sanitize_utf8_for_write(key_string, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
@@ -24239,6 +24245,11 @@ class binary_writer
itself in every case but a sanitized `replace`/`ignore` one, so @a
storage must outlive the returned reference only then.
@a s must be an lvalue that outlives the returned reference. An object key
whose `key_type` is not @ref string_t must therefore first be converted
into a named string_t (see @ref object_key_string_t); the deleted overload
below enforces this at compile time.
@param[in] s the string (value or object key) to write
@param[in] context the value @a s belongs to (for diagnostics)
@param[out] storage backing storage for a sanitized copy
@@ -24268,6 +24279,10 @@ class binary_writer
}
}
/// deleted: anything but a string_t would bind a temporary that dies before the returned reference is used
template < typename T, enable_if_t < !std::is_same<T, string_t>::value, int > = 0 >
const string_t& sanitize_utf8_for_write(const T& /*s*/, const BasicJsonType& /*context*/, string_t& /*storage*/) const = delete;
/*!
@brief write an integer in the shortest encoding
+2 -1
View File
@@ -138,7 +138,8 @@ json_test_set_test_options(test-disabled_exceptions
# only the #972 regression test needs thirdparty/fifo_map on its include path
json_test_set_test_options(test-regression1 LINK_LIBRARIES fifo_map_include)
# GCC's false -Warray-bounds error with JSON_DIAGNOSTICS only shows up when optimizing (#5742).
# Regression test for GCC's false -Warray-bounds error with JSON_DIAGNOSTICS (#5742, fixed in #5585). It only
# showed up when optimizing, so build this test with -O3 and the warning as an error.
# -O3 makes the optimizer-driven warnings of the ci_test_gcc flag set (-Winline,
# -Wsuggest-attribute=...) fire on the library's inline functions; they are not
# what this test checks, so turn them off for it.
@@ -12,7 +12,10 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <map>
#include <memory>
#include <string>
#include <utility>
#include <vector>
namespace
@@ -55,6 +58,49 @@ std::string dump_and_parse(const std::string& raw, eh error_handler)
return json::parse(json(raw).dump(-1, ' ', false, error_handler)).get<std::string>();
}
// an object key type that is not string_t, but converts implicitly to it
class converting_key
{
public:
converting_key(const char* s) : m_value(s) {} // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
converting_key(std::string s) : m_value(std::move(s)) {} // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
// the conversion yields a temporary string_t
operator std::string() const // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
{
return m_value;
}
// read by the exception messages when JSON_DIAGNOSTICS is enabled
const char* data() const noexcept
{
return m_value.data();
}
friend bool operator<(const converting_key& lhs, const converting_key& rhs)
{
return lhs.m_value < rhs.m_value;
}
private:
std::string m_value;
};
// ObjectType using converting_key; the Key template argument is ignored
template<typename Key, typename Value, typename Compare, typename Allocator>
class converting_key_object : public std::map<converting_key, Value, std::less<converting_key>, // NOLINT(modernize-use-transparent-functors)
typename std::allocator_traits<Allocator>::template rebind_alloc<std::pair<const converting_key, Value>>>
{
using base_type = std::map<converting_key, Value, std::less<converting_key>, // NOLINT(modernize-use-transparent-functors)
typename std::allocator_traits<Allocator>::template rebind_alloc<std::pair<const converting_key, Value>>>;
public:
using base_type::base_type;
using base_type::operator=;
};
using converting_key_json = nlohmann::basic_json<converting_key_object>;
} // namespace
TEST_CASE("UTF-8 error_handler for the binary readers and writers")
@@ -370,3 +416,109 @@ TEST_CASE("UTF-8 error_handler for the binary readers and writers")
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == ill_formed_cases()[0].bytes);
}
}
// The UBJSON and BJData writers bind the (possibly sanitized) key to a const
// string_t&. If key_type is not string_t but converts to it, the converted
// temporary must outlive that reference; this was a use-after-scope found by
// AddressSanitizer. Keys exceed the small string optimization on purpose.
TEST_CASE("UBJSON and BJData writers with an object_t whose key_type is not string_t")
{
const std::string long_prefix(70, 'k');
SECTION("well-formed keys, every error_handler")
{
const std::string key1 = long_prefix + "-first";
const std::string key2 = long_prefix + "-second";
converting_key_json::object_t o;
o.emplace(converting_key(key1), 1);
o.emplace(converting_key(key2), "value");
const converting_key_json v(std::move(o));
json expected;
expected[key1] = 1;
expected[key2] = "value";
const bool combos[3][2] = {{false, false}, {true, false}, {true, true}};
for (const auto h : all_handlers())
{
CAPTURE(static_cast<int>(h))
for (const auto& combo : combos)
{
const bool use_count = combo[0];
const bool use_type = combo[1];
CAPTURE(use_count)
CAPTURE(use_type)
CHECK(json::from_ubjson(converting_key_json::to_ubjson(v, use_count, use_type, h)) == expected);
CHECK(json::from_bjdata(converting_key_json::to_bjdata(v, use_count, use_type, json::bjdata_version_t::draft2, h)) == expected);
CHECK(json::from_bjdata(converting_key_json::to_bjdata(v, use_count, use_type, json::bjdata_version_t::draft3, h)) == expected);
}
}
}
SECTION("ill-formed keys")
{
for (const auto& c : ill_formed_cases())
{
CAPTURE(c.name)
const std::string key = long_prefix + c.bytes;
converting_key_json::object_t o;
o.emplace(converting_key(key), 1);
const converting_key_json v(std::move(o));
CHECK_THROWS_AS(converting_key_json::to_ubjson(v, false, false, eh::strict), converting_key_json::type_error&);
CHECK_THROWS_AS(converting_key_json::to_bjdata(v, false, false, json::bjdata_version_t::draft2, eh::strict), converting_key_json::type_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h))
const std::string expected = dump_and_parse(key, h);
CHECK(json::from_ubjson(converting_key_json::to_ubjson(v, false, false, h)).begin().key() == expected);
CHECK(json::from_bjdata(converting_key_json::to_bjdata(v, false, false, json::bjdata_version_t::draft2, h)).begin().key() == expected);
}
CHECK(json::from_ubjson(converting_key_json::to_ubjson(v, false, false, eh::keep)).begin().key() == key);
CHECK(json::from_bjdata(converting_key_json::to_bjdata(v, false, false, json::bjdata_version_t::draft2, eh::keep)).begin().key() == key);
}
}
SECTION("nested deeper than the recursion limit")
{
// wrap the previous value, innermost first
converting_key_json v = 42;
json expected = 42;
for (int i = 199; i >= 0; --i)
{
const std::string key = "level-" + std::to_string(i) + "-" + std::string(64, 'x');
converting_key_json::object_t o;
o.emplace(converting_key(key), std::move(v));
v = converting_key_json(std::move(o));
json e;
e[key] = std::move(expected);
expected = std::move(e);
}
for (const auto h : all_handlers())
{
CAPTURE(static_cast<int>(h))
for (const bool use_count :
{
false, true
})
{
CAPTURE(use_count)
CHECK(json::from_ubjson(converting_key_json::to_ubjson(v, use_count, false, h)) == expected);
CHECK(json::from_bjdata(converting_key_json::to_bjdata(v, use_count, false, json::bjdata_version_t::draft2, h)) == expected);
}
}
}
}
-66
View File
@@ -466,72 +466,6 @@ TEST_CASE("JSON pointers")
}
}
SECTION("creating intermediate levels")
{
SECTION("tokens that are valid array indices create arrays")
{
json j;
j["/0"_json_pointer] = 1;
CHECK(j == json({1}));
json j2;
j2["/2"_json_pointer] = 1;
CHECK(j2 == json({nullptr, nullptr, 1}));
json j3;
j3["/-"_json_pointer] = 1;
CHECK(j3 == json({1}));
json j4;
j4["/foo/0/0"_json_pointer] = 1;
CHECK(j4 == json({{"foo", {{1}}}}));
}
SECTION("tokens that are no valid array indices create objects")
{
json j;
j["/one"_json_pointer] = 1;
CHECK(j == json({{"one", 1}}));
// leading '0' can never be a valid array index (RFC 6901, Sect. 4)
json j2;
j2["/01"_json_pointer] = 1;
CHECK(j2 == json({{"01", 1}}));
// the empty token is a valid object key, but no valid array index
json j3;
j3["/"_json_pointer] = 1;
CHECK(j3 == json({{"", 1}}));
}
SECTION("creating a level yields the same result as reusing it (#5357)")
{
json j;
j["/a/b/01/d"_json_pointer] = "value";
json j_init = json::object();
j_init["/a/b"_json_pointer] = json::object();
j_init["/a/b/01/d"_json_pointer] = "value";
const json expected = json::parse(R"({"a":{"b":{"01":{"d":"value"}}}})");
CHECK(j == expected);
CHECK(j_init == expected);
// unflatten uses the same key
const json flat = {{"/a/b/01/d", "value"}};
CHECK(flat.unflatten() == expected);
}
SECTION("existing arrays still reject invalid indices")
{
json j = {1, 2, 3};
CHECK_THROWS_WITH_AS(j["/01"_json_pointer],
"[json.exception.parse_error.106] parse error: array index '01' must not begin with '0'", json::parse_error&);
CHECK_THROWS_WITH_AS(j.at("/01"_json_pointer),
"[json.exception.parse_error.106] parse error: array index '01' must not begin with '0'", json::parse_error&);
}
}
SECTION("flatten")
{
json j =