A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.
Fixes:
- wrong signatures and version histories (operator!= C++20 member,
binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed
Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
Mermaid diagrams, new API pages for json_pointer::operator<=> and
byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release
Checks:
- strict documentation build (broken links/anchors fail it); CI and
the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
header links, docset index coverage; its unused-example check works
again
- all examples produce the same output on every platform
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the key type when rejecting non-string CBOR/MessagePack map keys
CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:
syntax error while parsing MessagePack object key: only string keys
are supported, but found nil; last byte: 0xC0
The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.
Refs #2766, #3381
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point the MessagePack key note to the spec's profile section
The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Throw instead of writing MessagePack lengths beyond UINT32_MAX
MessagePack stores the length of a string, binary value, array, or
object in at most 32 bits. For a larger value, to_msgpack wrote no length
at all, so the output could not be read back. It now throws
out_of_range.412, which BSON already uses for its 32-bit length fields.
The check lives in one function, so each length is written by an
if/else chain that ends in a plain else, without a condition that can
never be false. It is tested with string and binary types that report a
size beyond UINT32_MAX without allocating it, like the BSON tests do.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the MessagePack length check
- mark to_msgpack_length's value as used when exceptions are disabled
(-Wunused-parameter, misc-unused-parameters)
- put "Exception safety" before "Exceptions" in to_msgpack.md, as the
documentation style check requires
- create the test's string value from its type: constructing it from a
beyond_uint32_string_t considers the std::filesystem::path conversion,
which libstdc++ 10 reports as ambiguous for a class derived from
std::string (clang 13)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the MessagePack string length test for clang with libstdc++ 10
C++17 builds consider the std::filesystem::path conversion for the
string type, and with clang and libstdc++ 10 that conversion is
ambiguous for a class derived from std::string. Creating the value from
its type did not avoid it, since any basic_json with that string type
instantiates the check. The binary and ext cases are still tested there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the MessagePack string test type and its alias in one block
astyle indented the alias oddly when it had an #ifdef of its own after
the binary alias; declare it right after the string type, in the same
block.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)
from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).
Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.
There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* ⚡ validate only newly read bytes of binary-format strings
get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Reject MessagePack/BSON binary subtypes that don't fit their wire format
Both formats store byte_container_with_subtype's subtype (a uint64_t)
in a single byte. The writers cast to std::int8_t/std::uint8_t without
a range check, so subtypes above 255 were silently truncated modulo
256 instead of raising an error. Throw out_of_range.413 instead when
the subtype exceeds the representable range of 0-255.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the new binary-subtype regression test out of unit-regression2.cpp
unit-regression2.cpp is already at the edge of what the MinGW linker
can relocate; adding this test's ~26 lines tips test-regression2_cpp20
(clang, Windows) over into "relocation truncated to fit:
IMAGE_REL_AMD64_REL32 against `.rdata'" (see 8ce64b9c1 / b82717c8a for
the same failure mode). Split the test along format lines instead:
MessagePack assertions move to unit-msgpack.cpp, BSON assertions to
unit-bson.cpp. The CBOR round-trip guard is dropped as redundant --
unit-cbor.cpp's "Tagged values" section already round-trips subtypes
up to 8589934590, far past the 70000 checked here.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Bound UBJSON optimized arrays of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized UBJSON array of one of those has no
payload: reading an element consumes no input at all. Its declared count is
therefore the only thing that decides how much is allocated, and nothing
bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing
two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and
OSS-Fuzz has an out-of-memory and a timeout report for the same shape.
Every other type costs at least one byte per element, so the end of the input
bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not
affected either: each element is preceded by its key, which costs bytes. And
BJData already refuses these markers as an optimized type, so this is a plain
UBJSON matter.
Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it
will not honour. The check runs before the SAX start event, so no container
is opened and then abandoned.
Rejecting on the read side alone would break the guarantee that anything
to_ubjson() writes can be read back, and would trip the round-trip assertion
in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized
encoding, one byte per element, for arrays of these types above the same
limit. Its decision depends only on the array's size, which is identical for
a value and for anything parsed back from it, so the round trip is stable.
No existing test changes: the largest such count in the test suite is 65,793.
The excessive-size test that already used this shape still passes, now
rejected a little earlier than by the max_size() check it used to reach.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Note the 1,048,576 valueless-array limit as (1 << 20) in the docs
Addresses review feedback from @gregmarr on PR #5504: spell out the
binary/hex form next to the decimal count so it reads as the round
power-of-two it is, matching how include/nlohmann/detail/input/binary_reader.hpp
defines max_valueless_container_size. Applied in both docs/exceptions.md
and ubjson.md, as requested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Reject JSON Patch move when from is a proper prefix of path
RFC 6902 (section 4.4) forbids "from" from being a proper prefix of
"path" for a "move" operation: "a location cannot be moved into one
of its children." "move" is implemented as remove-then-add with no
check for this. For object targets, the subsequent "add" happened to
throw as a side effect of resolving through the now-removed parent,
but for array targets, removing the "from" element shifts subsequent
indices, so "path" silently re-resolves to a different element and
the operation "succeeds" with a silently corrupted document.
Add a check, before performing the remove/add, for whether "from" is
a proper prefix of "path" at the reference-token level. This compares
json_pointer's already-unescaped reference_tokens vectors (basic_json
is a friend of json_pointer) rather than the raw pointer strings, so
that tokens containing escaped '/' or '~' characters are compared
correctly, and a token that merely looks like a string prefix (e.g.
"/ab" vs "/abc/x") is not mistaken for a pointer-token prefix. When
"from" is a proper prefix of "path", throw out_of_range.414.
Fixes#5397.
Stacked on top of the fix for #5396 (branch
issue-5396-patch-remove-primitive-parent), since both touch the same
patch_inplace move/remove handling in include/nlohmann/json.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add root-pointer and array-append-token edge case tests for the move prefix check
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Replace std::equal with an explicitly-bounded loop in the move prefix check
The three-iterator std::equal(first1, last1, first2) form has no
explicit end iterator for the second range, which a static analyzer
(Flawfinder, CWE-126) flags as a potential over-read even though the
preceding size comparison already guarantees the second range is long
enough. Rather than argue the point, make the bound visible in the code
itself via an explicit loop -- every access to ptr.reference_tokens is
now guarded by the same index the loop condition bounds against
from_size.
(The C++14 four-iterator std::equal(first1, last1, first2, last2) form
was tried first as a more minimal fix, but this codebase targets C++11
and that overload is not safely usable under -std=c++11 with all
supported standard library implementations.)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Extract the move prefix check into a named helper lambda
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Account for JSON_DIAGNOSTIC_POSITIONS in the move-prefix-check error messages
out_of_range::create() includes a "(bytes X-Y)" position annotation when
JSON_DIAGNOSTIC_POSITIONS is enabled, which the ci_test_diagnostic_positions
CI job builds the whole suite with. The five new out_of_range.414
assertions only checked the annotation-free message. Confirmed
JSON_DIAGNOSTICS produces the same (annotation-free) message as the
default build for this particular throw site (its path-based annotation
is empty at the root, where &result always points here), so only two
message variants are needed, not three.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
RFC 6902 (section 4.2) requires the target location of a "remove"
operation to exist. operation_remove handled parent.is_object() and
parent.is_array(), but had no final else branch: when the resolved
parent was a primitive value or null, neither branch matched and the
operation silently did nothing instead of failing.
Add the missing else branch, throwing out_of_range.413 with wording
that matches the existing out_of_range.411 thrown by the analogous
"add" case (operation_add) for the same kind of invalid parent.
Fixes#5396.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Throw other_error.502 when UBJSON use_type is set without use_size
Fixes#5321
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Scope UBJSON use_type check to container branches and expand tests
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Re-amalgamate single_include/json.hpp
The previous commit updated the split headers but the amalgamated
file didn't go back through astyle before I committed it, so CI's
amalgamation check caught formatting drift in json_fwd.hpp and a
few noexcept clauses in basic_json, plus one doc example. None of
it touches the UBJSON logic. Applied the patch CI generated to
bring single_include back in sync.
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
---------
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* 📝 Document exceptions newly thrown by the binary-format hardening
A round of binary-format input validation (#5274, #5284, #5287, #5332)
added new failure modes without updating exceptions.md, and left two
descriptions factually narrower than the code:
- parse_error.110 said "CBOR or MessagePack"; BSON and UBJSON also
throw it. Generalized, and added the BSON EOF example (#5332).
- parse_error.112: added the BSON document-size mismatch example
(#5287).
- parse_error.113 said "while parsing a map key", but its own existing
UBJSON char example already contradicted that. Broadened to cover
invalid length specifications, and added the negative-string-length
example (#5284).
- out_of_range.408 said "of an UBJSON array or object"; CBOR now throws
it too (#5274). Generalized and added both CBOR examples.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 📝 Correct which types may be referenced in a tuple extraction
The note added in #5271 said a referenced type must be one the library
stores "or an arithmetic type it can convert to/from". The parenthetical
is wrong: is_compatible_reference_type requires an exact match against
the stored types, so std::tuple<int&> is rejected by static_assert even
though int converts fine as a value. Only the value case is permissive.
Spell out the eight admissible types, give the int& counter-example, and
separate the reference restriction from by-value conversion.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 📝 Fix documentation gaps for 3.13.0 release (todos 138-142)
- Todo 138: Add "Known issues" section to modules.md with compiler-specific troubleshooting (GCC redefinition, MSVC symbol export). Add pointer note to quality_assurance.md.
- Todo 139: Document CBOR/MessagePack half-precision float encoding for NaN/Infinity (0xF9/0xCA with exact byte sequences). Explain pre-3.13.0 double-precision bug mechanism without issue citations.
- Todo 140: Document CBOR negative-integer-overflow rejection (parse_error.112) for magnitudes exceeding int64_t range (already implemented in rev 1).
- Todo 141: Update version history in value.md and operator[].md with behavior-change details, removing issue citations per citation policy (prose is self-contained).
- Todo 142: Global sed replace of 3.12.x → 3.13.0 placeholder across all 20 documentation files.
Revision 2 incorporates feedback to reduce changelog-like issue citations. Only citations that add unique troubleshooting value are retained (#5103 for GCC workaround, #3970 for MSVC symbol export). "Known issues" section follows PR #5252's visual pattern (info admonition with bold-bullet format).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 📝 Document integer type selection, type_name() invalid value, and std::optional get() fix
- number_handling.md: clarify that positive/negative integers select
unsigned/signed storage based on the leading minus sign (todo 143).
- type_name.md: document the new "invalid" return value for corrupted
JSON values (todo 145).
- get.md: note that get<std::optional<T>>() was unreachable in every
configuration prior to 3.13.0 due to an internal macro-guard bug,
unrelated to JSON_USE_IMPLICIT_CONVERSIONS's actual effect (todo 144).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Added NLOHNMANN_JSON_SERIALIZE_ENUM_STRICT
- duplicate of NLOHMANN_JSON_SERIALIZE_ENUM
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* Added failing tests for NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* modified NLOHMANN_JSON_SERIALIZE_STRICT to throw
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* added documentation and changed readme to include NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* ran amalgamate
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* docs(macros): add page for JSON_SERIALIZE_ENUM_STRICT
- added page to nav
- added links to new page where appropriate
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* refactor(macros): make JSON_SERIALIZE_ENUM_STRICT use JSON_THROW
- added templated wrapper function to fix scope error in calling JSON_THROW
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* refactor(macros): make NLOHMANN_SERIALIZE_ENUM_STRICT use error code 410
- added error code 410 to docs
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* tests(macros): add test for to_json with enum value not mentioned
in mapping for NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* Apply suggestions from code review
Co-authored-by: Niels Lohmann <niels.lohmann@gmail.com>
Signed-off-by: Caillin Nugent <nugentcaillin@gmail.com>
* fix(macro): prevent compilation error with -Werror and -Wunused-parameter
with NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
- casted exception to void to avoid warning
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* fix(docs): add link to NLOHMANN_SERIALIZE_ENUM_STRICT docs to exception page
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* docs(macros): add example of exception throwing for NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
* refactor(macros): add more in-depth error message to NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
- changed error message to follow style of nlohmann/json#4989
- made description of throw wrapper more general
- updated tests and example of exceptions
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
---------
Signed-off-by: Caillin Nugent <caillinn@student.unimelb.edu.au>
Signed-off-by: Caillin Nugent <nugentcaillin@gmail.com>
Co-authored-by: Niels Lohmann <niels.lohmann@gmail.com>
PR #4873 introduced a safety check in sax_parse functions to catch
nullptr passed as SAX parser object, which had been already annotated by
JSON_HEDLEY_NON_NULL macro.
Compilers (e.g. clang) which respected the non-null annotation tended to
eliminate the safety check completely in optimized builds, while
compilers which did not, compiled the safety check in. This led to
different behaviors accross different compilers/platforms and/or build
types (debug, release).
This commit reverts PR #4873 to remove this discrepancy. Passing null to
non-null annotated parameter is considered to be undefined behavior.
Fixes#5048
Signed-off-by: Richard Musil <risa2000x@gmail.com>
Co-authored-by: Richard Musil <risa2000x@gmail.com>