Review and extend the documentation, and check it in CI (#5638)

* Review and extend the documentation, and check it in CI

A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.

Fixes:
- wrong signatures and version histories (operator!= C++20 member,
  binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
  example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
  suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed

Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
  Mermaid diagrams, new API pages for json_pointer::operator<=> and
  byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release

Checks:
- strict documentation build (broken links/anchors fail it); CI and
  the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
  header links, docset index coverage; its unused-example check works
  again
- all examples produce the same output on every platform

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep the customer links that could not be fixed

A dead link on the customers page is still the evidence of where the
use of the library was documented. Keep the original URLs of the entries
without a working replacement (Marne, Cisco Webex Desk Camera, Philips
Hue, CyberArk) and exclude exactly these URLs from the link check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the duplicate-key recipe's claim about SAX positions

The SAX interface's key() receives no position either; only parse_error()
does. Also note that the recipe does not report the path to the repeated
key (see discussion #5085).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Say the library is available as a single header and mention json_fwd.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct documentation errors found while hunting for bugs

- patch/patch_inplace: list the JSON pointer errors parse_error.106-109
  and out_of_range.402/404, and quote the actual parse_error.105 message.
- unflatten: list parse_error.106/107/108 and out_of_range.404.
- to_bson: list out_of_range.415 (binary subtype above 255) and note
  that 412 and 415 are new in 3.13.0.
- to_string: state that string_t must be convertible to std::string, also
  in the StringType requirements table.
- JSON Lines: a `while (input >> j)` loop also throws after the last value
  for concatenated JSON values; show a loop that works for both.
- BON8: a string gets 0xFF only if nothing follows it in the message; a
  string at the end of an array or object is ended by 0xFE.
- custom_string_type.hpp: add operator+=(char), which the "Always
  required" list asks for (json_pointer::to_string, flatten, unflatten,
  and diff did not compile), and an ADL int_to_string for diff and items.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cache the release headers with functools.lru_cache

Codacy (Pylint) flagged the mutable default argument that header() used
as its cache. functools.lru_cache keeps the same memoization without it.
The script's output is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-10-02 11:32:15 +02:00
committed by GitHub
parent 48ff79647f
commit 63c10a51fc
276 changed files with 5318 additions and 624 deletions
@@ -3,6 +3,16 @@
This library can create a JSON value from a wide range of inputs. This page gives an overview of the available parsing
functions and how they behave; the linked pages go into more detail.
```mermaid
flowchart LR
I["JSON input"] --> P["parse()"]
I --> S["sax_parse()"]
I --> A["accept()"]
P -->|"optional parser callback filters values"| D["basic_json value (DOM)"]
S --> H["events delivered to a user SAX handler"]
A --> V["bool: is the input valid JSON?"]
```
## Input
The [`parse`](../../api/basic_json/parse.md) function reads a JSON value from an input. The input can be
@@ -76,3 +86,4 @@ options.
- [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [parsing and exceptions](parse_exceptions.md) - control error handling
- [parsing untrusted input](untrusted_input.md) - what to consider when parsing input from untrusted sources
@@ -46,8 +46,20 @@ JSON Lines input with more than one value is treated as invalid JSON by the [`pa
}
```
with a JSON Lines input does not work, because the parser will try to parse one value after the last one.
with a JSON Lines input does not work, because the parser will try to parse one value after the last one and throw
a [`parse_error.101`](../../home/exceptions.md#jsonexceptionparse_error101) exception. The same happens for a
stream of *concatenated* (non-newline-delimited) JSON values: `operator>>` reads them one at a time, but the loop
above throws after the last value. To read either format with `operator>>`, check for the end of the stream before
each read:
This is different from parsing a stream of *concatenated* (non-newline-delimited) JSON values, for which
`operator>>` does work, provided that a value that is a number is followed by whitespace -- see its
[notes](../../api/operator_gtgt.md#notes) for details.
```cpp
json j;
while (input >> std::ws && input.peek() != std::char_traits<char>::eof())
{
input >> j;
std::cout << j << std::endl;
}
```
A value that is a number must be followed by whitespace -- see the [notes](../../api/operator_gtgt.md#notes) of
`operator>>` for details.
@@ -23,9 +23,9 @@ In case exceptions are undesired or not supported by the environment, there are
## Switch off exceptions
The `parse()` function accepts a `#!cpp bool` parameter `allow_exceptions` which controls whether an exception is
thrown when a parse error occurs (`#!cpp true`, default) or whether a discarded value should be returned
(`#!cpp false`).
The [`parse()`](../../api/basic_json/parse.md) function accepts a `#!cpp bool` parameter `allow_exceptions` which
controls whether an exception is thrown when a parse error occurs (`#!cpp true`, default) or whether a discarded value
should be returned (`#!cpp false`).
```cpp
json j = json::parse(my_input, nullptr, false);
@@ -39,8 +39,8 @@ Note there is no diagnostic information available in this scenario.
## Use accept() function
Alternatively, function `accept()` can be used which does not return a `json` value, but a `#!cpp bool` indicating
whether the input is valid JSON.
Alternatively, function [`accept()`](../../api/basic_json/accept.md) can be used which does not return a `json` value,
but a `#!cpp bool` indicating whether the input is valid JSON.
```cpp
if (!json::accept(my_input))
@@ -66,56 +66,18 @@ bool parse_error(std::size_t position,
The return value indicates whether the parsing should continue, so the function should usually return `#!cpp false`.
??? example
??? example "Example: report parse errors without exceptions"
The example derives from the library's DOM parser and overrides `parse_error` to print the error instead of
throwing. Note the DOM parser is an implementation detail (`nlohmann::detail`) and may change between releases;
see [Do not use the `detail` namespace](../../integration/migration_guide.md#do-not-use-the-detail-namespace).
```cpp
#include <iostream>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
class sax_no_exception : public nlohmann::detail::json_sax_dom_parser<json>
{
public:
sax_no_exception(json& j)
: nlohmann::detail::json_sax_dom_parser<json>(j, false)
{}
bool parse_error(std::size_t position,
const std::string& last_token,
const json::exception& ex)
{
std::cerr << "parse error at input byte " << position << "\n"
<< ex.what() << "\n"
<< "last read: \"" << last_token << "\""
<< std::endl;
return false;
}
};
int main()
{
std::string myinput = "[1,2,3,]";
json result;
sax_no_exception sax(result);
bool parse_result = json::sax_parse(myinput, &sax);
if (!parse_result)
{
std::cerr << "parsing unsuccessful!" << std::endl;
}
std::cout << "parsed value: " << result << std::endl;
}
--8<-- "examples/sax_no_exception.cpp"
```
Output:
```
parse error at input byte 8
[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal
last read: "3,]"
parsing unsuccessful!
parsed value: [1,2,3]
--8<-- "examples/sax_no_exception.output"
```
@@ -2,8 +2,9 @@
## Overview
With a parser callback function, the result of parsing a JSON text can be influenced. When passed to `parse`, it is
called on certain events (passed as `parse_event_t` via parameter `event`) with a set recursion depth `depth` and
With a parser callback function, the result of parsing a JSON text can be influenced. When passed to
[`parse`](../../api/basic_json/parse.md), it is called on certain events (passed as
[`parse_event_t`](../../api/basic_json/parse_event_t.md) via parameter `event`) with a set recursion depth `depth` and
context JSON value `parsed`. The return value of the callback function is a boolean indicating whether the element that
emitted the callback shall be kept or not.
@@ -30,7 +31,7 @@ table describes the values of the parameters `depth`, `event`, and `parsed`.
| `parse_event_t::array_end` | the parser read `]` and finished processing a JSON array | depth of the parent of the JSON array | the parsed JSON array |
| `parse_event_t::value` | the parser finished reading a JSON value | depth of the value | the parsed JSON value |
??? example
??? example "Example: sequence of callback events"
When parsing the following JSON text,
@@ -76,7 +77,7 @@ was called:
- In case a value outside a structured type is skipped, it is replaced with `#!json null`. This case happens if the
top-level element is skipped.
??? example
??? example "Example: skip an object key while parsing"
The example below demonstrates the `parse()` function with and without callback function.
@@ -98,7 +99,7 @@ the resulting `#!c json` value -- once parsing has produced that value, the dupl
storage maps each key to a single value. If duplicate keys should instead be treated as an error, a parser callback
can detect them while the object is still being read, before that ambiguity ever applies.
??? example
??? example "Example: reject duplicate object keys"
```cpp
--8<-- "examples/reject_duplicate_keys.cpp"
@@ -110,16 +111,18 @@ can detect them while the object is still being read, before that ambiguity ever
--8<-- "examples/reject_duplicate_keys.output"
```
This approach has two limitations:
This approach has three limitations:
- The depth-indexed bookkeeping must account for the fact that `object_start` reports the depth of the *parent* of
the object, while the `key` events inside that object are reported one depth deeper (see the event table above);
it is easy to get this off by one for nested objects.
- The thrown exception cannot carry a `parse_error`-style byte offset, because position tracking only exists inside
the parser and lexer, not at the callback layer.
- The exception only names the repeated key, not where it occurs in the document. Reporting its full path requires
maintaining a stack of the enclosing keys and array indices in the callback as well.
For strict validation with precise error positions, implementing a [SAX interface](sax_interface.md) instead gives
access to the parser's position information directly.
A [SAX interface](sax_interface.md) does not lift the position limitation: its `key` function receives no position
either -- only `parse_error` is passed the byte position.
## Recipe: streaming a large homogeneous array
@@ -129,7 +132,7 @@ discard it, so memory usage stays bounded by a single element (plus the not-yet-
than the whole document. Since the top-level array's `array_start`/`array_end` are reported at `depth == 0` (its
parent is the document root), the object elements it contains are reported at `depth == 1`:
??? example
??? example "Example: stream a large top-level array"
```cpp
std::ifstream input("large_array.json");
@@ -154,7 +157,7 @@ homogeneous values by checking `object_end`/`value` events at `depth == 1` there
Since there is no built-in nesting-depth limit (see the note above), a callback can enforce one manually by
tracking the maximum `depth` seen and throwing once it is exceeded:
??? example
??? example "Example: limit the nesting depth"
```cpp
constexpr int max_depth = 32;
@@ -0,0 +1,163 @@
# Parsing Untrusted Input
This page is for applications that parse JSON -- or one of the supported [binary formats](../binary_formats/index.md)
(BJData, BON8, BSON, CBOR, MessagePack, UBJSON) -- from a source they do not fully control, such as a network
connection, an uploaded file, or another process. It summarizes what the library already does for such input and what
remains the caller's responsibility, linking to the pages that cover each aspect in detail rather than repeating them.
For the project's threat model and the countermeasures behind these behaviors, see the
[assurance case](../../community/assurance_case.md); to report a vulnerability, see the
[security policy](../../community/security_policy.md).
## Errors without exceptions
By default, [`parse()`](../../api/basic_json/parse.md) throws a
[`parse_error`](../../home/exceptions.md#jsonexceptionparse_error101) (for instance `parse_error.101` for a syntax
error) when the input is not valid. If your environment cannot use exceptions for untrusted input, the library offers
several alternatives; see [Parsing and exceptions](parse_exceptions.md) for the full comparison:
- Pass `#!cpp false` as the third argument to `parse()` to get a discarded value
(checked with [`is_discarded()`](../../api/basic_json/is_discarded.md)) instead of a thrown exception, with no
diagnostic information.
- Use [`accept()`](../../api/basic_json/accept.md) to only check whether the input is valid JSON, without building a
value.
- Implement the [SAX interface](sax_interface.md) and override `parse_error()` to react to an error yourself, with the
byte position and the exception that would otherwise have been thrown; see the
[example](parse_exceptions.md#user-defined-sax-interface) that overrides it to print instead of throw.
If exceptions are unavailable entirely (`-fno-exceptions`, or [`JSON_NOEXCEPTION`](../../api/macros/json_noexception.md)
defined), every `#!cpp throw` in the library becomes a call to `std::abort()` -- there is no way to recover from a
parse error of untrusted input in that configuration; see
[Switch off exceptions](../../home/exceptions.md#switch-off-exceptions) for the details and for overriding this with
`JSON_THROW_USER`.
## Nesting depth
The JSON parser and the binary readers are iterative: they keep the containers they are currently inside of on a
heap-allocated stack instead of calling themselves once per nesting level, so the native call stack does not grow with
the nesting depth of the input. A deeply nested document is therefore bounded by available memory, not by the call
stack, however deeply it is nested.
!!! warning "No built-in depth limit while parsing"
Neither the parser nor the binary readers impose a limit on how deep the input may nest. An attacker can still
exhaust memory (though not the call stack) with a sufficiently deep document. If you need to reject over-deep
untrusted input outright, track the depth yourself, either with a
[parser callback](parser_callbacks.md#recipe-max-nesting-depth-via-a-callback) for the JSON parser, or by counting
`start_object`/`start_array` and `end_object`/`end_array` calls in a
[SAX handler](sax_interface.md) (for the JSON parser or a binary format alike) and throwing once your limit is
exceeded.
Once a value has been parsed, operations that walk it recursively -- serializing it with
[`dump`](../../api/basic_json/dump.md), hashing it, copying it, comparing two values with `#!cpp ==`, `#!cpp <`, or (in
C++20) `#!cpp <=>`, merging with [`update`](../../api/basic_json/update.md), and applying a
[`merge_patch`](../../api/basic_json/merge_patch.md) -- descend at most 128 levels on the call stack and continue
below that with an explicit stack instead, so none of them can exhaust the stack either, however deeply the value is
nested. Destroying a value (its destructor) never recurses at all, regardless of nesting depth, for the same reason.
!!! note "Not every operation is bounded yet"
[`diff`](../../api/basic_json/diff.md), [`flatten`](../../api/basic_json/flatten.md), and the binary writers
(`to_cbor`, `to_msgpack`, ...) still recurse once per nesting level; this is called out as work in progress in the
[assurance case](../../community/assurance_case.md#secure-design). A value deep enough to matter for these
operations would typically first have to survive parsing without hitting a self-imposed depth limit, as described
above.
## Input size
The library does not limit the overall size of a JSON text; a value nested or wide enough will use memory
proportional to the input. If you parse untrusted input of unbounded size, check the size of the file or stream
yourself before -- or while -- handing it to `parse()`.
For the binary formats, an announced size is never trusted outright:
- Reading a string or binary value copies the input in bounded 4096-byte chunks and grows the result as bytes are
actually consumed, rather than allocating the announced length up front -- a truncated input runs out of bytes
(reported as a parse error) instead of triggering an oversized allocation.
- When an array announces its number of elements and the array container supports `reserve()` (as `#!cpp std::vector`,
the default, does), the library reserves storage for at most 16384 of them upfront, regardless of how large the
announced count is; further elements still grow the container normally as they are read.
- An announced array or object size that exceeds what the target container could ever hold (its `max_size()`) is
rejected immediately as [`out_of_range.408`](../../home/exceptions.md#jsonexceptionout_of_range408), without
attempting to allocate anything.
## Strings
Invalid UTF-8 is rejected while parsing, not just while serializing:
- In JSON text, an ill-formed UTF-8 byte in a string is a
[`parse_error.101`](../../home/exceptions.md#jsonexceptionparse_error101) ("invalid string: ill-formed UTF-8 byte").
- In a binary format, a string that is not valid UTF-8 is a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113).
A `#!cpp '\0'` (NUL) byte *inside* a quoted JSON string is always rejected (it must be escaped as `\u0000`). A NUL byte
*outside* of a string is different: by default it is silently treated as the end of the input, so trailing bytes after
it -- including further, otherwise well-formed JSON -- are silently ignored rather than rejected. Since untrusted input
that happens to embed a NUL is a way to make part of it disappear without a parse error, see the
[FAQ entry](../../home/faq.md#nul-bytes-in-the-input) and consider defining
[`JSON_STRICT_NUL_HANDLING`](../../api/macros/json_strict_nul_handling.md) to `1` to reject a NUL byte like any other
unexpected byte instead.
Parsing is not the only place invalid UTF-8 matters: a string that reached a `#!cpp json` value some other way (for
example, constructed by application code, or, before JSON_STRICT_NUL_HANDLING existed, read from a binary format that
does not validate strings) still has to round-trip back to JSON text. By default,
[`dump()`](../../api/basic_json/dump.md) throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
if the string is not valid UTF-8; passing
[`error_handler_t::replace`](../../api/basic_json/error_handler_t.md) or `error_handler_t::ignore` avoids the exception
instead of crashing an application that forgot to catch it. See
[Handling invalid UTF-8](../serialization.md#handling-invalid-utf-8) for the options and an example.
## Duplicate object keys
The JSON specification leaves the handling of repeated keys in an object up to the implementation, and this library
does too: as described in [`object_t`](../../api/basic_json/object_t.md#behavior), it is unspecified which of the
values for a repeated key ends up in the parsed object. If your application must reject duplicate keys instead of
silently resolving them one way or another, see the
[parser callback recipe for rejecting duplicate keys](parser_callbacks.md#recipe-rejecting-duplicate-object-keys).
## Numbers
A number whose value cannot be represented -- for instance `1E1000`, which overflows `double` -- is rejected while
parsing as [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) rather than silently becoming
infinity. An integer that is syntactically valid but does not fit the 64-bit integer types is not rejected; it is
instead stored as a `double`, which may lose precision for very large values. See
[number limits](../types/number_handling.md#number-limits) for the exact ranges and an example.
## Comments and trailing commas
Both [comments](../comments.md) and [trailing commas](../trailing_commas.md) are rejected by default, matching the
JSON specification; they must be explicitly enabled per call with the `ignore_comments` and `ignore_trailing_commas`
parameters of [`parse()`](../../api/basic_json/parse.md) or [`accept()`](../../api/basic_json/accept.md). Do not
enable either for input whose conformance you cannot otherwise control, since interoperability with strictly
conforming JSON consumers is exactly what the default rejects.
## Checklist
- Wrap parsing in a `#!cpp try`/`#!cpp catch` block, or use `allow_exceptions=false`/`accept()` if your environment
cannot use exceptions; never let `JSON_NOEXCEPTION`'s `abort()` be the first time you think about error handling.
- If the input's nesting depth matters to you, enforce your own limit with a
[parser callback](parser_callbacks.md#recipe-max-nesting-depth-via-a-callback) or a
[SAX handler](sax_interface.md); the library bounds the call stack but not memory use.
- Bound the size of the input itself before parsing, independent of the library's own bounded allocations for binary
format lengths.
- Decide up front how you want a string with invalid UTF-8 -- from any source, not only parsing -- to be serialized
(`strict`, `replace`, or `ignore`), rather than discovering it from an uncaught `type_error.316`.
- If a stray NUL byte silently truncating trailing input is a problem for your input format, define
`JSON_STRICT_NUL_HANDLING`.
- Decide whether duplicate object keys should be an error for your application, and add a callback if so.
- Do not enable `ignore_comments` or `ignore_trailing_commas` for input that must be strictly conforming JSON.
For the broader design rationale and how it is tested (fuzzing, sanitizers, static analysis), see the
[assurance case](../../community/assurance_case.md) and [quality assurance](../../community/quality_assurance.md). To
report a security issue in the library itself, follow the [security policy](../../community/security_policy.md).
## See also
- [Parsing](index.md) - overview of the parsing functions
- [Parsing and exceptions](parse_exceptions.md) - error handling without exceptions
- [Parser callbacks](parser_callbacks.md) - depth limits, duplicate-key rejection, and streaming recipes
- [SAX interface](sax_interface.md) - implement a custom handler with access to parse errors and positions
- [Serialization](../serialization.md) - handling invalid UTF-8 when dumping
- [Number handling](../types/number_handling.md) - number ranges and overflow behavior
- [Assurance case](../../community/assurance_case.md) - the library's threat model and countermeasures
- [Security policy](../../community/security_policy.md) - how to report a vulnerability