Review and extend the documentation, and check it in CI (#5638)

* Review and extend the documentation, and check it in CI

A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.

Fixes:
- wrong signatures and version histories (operator!= C++20 member,
  binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
  example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
  suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed

Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
  Mermaid diagrams, new API pages for json_pointer::operator<=> and
  byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release

Checks:
- strict documentation build (broken links/anchors fail it); CI and
  the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
  header links, docset index coverage; its unused-example check works
  again
- all examples produce the same output on every platform

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep the customer links that could not be fixed

A dead link on the customers page is still the evidence of where the
use of the library was documented. Keep the original URLs of the entries
without a working replacement (Marne, Cisco Webex Desk Camera, Philips
Hue, CyberArk) and exclude exactly these URLs from the link check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the duplicate-key recipe's claim about SAX positions

The SAX interface's key() receives no position either; only parse_error()
does. Also note that the recipe does not report the path to the repeated
key (see discussion #5085).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Say the library is available as a single header and mention json_fwd.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct documentation errors found while hunting for bugs

- patch/patch_inplace: list the JSON pointer errors parse_error.106-109
  and out_of_range.402/404, and quote the actual parse_error.105 message.
- unflatten: list parse_error.106/107/108 and out_of_range.404.
- to_bson: list out_of_range.415 (binary subtype above 255) and note
  that 412 and 415 are new in 3.13.0.
- to_string: state that string_t must be convertible to std::string, also
  in the StringType requirements table.
- JSON Lines: a `while (input >> j)` loop also throws after the last value
  for concatenated JSON values; show a loop that works for both.
- BON8: a string gets 0xFF only if nothing follows it in the message; a
  string at the end of an array or object is ended by 0xFE.
- custom_string_type.hpp: add operator+=(char), which the "Always
  required" list asks for (json_pointer::to_string, flatten, unflatten,
  and diff did not compile), and an ADL int_to_string for diff and items.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cache the release headers with functools.lru_cache

Codacy (Pylint) flagged the mutable default argument that header() used
as its cache. functools.lru_cache keeps the same memoization without it.
The script's output is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-10-02 11:32:15 +02:00
committed by GitHub
parent 48ff79647f
commit 63c10a51fc
276 changed files with 5318 additions and 624 deletions
+86 -14
View File
@@ -80,6 +80,29 @@ Some important things:
* In function `from_json`, use function [`at()`](../api/basic_json/at.md) to access the object values rather than `operator[]`. In case a key does not exist, `at` throws an exception that you can handle, whereas `operator[]` exhibits undefined behavior.
* You do not need to add serializers or deserializers for STL types like `std::vector`: the library already implements these.
??? example "Example: serialize a `person` to JSON with `to_json`"
```cpp
--8<-- "examples/to_json.cpp"
```
Output:
```json
--8<-- "examples/to_json.output"
```
??? example "Example: deserialize a `person` from JSON with `from_json`"
```cpp
--8<-- "examples/from_json__default_constructible.cpp"
```
Output:
```
--8<-- "examples/from_json__default_constructible.output"
```
## Simplify your life with macros
@@ -98,7 +121,29 @@ There are several macros to make your life easier as long as you want to use a J
For all the macros, the first parameter is the name of the class/struct. The `DERIVED_TYPE` macros require a second parameter of a base class. All the remaining parameters name the member variables. The `WITH_NAMES` macros require a JSON name before each of the variables.
| Need access to private members | Need only de-serialization | Allow missing values when de-serializing | macro |
```mermaid
flowchart TD
A["choosing a NLOHMANN_DEFINE_* macro"] --> B{"adding fields to a base class?"}
B -->|"yes"| C["...DERIVED_TYPE..."]
B -->|"no"| D["...TYPE..."]
C --> E{"need access to private members?"}
D --> E
E -->|"yes"| F["...INTRUSIVE... (used inside the class)"]
E -->|"no"| G["...NON_INTRUSIVE... (used in the namespace)"]
F --> H{"only serializing, never parsing back?"}
G --> H
H -->|"yes"| I["...ONLY_SERIALIZE"]
H -->|"no"| J{"allow missing keys when parsing?"}
J -->|"yes"| K["...WITH_DEFAULT"]
J -->|"no"| L["plain (missing keys throw)"]
I --> M{"need custom JSON key names?"}
K --> M
L --> M
M -->|"yes"| N["...WITH_NAMES"]
M -->|"no"| O["done"]
```
| Need access to private members | Need only serialization | Allow missing values when de-serializing | macro |
|------------------------------------------------------------------|------------------------------------------------------------------|------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------|
| <div style="color: green;">:octicons-check-circle-fill-24:</div> | <div style="color: red;">:octicons-x-circle-fill-24:</div> | <div style="color: red;">:octicons-x-circle-fill-24:</div> | [**NLOHMANN_DEFINE_TYPE_INTRUSIVE**](../api/macros/nlohmann_define_type_intrusive.md) |
| <div style="color: green;">:octicons-check-circle-fill-24:</div> | <div style="color: red;">:octicons-x-circle-fill-24:</div> | <div style="color: green;">:octicons-check-circle-fill-24:</div> | [**NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT**](../api/macros/nlohmann_define_type_intrusive.md) |
@@ -109,7 +154,7 @@ For all the macros, the first parameter is the name of the class/struct. The `DE
For _derived_ classes and structs, use the following macros
| Need access to private members | Need only de-serialization | Allow missing values when de-serializing | macro |
| Need access to private members | Need only serialization | Allow missing values when de-serializing | macro |
|------------------------------------------------------------------|------------------------------------------------------------------|------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------|
| <div style="color: green;">:octicons-check-circle-fill-24:</div> | <div style="color: red;">:octicons-x-circle-fill-24:</div> | <div style="color: red;">:octicons-x-circle-fill-24:</div> | [**NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE**](../api/macros/nlohmann_define_derived_type.md) |
| <div style="color: green;">:octicons-check-circle-fill-24:</div> | <div style="color: red;">:octicons-x-circle-fill-24:</div> | <div style="color: green;">:octicons-check-circle-fill-24:</div> | [**NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT**](../api/macros/nlohmann_define_derived_type.md) |
@@ -124,7 +169,7 @@ For _derived_ classes and structs, use the following macros
types with more than 63 member variables, you need to define the `to_json`/`from_json` functions manually.
- For the `WITH_NAMES` variants the limit is halved to 31 member variables.
??? example
??? example "Example: using the `NLOHMANN_DEFINE_TYPE_*` macros"
The `to_json`/`from_json` functions for the `person` struct above can be created with:
@@ -245,6 +290,14 @@ For _derived_ classes and structs, use the following macros
This requires a bit more advanced technique. But first, let us see how this conversion mechanism works:
```mermaid
flowchart LR
A["construct json j = t, or call j.get() for T"] --> B["JSONSerializer for T: to_json / from_json"]
B -->|"default JSONSerializer"| C["adl_serializer for T: to_json / from_json"]
C -->|"unqualified call, found via ADL"| D["free to_json(j, t) / from_json(j, t) in T's namespace"]
B -->|"user specialization replaces the default"| E["user's adl_serializer specialization for T"]
```
The library uses **JSON Serializers** to convert types to JSON.
The default serializer for `nlohmann::json` is `nlohmann::adl_serializer` (ADL means [Argument-Dependent Lookup](https://en.cppreference.com/w/cpp/language/adl)).
@@ -300,7 +353,24 @@ NLOHMANN_JSON_NAMESPACE_END
## How can I use `get()` for non-default constructible/non-copyable types?
There is a way if your type is [MoveConstructible](https://en.cppreference.com/w/cpp/named_req/MoveConstructible). You will need to specialize the `adl_serializer` as well, but with a special `from_json` overload:
For a type that is not [DefaultConstructible](https://en.cppreference.com/w/cpp/named_req/DefaultConstructible) but is
otherwise an ordinary value type, specialize `adl_serializer` with a `from_json` overload that returns the value instead
of writing into a reference:
??? example "Example: `get()` for a non-default-constructible type"
```cpp
--8<-- "examples/from_json__non_default_constructible.cpp"
```
Output:
```
--8<-- "examples/from_json__non_default_constructible.output"
```
The same technique also works if your type is not copyable, as long as it is
[MoveConstructible](https://en.cppreference.com/w/cpp/named_req/MoveConstructible):
```cpp
struct move_only_type {
@@ -359,15 +429,10 @@ json any_to_json(const std::any& a) {
## Why does serializing a `std::map`/`std::unordered_map` with non-string keys produce an array?
A `std::map`/`std::unordered_map` whose key type is not string-like (e.g., `std::map<int, std::string>`) is
serialized as a JSON *array* of 2-element `[key, value]` arrays, not as a JSON object -- JSON object keys must be
strings, so the library cannot represent an integer-keyed map as an object.
```cpp
std::map<int, std::string> m{{1, "one"}, {2, "two"}};
json j = m;
// j is [[1,"one"],[2,"two"]], not {"1":"one","2":"two"}
```
A `std::map`/`std::unordered_map` whose key type is not string-like (e.g., `std::map<int, std::string>`) cannot be
serialized as a JSON object, because JSON object keys must be strings. See
[Converting maps with non-string keys](types/index.md#converting-maps-with-non-string-keys) in the types article for
what the library does instead.
## Why does `std::wstring` convert or dump incorrectly?
@@ -411,7 +476,7 @@ struct less_than_32_serializer {
Be **very** careful when reimplementing your serializer, you can stack overflow if you don't pay attention:
```cpp
template <typename T, void>
template <typename T, typename = void>
struct bad_serializer
{
template <typename BasicJsonType>
@@ -429,3 +494,10 @@ struct bad_serializer
}
};
```
## See also
- [Converting values](conversions.md) - the general overview of `get`/`get_to` and implicit conversions
- [Specializing enum conversion](enum_conversion.md) - map enums to JSON strings instead of integers
- [Supported macros](macros.md) - reference for `NLOHMANN_DEFINE_TYPE_*` and related macros
- [`adl_serializer`](../api/adl_serializer/index.md) - the default `JSONSerializer` used in the conversion dispatch
+4 -4
View File
@@ -27,7 +27,7 @@ If you are not sure whether an element in an object exists, use checked access w
See also the documentation on [element access](element_access/index.md).
??? example "Example 1: Missing object key"
??? example "Example: missing object key"
The following code will trigger an assertion at runtime:
@@ -54,7 +54,7 @@ See also the documentation on [element access](element_access/index.md).
Constructing a JSON value from an iterator range (see [constructor](../api/basic_json/basic_json.md)) with an
uninitialized iterator is undefined behavior and yields a runtime assertion.
??? example "Example 2: Uninitialized iterator range"
??? example "Example: uninitialized iterator range"
The following code will trigger an assertion at runtime:
@@ -81,7 +81,7 @@ uninitialized iterator is undefined behavior and yields a runtime assertion.
Any operation on uninitialized iterators (i.e., iterators that are not associated with any JSON value) is undefined
behavior and yields a runtime assertion.
??? example "Example 3: Uninitialized iterator"
??? example "Example: uninitialized iterator"
The following code will trigger an assertion at runtime:
@@ -112,7 +112,7 @@ library asserted that the pointer was not `nullptr` using a runtime assertion. I
result in undefined behavior. Since version 3.12.0, this library checks for `nullptr` and throws a
[`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) to prevent the undefined behavior.
??? example "Example 4: Reading from null pointer"
??? example "Example: reading from null pointer"
The following code will trigger an assertion at runtime:
@@ -73,7 +73,7 @@ The library uses the following mapping from JSON values types to BJData types ac
!!! info "NaN/infinity handling"
If NaN or Infinity are stored inside a JSON number, they are serialized properly. This behavior differs from the
`dump()` function which serializes NaN or Infinity to `#!json null`.
[`dump()`](../../api/basic_json/dump.md) function which serializes NaN or Infinity to `#!json null`.
!!! info "Endianness"
@@ -163,7 +163,7 @@ The library uses the following mapping from JSON values types to BJData types ac
[BJDataBinArr]: https://github.com/NeuroJSON/bjdata/blob/master/Binary_JData_Specification.md#optimized-binary-array
??? example
??? example "Example: serialize JSON values to BJData, with and without size/type optimization"
```cpp
--8<-- "examples/to_bjdata.cpp"
@@ -218,7 +218,7 @@ The library maps BJData types to JSON value types as follows:
binary values above), and serializing such an array again may choose different, but equally valid, type markers.
The bytes can then differ, but parsing them again yields the same value.
??? example
??? example "Example: deserialize a JSON value from BJData"
```cpp
--8<-- "examples/from_bjdata.cpp"
@@ -53,8 +53,9 @@ The library uses the following mapping from JSON values types to BON8 types acco
An integer that takes 2 to 4 bytes starts with a UTF-8 lead byte (0xC2..0xF7) that is followed by a byte that cannot
continue a UTF-8 character: 0x00..0x7F for positive and 0xC0..0xFF for negative integers. A string is terminated by
0xFF only if it is empty, if another string follows it, or if it is the last value of the message; otherwise, the first
byte of the next value ends it.
0xFF only if it is empty, if another string follows it, or if nothing follows it in the message; otherwise, the byte
after it ends it: the first byte of the next value, or the 0xFE that ends an array or object. For example, `["e"]` is
serialized as 0x81 0x65 0xFF, but `[1,2,3,4,"e"]` as 0x85 0x91 0x92 0x93 0x94 0x65 0xFE.
!!! success "Complete mapping"
@@ -92,7 +93,7 @@ byte of the next value ends it.
- Object keys are written in the order of the object type, which is sorted for `json`, but not for
[`ordered_json`](../../api/ordered_json.md).
??? example
??? example "Example: serialize a JSON value to BON8"
```cpp
--8<-- "examples/to_bon8.cpp"
@@ -140,13 +141,13 @@ Non-negative integers are read as number_unsigned, negative integers as number_i
arrays and objects with up to four elements that are terminated by 0xFE, unsorted object keys, or a 0xFF after a
string that would also end without it, are accepted. A second 0xFF is not a terminator but an empty string.
Strings must be valid UTF-8, and the last string of a message must be terminated by 0xFF.
Strings must be valid UTF-8, and a string at the very end of a message must be terminated by 0xFF.
!!! info
Any BON8 output created by `to_bon8` can be successfully parsed by `from_bon8`.
??? example
??? example "Example: deserialize a JSON value from BON8"
```cpp
--8<-- "examples/from_bon8.cpp"
@@ -48,7 +48,7 @@ The library uses the following mapping from JSON values types to BSON types:
As a result, serializing and deserializing a JSON object containing such a value produces a different JSON object,
even though the binary data is unchanged.
??? example
??? example "Example: serialize a JSON value to BSON"
```cpp
--8<-- "examples/to_bson.cpp"
@@ -118,7 +118,7 @@ The library maps BSON record types to JSON value types as follows:
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
byte-by-byte as a C string, or are not required to hold text, respectively.
??? example
??? example "Example: deserialize a JSON value from BSON"
```cpp
--8<-- "examples/from_bson.cpp"
@@ -98,7 +98,7 @@ see "binary" cells in the table above.
Binary subtypes will be serialized as tagged items. See [binary values](../binary_values.md#cbor) for an example.
??? example
??? example "Example: serialize a JSON value to CBOR"
```cpp
--8<-- "examples/to_cbor.cpp"
@@ -203,7 +203,7 @@ The library maps CBOR types to JSON value types as follows:
Tagged items (0xC0..0xDB) will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`, in which case the tag is skipped and the enclosed data item is parsed on its own. Passing `cbor_tag_handler_t::store` to function `from_cbor` stores tagged byte strings (for bytes 0xd8..0xdb) as binary values with the tag as subtype; other tagged values are read as if the tag were ignored. If several tags precede a byte string, only the innermost one is stored. Note that no tag is ever interpreted: for instance, a text string tagged with tag 0 (date/time) stays a string.
??? example
??? example "Example: deserialize a JSON value from CBOR"
```cpp
--8<-- "examples/from_cbor.cpp"
@@ -79,7 +79,7 @@ specification:
total), because the check used to select the smaller float 32 encoding compared magnitudes with NaN, which is
always `false` and caused the float 32 path to be skipped.
??? example
??? example "Example: serialize a JSON value to MessagePack"
```cpp
--8<-- "examples/to_msgpack.cpp"
@@ -162,7 +162,7 @@ The library maps MessagePack types to JSON value types as follows:
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
to hold text.
??? example
??? example "Example: deserialize a JSON value from MessagePack"
```cpp
--8<-- "examples/from_msgpack.cpp"
@@ -11,29 +11,29 @@ achieve the generality of JSON, combined with being much easier to process than
The library uses the following mapping from JSON values types to UBJSON types according to the UBJSON specification:
| JSON value type | value/range | UBJSON type | marker |
|-----------------|-----------------------------------|----------------|--------|
| null | `null` | null | `Z` |
| boolean | `true` | true | `T` |
| boolean | `false` | false | `F` |
| number_integer | -9223372036854775808..-2147483649 | int64 | `L` |
| number_integer | -2147483648..-32769 | int32 | `l` |
| number_integer | -32768..-129 | int16 | `I` |
| number_integer | -128..127 | int8 | `i` |
| number_integer | 128..255 | uint8 | `U` |
| number_integer | 256..32767 | int16 | `I` |
| number_integer | 32768..2147483647 | int32 | `l` |
| number_integer | 2147483648..9223372036854775807 | int64 | `L` |
| number_unsigned | 0..127 | int8 | `i` |
| number_unsigned | 128..255 | uint8 | `U` |
| number_unsigned | 256..32767 | int16 | `I` |
| number_unsigned | 32768..2147483647 | int32 | `l` |
| number_unsigned | 2147483648..9223372036854775807 | int64 | `L` |
| number_unsigned | 2147483649..18446744073709551615 | high-precision | `H` |
| number_float | *any value* | float64 | `D` |
| string | *with shortest length indicator* | string | `S` |
| array | *see notes on optimized format* | array | `[` |
| object | *see notes on optimized format* | map | `{` |
| JSON value type | value/range | UBJSON type | marker |
|-----------------|-------------------------------------------|----------------|--------|
| null | `null` | null | `Z` |
| boolean | `true` | true | `T` |
| boolean | `false` | false | `F` |
| number_integer | -9223372036854775808..-2147483649 | int64 | `L` |
| number_integer | -2147483648..-32769 | int32 | `l` |
| number_integer | -32768..-129 | int16 | `I` |
| number_integer | -128..127 | int8 | `i` |
| number_integer | 128..255 | uint8 | `U` |
| number_integer | 256..32767 | int16 | `I` |
| number_integer | 32768..2147483647 | int32 | `l` |
| number_integer | 2147483648..9223372036854775807 | int64 | `L` |
| number_unsigned | 0..127 | int8 | `i` |
| number_unsigned | 128..255 | uint8 | `U` |
| number_unsigned | 256..32767 | int16 | `I` |
| number_unsigned | 32768..2147483647 | int32 | `l` |
| number_unsigned | 2147483648..9223372036854775807 | int64 | `L` |
| number_unsigned | 9223372036854775808..18446744073709551615 | high-precision | `H` |
| number_float | *any value* | float64 | `D` |
| string | *with shortest length indicator* | string | `S` |
| array | *see notes on optimized format* | array | `[` |
| object | *see notes on optimized format* | map | `{` |
!!! success "Complete mapping"
@@ -57,7 +57,7 @@ The library uses the following mapping from JSON values types to UBJSON types ac
!!! info "NaN/infinity handling"
If NaN or Infinity are stored inside a JSON number, they are serialized properly. This behavior differs from the
`dump()` function which serializes NaN or Infinity to `null`.
[`dump()`](../../api/basic_json/dump.md) function which serializes NaN or Infinity to `null`.
!!! info "Optimized formats"
@@ -82,7 +82,7 @@ The library uses the following mapping from JSON values types to UBJSON types ac
documentation. In particular, this means that serialization and the deserialization of a JSON containing binary
values into UBJSON and back will result in a different JSON object.
??? example
??? example "Example: serialize JSON values to UBJSON, with and without size/type optimization"
```cpp
--8<-- "examples/to_ubjson.cpp"
@@ -120,7 +120,7 @@ The library maps UBJSON types to JSON value types as follows:
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
??? example
??? example "Example: deserialize a JSON value from UBJSON"
```cpp
--8<-- "examples/from_ubjson.cpp"
+16 -12
View File
@@ -27,7 +27,7 @@ vector <|-- binary_t
By default, binary values are stored as `std::vector<std::uint8_t>`. This type can be changed by providing a template
parameter to the `basic_json` type. To store binary subtypes, the storage type is extended and exposed as
`json::binary_t`:
[`json::binary_t`](../api/basic_json/binary_t.md):
```cpp
auto binary = json::binary_t({0xCA, 0xFE, 0xBA, 0xBE});
@@ -62,21 +62,23 @@ JSON values can be constructed from `json::binary_t`:
json j = binary;
```
Binary values are primitive values just like numbers or strings:
Binary values are primitive values just like numbers or strings, as reflected by
[`is_binary()`](../api/basic_json/is_binary.md) and [`is_primitive()`](../api/basic_json/is_primitive.md):
```cpp
j.is_binary(); // returns true
j.is_primitive(); // returns true
```
Given a binary JSON value, the `binary_t` can be accessed by reference as via `get_binary()`:
Given a binary JSON value, the `binary_t` can be accessed by reference via
[`get_binary()`](../api/basic_json/get_binary.md):
```cpp
j.get_binary().has_subtype(); // returns true
j.get_binary().size(); // returns 4
```
For convenience, binary JSON values can be constructed via `json::binary`:
For convenience, binary JSON values can be constructed via [`json::binary`](../api/basic_json/binary.md):
```cpp
auto j2 = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 23);
@@ -99,7 +101,7 @@ JSON does not have a binary type, and this library does not introduce a new type
Instead, binary values are serialized as an object with two keys: `bytes` holds an array of integers, and `subtype`
is an integer or `null`.
??? example
??? example "Example: serialize a binary value to JSON"
Code:
@@ -133,7 +135,7 @@ is an integer or `null`.
[BJData](binary_formats/bjdata.md) neither supports binary values nor subtypes and proposes to serialize binary values
as an array of uint8 values. The library implements this translation.
??? example
??? example "Example: serialize a binary value to BJData"
Code:
@@ -192,7 +194,7 @@ as an array of uint8 values. The library implements this translation.
[BON8](binary_formats/bon8.md) neither supports binary values nor subtypes. The library serializes binary values as an
array of integers.
??? example
??? example "Example: serialize a binary value to BON8"
Code:
@@ -227,7 +229,7 @@ array of integers.
[BSON](binary_formats/bson.md) supports binary values and subtypes. If a subtype is given, it is used and added as an
unsigned 8-bit integer. If no subtype is given, the generic binary subtype 0x00 is used.
??? example
??? example "Example: serialize a binary value to BSON"
Code:
@@ -269,7 +271,7 @@ unsigned 8-bit integer. If no subtype is given, the generic binary subtype 0x00
value will be serialized as byte strings. The library will choose the smallest representation using the length of the
byte array.
??? example
??? example "Example: serialize a binary value to CBOR"
Code:
@@ -294,7 +296,9 @@ byte array.
```
Note that the subtype is serialized as tag. However, parsing tagged values yield a parse error unless
`json::cbor_tag_handler_t::ignore` or `json::cbor_tag_handler_t::store` is passed to `json::from_cbor`.
`json::cbor_tag_handler_t::ignore` or `json::cbor_tag_handler_t::store` is passed to
[`json::from_cbor`](../api/basic_json/from_cbor.md) (see
[`cbor_tag_handler_t`](../api/basic_json/cbor_tag_handler_t.md)).
```json
{
@@ -313,7 +317,7 @@ ext32. The subtype is then added as a signed 8-bit integer.
If no subtype is given, the bin family (bin8, bin16, bin32) is used.
??? example
??? example "Example: serialize a binary value to MessagePack"
Code:
@@ -353,7 +357,7 @@ If no subtype is given, the bin family (bin8, bin16, bin32) is used.
[UBJSON](binary_formats/ubjson.md) neither supports binary values nor subtypes and proposes to serialize binary values
as an array of uint8 values. The library implements this translation.
??? example
??? example "Example: serialize a binary value to UBJSON"
Code:
+1 -1
View File
@@ -11,7 +11,7 @@ This library does not support comments *by default*. It does so for three reason
3. It is dangerous for interoperability if some libraries add comment support while others do not. Please check [The Harmful Consequences of the Robustness Principle](https://tools.ietf.org/html/draft-iab-protocol-maintenance-01) on this.
However, you can set parameter `ignore_comments` to `#!cpp true` in the [`parse`](../api/basic_json/parse.md) function to ignore `//` or `/* */` comments. Comments will then be treated as whitespace. Combined with `ignore_trailing_commas` (also a `parse` parameter), this covers what is commonly referred to as **JSONC** (JSON with Comments, as used e.g. by Visual Studio Code's `.jsonc` files) -- comments and trailing commas, nothing more. This is a different, smaller extension than [JSON5](https://json5.org), which additionally allows unquoted keys, single-quoted strings, and other syntax changes that this library does not support.
However, you can set parameter `ignore_comments` to `#!cpp true` in the [`parse`](../api/basic_json/parse.md) function to ignore `//` or `/* */` comments. Comments will then be treated as whitespace. Combined with [`ignore_trailing_commas`](trailing_commas.md) (also a `parse` parameter), this covers what is commonly referred to as **JSONC** (JSON with Comments, as used e.g. by Visual Studio Code's `.jsonc` files) -- comments and trailing commas, nothing more. This is a different, smaller extension than [JSON5](https://json5.org), which additionally allows unquoted keys, single-quoted strings, and other syntax changes that this library does not support.
For more information, see [JSON With Commas and Comments (JWCC)](https://nigeltao.github.io/blog/2021/json-with-commas-comments.html).
@@ -6,7 +6,7 @@ The [`at`](../../api/basic_json/at.md) member function performs checked access;
desired value if it exists and throws a [`basic_json::out_of_range` exception](../../home/exceptions.md#out-of-range)
otherwise.
??? example "Read access"
??? example "Example: read access"
Consider the following JSON value:
@@ -31,7 +31,7 @@ otherwise.
The return value is a reference, so it can be used to modify the original value.
??? example "Write access"
??? example "Example: write access"
```cpp
j.at("name") = "John Smith";
@@ -50,7 +50,7 @@ The return value is a reference, so it can be used to modify the original value.
When accessing an invalid index (i.e., an index greater than or equal to the array size) or the passed object key is
non-existing, an exception is thrown.
??? example "Accessing via invalid index or missing key"
??? example "Example: access via invalid index or missing key"
```cpp
j.at("hobbies").at(3) = "cooking";
@@ -41,9 +41,9 @@ you want to access and a default value in case there is no value stored with tha
The value function is a template, and the return type of the function is determined by the type of the provided
default value unless otherwise specified. This can have unexpected effects. In the example below, we store a 64-bit
unsigned integer. We get exactly that value when using [`operator[]`](../../api/basic_json/operator[].md). However,
when we call `value` and provide `#!c 0` as default value, then `#!c -1` is returned. This occurs, because `#!c 0`
has type `#!c int` which overflows when handling the value `#!c 18446744073709551615`.
unsigned integer. We get exactly that value when using [`operator[]`](../../api/basic_json/operator%5B%5D.md).
However, when we call `value` and provide `#!c 0` as default value, then `#!c -1` is returned. This occurs,
because `#!c 0` has type `#!c int` which overflows when handling the value `#!c 18446744073709551615`.
To address this issue, either provide a correctly typed default value or use the template parameter to specify the
desired return type. Note that this issue occurs even when a value is stored at the provided key, and the default
@@ -5,5 +5,19 @@ There are many ways elements in a JSON value can be accessed:
- unchecked access via [`operator[]`](unchecked_access.md)
- checked access via [`at`](checked_access.md)
- access with default value via [`value`](default_value.md)
- iterators
- JSON pointers
- [iterators](../iterators.md)
- [JSON pointers](../json_pointer.md)
Testing whether a key or index exists before accessing it is also possible, with
[`contains`](../../api/basic_json/contains.md) or [`find`](../../api/basic_json/find.md) (which returns an iterator to
the value, or `end()` if it is not found).
```mermaid
flowchart TD
A["accessing a value"] --> B{"must it exist?"}
B -->|"yes, missing is an error"| C["at() -- throws"]
B -->|"yes, but checking is my job"| D["operator[] -- unchecked"]
B -->|"no, a fallback is fine"| E["value() -- default value"]
A --> F{"just testing first?"}
F -->|"yes"| G["contains() / find()"]
```
@@ -5,7 +5,7 @@
Elements in a JSON object and a JSON array can be accessed via [`operator[]`](../../api/basic_json/operator%5B%5D.md)
similar to a `#!cpp std::map` and a `#!cpp std::vector`, respectively.
??? example "Read access"
??? example "Example: read access"
Consider the following JSON value:
@@ -31,7 +31,7 @@ similar to a `#!cpp std::map` and a `#!cpp std::vector`, respectively.
The return value is a reference, so it can modify the original value. In case the passed object key is non-existing, a
`#!json null` value is inserted which can immediately be overwritten.
??? example "Write access"
??? example "Example: write access"
```cpp
j["name"] = "John Smith";
@@ -52,7 +52,7 @@ The return value is a reference, so it can modify the original value. In case th
When accessing an invalid index (i.e., an index greater than or equal to the array size), the JSON array is resized such
that the passed index is the new maximal index. Intermediate values are filled with `#!json null`.
??? example "Filling up arrays with `#!json null` values"
??? example "Example: filling up arrays with `#!json null` values"
```cpp
j["hobbies"][0] = "running";
@@ -94,8 +94,8 @@ that the passed index is the new maximal index. Intermediate values are filled w
- It is **undefined behavior** to access a const object with a non-existing key.
- It is **undefined behavior** to access a const array with an invalid index.
- In debug mode, an **assertion** will fire in both cases. You can disable assertions by defining the preprocessor
symbol `#!cpp NDEBUG` or redefine the macro [`JSON_ASSERT(x)`](../macros.md#json_assertx). See the documentation
on [runtime assertions](../assertions.md) for more information.
symbol `#!cpp NDEBUG` or redefine the macro [`JSON_ASSERT(x)`](../../api/macros/json_assert.md). See the
documentation on [runtime assertions](../assertions.md) for more information.
!!! failure "Exceptions"
@@ -105,8 +105,9 @@ that the passed index is the new maximal index. Intermediate values are filled w
## Performance: reserving array capacity
There is no public `reserve(count)` member on `basic_json` for pre-allocating array capacity. If you are building
a large array incrementally (e.g., via repeated `push_back()`) and know its final size ahead of time, you can
reserve capacity via `get_ref()` to access the underlying `array_t` directly:
a large array incrementally (e.g., via repeated [`push_back()`](../../api/basic_json/push_back.md)) and know its final
size ahead of time, you can reserve capacity via [`get_ref()`](../../api/basic_json/get_ref.md) to access the
underlying `array_t` directly:
```cpp
json j = json::array();
+34 -3
View File
@@ -29,6 +29,9 @@ The [`NLOHMANN_JSON_SERIALIZE_ENUM()` macro](../api/macros/nlohmann_json_seriali
## Usage
Serialization converts an enum value to its mapped string, deserialization does the reverse, and an unrecognized JSON
value deserializes to the first pair in the map:
```cpp
// enum to JSON as string
json j = TS_STOPPED;
@@ -43,6 +46,18 @@ json jPi = 3.14;
assert(jPi.get<TaskState>() == TS_INVALID );
```
??? example "Example: serializing/deserializing enums, including a second enum type"
```cpp
--8<-- "examples/nlohmann_json_serialize_enum.cpp"
```
Output:
```json
--8<-- "examples/nlohmann_json_serialize_enum.output"
```
## Notes
Just as in [Arbitrary Type Conversions](arbitrary_types.md) above,
@@ -54,9 +69,25 @@ Just as in [Arbitrary Type Conversions](arbitrary_types.md) above,
Other Important points:
- When using `get<ENUM_TYPE>()`, undefined JSON values will default to the first pair specified in your map. Select this
default pair carefully. If you desire an exception in this circumstance use [`NLOHMANN_JSON_SERIALIZE_ENUM_STRICT()`](../api/macros/nlohmann_json_serialize_enum_strict.md)
which behaves identically except for throwing an exception on unrecognized values.
- When using [`get<ENUM_TYPE>()`](../api/basic_json/get.md), undefined JSON values will default to the first pair
specified in your map. Select this default pair carefully. If you desire an exception in this circumstance use
[`NLOHMANN_JSON_SERIALIZE_ENUM_STRICT()`](../api/macros/nlohmann_json_serialize_enum_strict.md) which behaves
identically except for throwing an
[`out_of_range.410`](../home/exceptions.md#jsonexceptionout_of_range410) exception on unrecognized values, both when
serializing an enum value not listed in the map and when deserializing a JSON value that matches none of the map's
entries.
- If an enum or JSON value is specified more than once in your map, the first matching occurrence from the top of the
map will be returned when converting to or from JSON.
- To disable the default serialization of enumerators as integers and force a compiler error instead, see [`JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md).
??? example "Example: `NLOHMANN_JSON_SERIALIZE_ENUM_STRICT` throwing on unrecognized values"
```cpp
--8<-- "examples/nlohmann_json_serialize_enum_strict_err.cpp"
```
Output:
```json
--8<-- "examples/nlohmann_json_serialize_enum_strict_err.output"
```
+5 -1
View File
@@ -10,7 +10,8 @@ C++ types, and finally serialize it again.
understand the `#!cpp {}` vs. `#!cpp []` ambiguity.
- [Parsing](parsing/index.md) — read a JSON value from a string, file, or stream, including
[JSON Lines](parsing/json_lines.md), [callbacks](parsing/parser_callbacks.md), the
[SAX interface](parsing/sax_interface.md), and [error handling](parsing/parse_exceptions.md).
[SAX interface](parsing/sax_interface.md), [error handling](parsing/parse_exceptions.md), and
[parsing untrusted input](parsing/untrusted_input.md).
- [Comments](comments.md) and [trailing commas](trailing_commas.md) — opt-in relaxations of the JSON grammar.
## Accessing and modifying values
@@ -43,7 +44,10 @@ C++ types, and finally serialize it again.
- [Types](types/index.md) and [number handling](types/number_handling.md) — how JSON types map to C++ types and how
numbers are treated.
- [Template parameter requirements](types/template_parameters.md) — what a type passed as one of `basic_json`'s
template parameters has to provide.
- [Object order](object_order.md) — keep insertion order with [`ordered_json`](../api/ordered_json.md).
- [Performance](performance.md) — practical advice on parsing, memory use, serialization, and compile times.
- [Runtime assertions](assertions.md), [supported macros](macros.md), the [`nlohmann` namespace](namespace.md), and
[C++ modules](modules.md) — build-time and runtime configuration.
+13 -7
View File
@@ -4,7 +4,10 @@
A `basic_json` value is a container and allows access via iterators. Depending on the value type, `basic_json` stores zero or more values.
As for other containers, `begin()` returns an iterator to the first value and `end()` returns an iterator to the value following the last value. The latter iterator is a placeholder and cannot be dereferenced. In case of null values, empty arrays, or empty objects, `begin()` will return `end()`.
As for other containers, [`begin()`](../api/basic_json/begin.md) returns an iterator to the first value and
[`end()`](../api/basic_json/end.md) returns an iterator to the value following the last value. The latter iterator is a
placeholder and cannot be dereferenced. In case of null values, empty arrays, or empty objects, `begin()` will return
`end()`.
![Illustration from cppreference.com](../images/range-begin-end.svg)
@@ -12,7 +15,7 @@ As for other containers, `begin()` returns an iterator to the first value and `e
When iterating over objects, values are ordered with respect to the `object_comparator_t` type which defaults to `std::less`. See the [types documentation](types/index.md#key-order) for more information.
??? example
??? example "Example: iteration order of object values"
```cpp
// create JSON object {"one": 1, "two": 2, "three": 3}
@@ -41,7 +44,7 @@ When iterating over objects, values are ordered with respect to the `object_comp
The JSON iterators have two member functions, `key()` and `value()` to access the object key and stored value, respectively. When calling `key()` on a non-object iterator, an [invalid_iterator.207](../home/exceptions.md#jsonexceptioninvalid_iterator207) exception is thrown.
??? example
??? example "Example: access object keys with `key()` and `value()`"
```cpp
// create JSON object {"one": 1, "two": 2, "three": 3}
@@ -76,7 +79,9 @@ for (auto it : j_object)
}
```
For this reason, the `items()` function allows accessing `iterator::key()` and `iterator::value()` during range-based for loops. In these loops, a reference to the JSON values is returned, so there is no access to the underlying iterator.
For this reason, the [`items()`](../api/basic_json/items.md) function allows accessing `iterator::key()` and
`iterator::value()` during range-based for loops. In these loops, a reference to the JSON values is returned, so there
is no access to the underlying iterator.
```cpp
for (auto& el : j_object.items())
@@ -104,11 +109,12 @@ for (auto& [key, val] : j_object.items())
### Reverse iteration order
`rbegin()` and `rend()` return iterators in the reverse sequence.
[`rbegin()`](../api/basic_json/rbegin.md) and [`rend()`](../api/basic_json/rend.md) return iterators in the reverse
sequence.
![Illustration from cppreference.com](../images/range-rbegin-rend.svg)
??? example
??? example "Example: reverse iteration with `rbegin()` and `rend()`"
```cpp
json j = {1, 2, 3, 4};
@@ -132,7 +138,7 @@ for (auto& [key, val] : j_object.items())
Note that "value" means a JSON value in this setting, not values stored in the underlying containers. That is, `*begin()` returns the complete string or binary array and is also safe if the underlying string or binary array is empty.
??? example
??? example "Example: iterate over a string value"
```cpp
json j = "Hello, world";
+28 -5
View File
@@ -3,10 +3,17 @@
## Patches
JSON Patch ([RFC 6902](https://tools.ietf.org/html/rfc6902)) defines a JSON document structure for expressing a sequence
of operations to apply to a JSON document. With the `patch` function, a JSON Patch is applied to the current JSON value
by executing all operations from the patch.
of operations to apply to a JSON document. Operations address locations in the document using
[JSON Pointer](json_pointer.md) paths. With the [`patch`](../api/basic_json/patch.md) function, a JSON Patch is applied
to the current JSON value by executing all operations from the patch, yielding the patched document as a new value.
??? example
!!! tip "Applying a patch without copying"
[`patch`](../api/basic_json/patch.md) leaves the original value unchanged and returns the patched result as a copy.
If the document is large and the original value is no longer needed,
[`patch_inplace`](../api/basic_json/patch_inplace.md) applies the same operations in place instead.
??? example "Example: apply a JSON Patch"
The following code shows how a JSON patch is applied to a value.
@@ -22,7 +29,15 @@ by executing all operations from the patch.
## Diff
The library can also calculate a JSON patch (i.e., a **diff**) given two JSON values.
The library can also calculate a JSON patch (i.e., a **diff**) given two JSON values with the
[`diff`](../api/basic_json/diff.md) function.
```mermaid
flowchart LR
S["source"] -->|"diff(source, target)"| P["patch"]
S -->|"source.patch(patch)"| T["target"]
P -.->|"applied to source, yields"| T
```
!!! success "Invariant"
@@ -32,7 +47,7 @@ The library can also calculate a JSON patch (i.e., a **diff**) given two JSON va
source.patch(diff(source, target)) == target;
```
??? example
??? example "Example: create a JSON Patch from the difference of two values"
The following code shows how a JSON patch is created as a diff for two JSON values.
@@ -45,3 +60,11 @@ The library can also calculate a JSON patch (i.e., a **diff**) given two JSON va
```json
--8<-- "examples/diff.output"
```
## See also
- [JSON Pointer](json_pointer.md) - the addressing scheme used for patch paths
- [JSON Merge Patch](merge_patch.md) - a simpler, less expressive alternative patch format
- [`patch`](../api/basic_json/patch.md) - apply a JSON Patch, returning the result as a copy
- [`patch_inplace`](../api/basic_json/patch_inplace.md) - apply a JSON Patch without copying
- [`diff`](../api/basic_json/diff.md) - compute a JSON Patch from two values
+2 -1
View File
@@ -128,4 +128,5 @@ auto j_original = j_flat.unflatten();
- Class [`json_pointer`](../api/json_pointer/index.md)
- Function [`flatten`](../api/basic_json/flatten.md)
- Function [`unflatten`](../api/basic_json/unflatten.md)
- [JSON Patch](json_patch.md)
- [JSON Patch](json_patch.md) - paths inside a patch are JSON Pointers
- [JSON Merge Patch](merge_patch.md) - an alternative patch format that does not use JSON Pointer
+2 -1
View File
@@ -179,7 +179,8 @@ See [full documentation of `JSON_USE_GLOBAL_UDLS`](../api/macros/json_use_global
## `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`
When defined to `1`, the library restores the legacy behavior in which a discarded value compared equal to itself. This
behavior is deprecated and switched off (`0`) by default.
behavior is [deprecated](../integration/migration_guide.md#miscellaneous-functions) and switched off (`0`) by
default.
See [full documentation of `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../api/macros/json_use_legacy_discarded_value_comparison.md).
+12 -2
View File
@@ -1,9 +1,13 @@
# JSON Merge Patch
The library supports JSON Merge Patch ([RFC 7386](https://tools.ietf.org/html/rfc7386)) as a patch format.
The merge patch format is primarily intended for use with the HTTP PATCH method as a means of describing a set of modifications to a target resource's content. This function applies a merge patch to the current JSON value.
The merge patch format is primarily intended for use with the HTTP PATCH method as a means of describing a set of
modifications to a target resource's content. This function applies a merge patch to the current JSON value.
Instead of using [JSON Pointer](json_pointer.md) to specify values to be manipulated, it describes the changes using a syntax that closely mimics the document being modified.
Instead of using [JSON Pointer](json_pointer.md) to specify values to be manipulated, it describes the changes using a
syntax that closely mimics the document being modified. Unlike [JSON Patch](json_patch.md), a JSON Merge Patch cannot
express every kind of change (e.g., it cannot reorder array elements or remove a specific array element), but it is
easier to read and write for object-shaped documents.
??? example
@@ -18,3 +22,9 @@ Instead of using [JSON Pointer](json_pointer.md) to specify values to be manipul
```json
--8<-- "examples/merge_patch.output"
```
## See also
- [JSON Patch and Diff](json_patch.md) - a more expressive alternative that describes a sequence of operations
- [JSON Pointer](json_pointer.md) - the addressing scheme used by JSON Patch
- Function [`merge_patch`](../api/basic_json/merge_patch.md)
+2 -2
View File
@@ -6,7 +6,7 @@ The [JSON standard](https://tools.ietf.org/html/rfc8259.html) defines objects as
The default type `nlohmann::json` uses a `std::map` to store JSON objects, and thus stores object keys **sorted alphabetically**.
??? example
??? example "Example: `json` sorts object keys"
```cpp
#include <iostream>
@@ -39,7 +39,7 @@ The default type `nlohmann::json` uses a `std::map` to store JSON objects, and t
If you do want to preserve the **insertion order**, you can use the type [`nlohmann::ordered_json`](../api/ordered_json.md).
??? example
??? example "Example: `ordered_json` preserves insertion order"
```cpp
--8<-- "examples/ordered_json.cpp"
@@ -3,6 +3,16 @@
This library can create a JSON value from a wide range of inputs. This page gives an overview of the available parsing
functions and how they behave; the linked pages go into more detail.
```mermaid
flowchart LR
I["JSON input"] --> P["parse()"]
I --> S["sax_parse()"]
I --> A["accept()"]
P -->|"optional parser callback filters values"| D["basic_json value (DOM)"]
S --> H["events delivered to a user SAX handler"]
A --> V["bool: is the input valid JSON?"]
```
## Input
The [`parse`](../../api/basic_json/parse.md) function reads a JSON value from an input. The input can be
@@ -76,3 +86,4 @@ options.
- [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [parsing and exceptions](parse_exceptions.md) - control error handling
- [parsing untrusted input](untrusted_input.md) - what to consider when parsing input from untrusted sources
@@ -46,8 +46,20 @@ JSON Lines input with more than one value is treated as invalid JSON by the [`pa
}
```
with a JSON Lines input does not work, because the parser will try to parse one value after the last one.
with a JSON Lines input does not work, because the parser will try to parse one value after the last one and throw
a [`parse_error.101`](../../home/exceptions.md#jsonexceptionparse_error101) exception. The same happens for a
stream of *concatenated* (non-newline-delimited) JSON values: `operator>>` reads them one at a time, but the loop
above throws after the last value. To read either format with `operator>>`, check for the end of the stream before
each read:
This is different from parsing a stream of *concatenated* (non-newline-delimited) JSON values, for which
`operator>>` does work, provided that a value that is a number is followed by whitespace -- see its
[notes](../../api/operator_gtgt.md#notes) for details.
```cpp
json j;
while (input >> std::ws && input.peek() != std::char_traits<char>::eof())
{
input >> j;
std::cout << j << std::endl;
}
```
A value that is a number must be followed by whitespace -- see the [notes](../../api/operator_gtgt.md#notes) of
`operator>>` for details.
@@ -23,9 +23,9 @@ In case exceptions are undesired or not supported by the environment, there are
## Switch off exceptions
The `parse()` function accepts a `#!cpp bool` parameter `allow_exceptions` which controls whether an exception is
thrown when a parse error occurs (`#!cpp true`, default) or whether a discarded value should be returned
(`#!cpp false`).
The [`parse()`](../../api/basic_json/parse.md) function accepts a `#!cpp bool` parameter `allow_exceptions` which
controls whether an exception is thrown when a parse error occurs (`#!cpp true`, default) or whether a discarded value
should be returned (`#!cpp false`).
```cpp
json j = json::parse(my_input, nullptr, false);
@@ -39,8 +39,8 @@ Note there is no diagnostic information available in this scenario.
## Use accept() function
Alternatively, function `accept()` can be used which does not return a `json` value, but a `#!cpp bool` indicating
whether the input is valid JSON.
Alternatively, function [`accept()`](../../api/basic_json/accept.md) can be used which does not return a `json` value,
but a `#!cpp bool` indicating whether the input is valid JSON.
```cpp
if (!json::accept(my_input))
@@ -66,56 +66,18 @@ bool parse_error(std::size_t position,
The return value indicates whether the parsing should continue, so the function should usually return `#!cpp false`.
??? example
??? example "Example: report parse errors without exceptions"
The example derives from the library's DOM parser and overrides `parse_error` to print the error instead of
throwing. Note the DOM parser is an implementation detail (`nlohmann::detail`) and may change between releases;
see [Do not use the `detail` namespace](../../integration/migration_guide.md#do-not-use-the-detail-namespace).
```cpp
#include <iostream>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
class sax_no_exception : public nlohmann::detail::json_sax_dom_parser<json>
{
public:
sax_no_exception(json& j)
: nlohmann::detail::json_sax_dom_parser<json>(j, false)
{}
bool parse_error(std::size_t position,
const std::string& last_token,
const json::exception& ex)
{
std::cerr << "parse error at input byte " << position << "\n"
<< ex.what() << "\n"
<< "last read: \"" << last_token << "\""
<< std::endl;
return false;
}
};
int main()
{
std::string myinput = "[1,2,3,]";
json result;
sax_no_exception sax(result);
bool parse_result = json::sax_parse(myinput, &sax);
if (!parse_result)
{
std::cerr << "parsing unsuccessful!" << std::endl;
}
std::cout << "parsed value: " << result << std::endl;
}
--8<-- "examples/sax_no_exception.cpp"
```
Output:
```
parse error at input byte 8
[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal
last read: "3,]"
parsing unsuccessful!
parsed value: [1,2,3]
--8<-- "examples/sax_no_exception.output"
```
@@ -2,8 +2,9 @@
## Overview
With a parser callback function, the result of parsing a JSON text can be influenced. When passed to `parse`, it is
called on certain events (passed as `parse_event_t` via parameter `event`) with a set recursion depth `depth` and
With a parser callback function, the result of parsing a JSON text can be influenced. When passed to
[`parse`](../../api/basic_json/parse.md), it is called on certain events (passed as
[`parse_event_t`](../../api/basic_json/parse_event_t.md) via parameter `event`) with a set recursion depth `depth` and
context JSON value `parsed`. The return value of the callback function is a boolean indicating whether the element that
emitted the callback shall be kept or not.
@@ -30,7 +31,7 @@ table describes the values of the parameters `depth`, `event`, and `parsed`.
| `parse_event_t::array_end` | the parser read `]` and finished processing a JSON array | depth of the parent of the JSON array | the parsed JSON array |
| `parse_event_t::value` | the parser finished reading a JSON value | depth of the value | the parsed JSON value |
??? example
??? example "Example: sequence of callback events"
When parsing the following JSON text,
@@ -76,7 +77,7 @@ was called:
- In case a value outside a structured type is skipped, it is replaced with `#!json null`. This case happens if the
top-level element is skipped.
??? example
??? example "Example: skip an object key while parsing"
The example below demonstrates the `parse()` function with and without callback function.
@@ -98,7 +99,7 @@ the resulting `#!c json` value -- once parsing has produced that value, the dupl
storage maps each key to a single value. If duplicate keys should instead be treated as an error, a parser callback
can detect them while the object is still being read, before that ambiguity ever applies.
??? example
??? example "Example: reject duplicate object keys"
```cpp
--8<-- "examples/reject_duplicate_keys.cpp"
@@ -110,16 +111,18 @@ can detect them while the object is still being read, before that ambiguity ever
--8<-- "examples/reject_duplicate_keys.output"
```
This approach has two limitations:
This approach has three limitations:
- The depth-indexed bookkeeping must account for the fact that `object_start` reports the depth of the *parent* of
the object, while the `key` events inside that object are reported one depth deeper (see the event table above);
it is easy to get this off by one for nested objects.
- The thrown exception cannot carry a `parse_error`-style byte offset, because position tracking only exists inside
the parser and lexer, not at the callback layer.
- The exception only names the repeated key, not where it occurs in the document. Reporting its full path requires
maintaining a stack of the enclosing keys and array indices in the callback as well.
For strict validation with precise error positions, implementing a [SAX interface](sax_interface.md) instead gives
access to the parser's position information directly.
A [SAX interface](sax_interface.md) does not lift the position limitation: its `key` function receives no position
either -- only `parse_error` is passed the byte position.
## Recipe: streaming a large homogeneous array
@@ -129,7 +132,7 @@ discard it, so memory usage stays bounded by a single element (plus the not-yet-
than the whole document. Since the top-level array's `array_start`/`array_end` are reported at `depth == 0` (its
parent is the document root), the object elements it contains are reported at `depth == 1`:
??? example
??? example "Example: stream a large top-level array"
```cpp
std::ifstream input("large_array.json");
@@ -154,7 +157,7 @@ homogeneous values by checking `object_end`/`value` events at `depth == 1` there
Since there is no built-in nesting-depth limit (see the note above), a callback can enforce one manually by
tracking the maximum `depth` seen and throwing once it is exceeded:
??? example
??? example "Example: limit the nesting depth"
```cpp
constexpr int max_depth = 32;
@@ -0,0 +1,163 @@
# Parsing Untrusted Input
This page is for applications that parse JSON -- or one of the supported [binary formats](../binary_formats/index.md)
(BJData, BON8, BSON, CBOR, MessagePack, UBJSON) -- from a source they do not fully control, such as a network
connection, an uploaded file, or another process. It summarizes what the library already does for such input and what
remains the caller's responsibility, linking to the pages that cover each aspect in detail rather than repeating them.
For the project's threat model and the countermeasures behind these behaviors, see the
[assurance case](../../community/assurance_case.md); to report a vulnerability, see the
[security policy](../../community/security_policy.md).
## Errors without exceptions
By default, [`parse()`](../../api/basic_json/parse.md) throws a
[`parse_error`](../../home/exceptions.md#jsonexceptionparse_error101) (for instance `parse_error.101` for a syntax
error) when the input is not valid. If your environment cannot use exceptions for untrusted input, the library offers
several alternatives; see [Parsing and exceptions](parse_exceptions.md) for the full comparison:
- Pass `#!cpp false` as the third argument to `parse()` to get a discarded value
(checked with [`is_discarded()`](../../api/basic_json/is_discarded.md)) instead of a thrown exception, with no
diagnostic information.
- Use [`accept()`](../../api/basic_json/accept.md) to only check whether the input is valid JSON, without building a
value.
- Implement the [SAX interface](sax_interface.md) and override `parse_error()` to react to an error yourself, with the
byte position and the exception that would otherwise have been thrown; see the
[example](parse_exceptions.md#user-defined-sax-interface) that overrides it to print instead of throw.
If exceptions are unavailable entirely (`-fno-exceptions`, or [`JSON_NOEXCEPTION`](../../api/macros/json_noexception.md)
defined), every `#!cpp throw` in the library becomes a call to `std::abort()` -- there is no way to recover from a
parse error of untrusted input in that configuration; see
[Switch off exceptions](../../home/exceptions.md#switch-off-exceptions) for the details and for overriding this with
`JSON_THROW_USER`.
## Nesting depth
The JSON parser and the binary readers are iterative: they keep the containers they are currently inside of on a
heap-allocated stack instead of calling themselves once per nesting level, so the native call stack does not grow with
the nesting depth of the input. A deeply nested document is therefore bounded by available memory, not by the call
stack, however deeply it is nested.
!!! warning "No built-in depth limit while parsing"
Neither the parser nor the binary readers impose a limit on how deep the input may nest. An attacker can still
exhaust memory (though not the call stack) with a sufficiently deep document. If you need to reject over-deep
untrusted input outright, track the depth yourself, either with a
[parser callback](parser_callbacks.md#recipe-max-nesting-depth-via-a-callback) for the JSON parser, or by counting
`start_object`/`start_array` and `end_object`/`end_array` calls in a
[SAX handler](sax_interface.md) (for the JSON parser or a binary format alike) and throwing once your limit is
exceeded.
Once a value has been parsed, operations that walk it recursively -- serializing it with
[`dump`](../../api/basic_json/dump.md), hashing it, copying it, comparing two values with `#!cpp ==`, `#!cpp <`, or (in
C++20) `#!cpp <=>`, merging with [`update`](../../api/basic_json/update.md), and applying a
[`merge_patch`](../../api/basic_json/merge_patch.md) -- descend at most 128 levels on the call stack and continue
below that with an explicit stack instead, so none of them can exhaust the stack either, however deeply the value is
nested. Destroying a value (its destructor) never recurses at all, regardless of nesting depth, for the same reason.
!!! note "Not every operation is bounded yet"
[`diff`](../../api/basic_json/diff.md), [`flatten`](../../api/basic_json/flatten.md), and the binary writers
(`to_cbor`, `to_msgpack`, ...) still recurse once per nesting level; this is called out as work in progress in the
[assurance case](../../community/assurance_case.md#secure-design). A value deep enough to matter for these
operations would typically first have to survive parsing without hitting a self-imposed depth limit, as described
above.
## Input size
The library does not limit the overall size of a JSON text; a value nested or wide enough will use memory
proportional to the input. If you parse untrusted input of unbounded size, check the size of the file or stream
yourself before -- or while -- handing it to `parse()`.
For the binary formats, an announced size is never trusted outright:
- Reading a string or binary value copies the input in bounded 4096-byte chunks and grows the result as bytes are
actually consumed, rather than allocating the announced length up front -- a truncated input runs out of bytes
(reported as a parse error) instead of triggering an oversized allocation.
- When an array announces its number of elements and the array container supports `reserve()` (as `#!cpp std::vector`,
the default, does), the library reserves storage for at most 16384 of them upfront, regardless of how large the
announced count is; further elements still grow the container normally as they are read.
- An announced array or object size that exceeds what the target container could ever hold (its `max_size()`) is
rejected immediately as [`out_of_range.408`](../../home/exceptions.md#jsonexceptionout_of_range408), without
attempting to allocate anything.
## Strings
Invalid UTF-8 is rejected while parsing, not just while serializing:
- In JSON text, an ill-formed UTF-8 byte in a string is a
[`parse_error.101`](../../home/exceptions.md#jsonexceptionparse_error101) ("invalid string: ill-formed UTF-8 byte").
- In a binary format, a string that is not valid UTF-8 is a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113).
A `#!cpp '\0'` (NUL) byte *inside* a quoted JSON string is always rejected (it must be escaped as `\u0000`). A NUL byte
*outside* of a string is different: by default it is silently treated as the end of the input, so trailing bytes after
it -- including further, otherwise well-formed JSON -- are silently ignored rather than rejected. Since untrusted input
that happens to embed a NUL is a way to make part of it disappear without a parse error, see the
[FAQ entry](../../home/faq.md#nul-bytes-in-the-input) and consider defining
[`JSON_STRICT_NUL_HANDLING`](../../api/macros/json_strict_nul_handling.md) to `1` to reject a NUL byte like any other
unexpected byte instead.
Parsing is not the only place invalid UTF-8 matters: a string that reached a `#!cpp json` value some other way (for
example, constructed by application code, or, before JSON_STRICT_NUL_HANDLING existed, read from a binary format that
does not validate strings) still has to round-trip back to JSON text. By default,
[`dump()`](../../api/basic_json/dump.md) throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
if the string is not valid UTF-8; passing
[`error_handler_t::replace`](../../api/basic_json/error_handler_t.md) or `error_handler_t::ignore` avoids the exception
instead of crashing an application that forgot to catch it. See
[Handling invalid UTF-8](../serialization.md#handling-invalid-utf-8) for the options and an example.
## Duplicate object keys
The JSON specification leaves the handling of repeated keys in an object up to the implementation, and this library
does too: as described in [`object_t`](../../api/basic_json/object_t.md#behavior), it is unspecified which of the
values for a repeated key ends up in the parsed object. If your application must reject duplicate keys instead of
silently resolving them one way or another, see the
[parser callback recipe for rejecting duplicate keys](parser_callbacks.md#recipe-rejecting-duplicate-object-keys).
## Numbers
A number whose value cannot be represented -- for instance `1E1000`, which overflows `double` -- is rejected while
parsing as [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) rather than silently becoming
infinity. An integer that is syntactically valid but does not fit the 64-bit integer types is not rejected; it is
instead stored as a `double`, which may lose precision for very large values. See
[number limits](../types/number_handling.md#number-limits) for the exact ranges and an example.
## Comments and trailing commas
Both [comments](../comments.md) and [trailing commas](../trailing_commas.md) are rejected by default, matching the
JSON specification; they must be explicitly enabled per call with the `ignore_comments` and `ignore_trailing_commas`
parameters of [`parse()`](../../api/basic_json/parse.md) or [`accept()`](../../api/basic_json/accept.md). Do not
enable either for input whose conformance you cannot otherwise control, since interoperability with strictly
conforming JSON consumers is exactly what the default rejects.
## Checklist
- Wrap parsing in a `#!cpp try`/`#!cpp catch` block, or use `allow_exceptions=false`/`accept()` if your environment
cannot use exceptions; never let `JSON_NOEXCEPTION`'s `abort()` be the first time you think about error handling.
- If the input's nesting depth matters to you, enforce your own limit with a
[parser callback](parser_callbacks.md#recipe-max-nesting-depth-via-a-callback) or a
[SAX handler](sax_interface.md); the library bounds the call stack but not memory use.
- Bound the size of the input itself before parsing, independent of the library's own bounded allocations for binary
format lengths.
- Decide up front how you want a string with invalid UTF-8 -- from any source, not only parsing -- to be serialized
(`strict`, `replace`, or `ignore`), rather than discovering it from an uncaught `type_error.316`.
- If a stray NUL byte silently truncating trailing input is a problem for your input format, define
`JSON_STRICT_NUL_HANDLING`.
- Decide whether duplicate object keys should be an error for your application, and add a callback if so.
- Do not enable `ignore_comments` or `ignore_trailing_commas` for input that must be strictly conforming JSON.
For the broader design rationale and how it is tested (fuzzing, sanitizers, static analysis), see the
[assurance case](../../community/assurance_case.md) and [quality assurance](../../community/quality_assurance.md). To
report a security issue in the library itself, follow the [security policy](../../community/security_policy.md).
## See also
- [Parsing](index.md) - overview of the parsing functions
- [Parsing and exceptions](parse_exceptions.md) - error handling without exceptions
- [Parser callbacks](parser_callbacks.md) - depth limits, duplicate-key rejection, and streaming recipes
- [SAX interface](sax_interface.md) - implement a custom handler with access to parse errors and positions
- [Serialization](../serialization.md) - handling invalid UTF-8 when dumping
- [Number handling](../types/number_handling.md) - number ranges and overflow behavior
- [Assurance case](../../community/assurance_case.md) - the library's threat model and countermeasures
- [Security policy](../../community/security_policy.md) - how to report a vulnerability
+215
View File
@@ -0,0 +1,215 @@
# Performance
Speed was never the primary goal of this library. The [design goals](../home/design_goals.md) page says so plainly:
"There are certainly faster JSON libraries out there." Intuitive syntax, trivial integration, and thorough testing came
first. If a hard real-time budget or the last percent of throughput matters more than convenience, a
[faster, more specialized library](https://github.com/miloyip/nativejson-benchmark#parsing-time) may be a better fit.
That said, how you use this library still makes a measurable difference. This page collects practical, code-verified
techniques for reducing time, memory, and compile-time cost -- without repeating the detailed pages it links to.
## Parsing input
[`parse`](../api/basic_json/parse.md) accepts a string, a pair of iterators, a container, a `#!cpp std::istream`, or a
`#!cpp FILE*` (see [Parsing](parsing/index.md#input)). Internally, every input is wrapped in an
[input adapter](../home/architecture.md#input-adapters), and not all adapters are equally fast.
For inputs backed by contiguous, single-byte memory -- a `#!cpp std::string`, a `#!cpp std::vector<char>`, a string
literal, or a pointer range -- the library uses `iterator_input_adapter`, wrapped in a raw pointer so the fast paths
below apply on every supported standard. This adapter exposes two optimizations the lexer detects at compile time:
- it can reconstruct already-consumed input on demand for error messages, instead of copying every character as it is
read, and
- the lexer can scan ordinary string characters directly out of the buffer, several bytes at a time, rather than one
character (and one function call) at a time.
A `#!cpp std::istream` (including `#!cpp std::ifstream`) or `#!cpp FILE*`, by contrast, is read through
`input_stream_adapter` or `file_input_adapter`, which read one character (or one block, for binary formats) at a time
and expose neither optimization -- the lexer falls back to the same byte-at-a-time path it uses for any
non-contiguous, general-purpose iterator range. An iterator pair over non-contiguous but random-access storage (e.g.
`#!cpp std::deque<char>::iterator`) gets the first optimization but not the second, since the byte-scanning fast path
additionally requires contiguous storage.
Practically: if the JSON text is already in memory, or small enough to read into memory, prefer passing a
`#!cpp std::string`, a `#!cpp std::vector<char>`, or a pointer range to `parse` over a `#!cpp std::istream`. For a
file, that means reading it into a string first and then parsing the string, rather than passing a
`#!cpp std::ifstream` directly to `parse` -- the latter never benefits from either optimization:
```cpp
// gets the contiguous fast paths
std::ifstream f("example.json");
std::string contents((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
json j = json::parse(contents);
// does not: input_stream_adapter has no fast path
std::ifstream f2("example.json");
json j2 = json::parse(f2);
```
For contiguous input with many non-ASCII characters, [`JSON_USE_SIMDUTF`](../api/macros/json_use_simdutf.md) can
additionally speed up UTF-8 validation by using the [simdutf](https://github.com/simdutf/simdutf) library instead of
the built-in scalar validator; streaming inputs (files, `#!cpp std::istream`, wide strings, user-defined adapters)
always use the scalar path regardless of this macro.
## Large documents
Parsing always produces SAX events internally; [`parse`](../api/basic_json/parse.md) simply feeds them to a consumer
that builds a complete `basic_json` value tree (a DOM) in memory. For documents too large to comfortably hold as
a DOM, two alternatives avoid building it:
- Implement the [SAX interface](parsing/sax_interface.md) directly and pass it to
[`sax_parse`](../api/basic_json/sax_parse.md); only the parts of the input you choose to keep ever become
`basic_json` values.
- Pass a [parser callback](parsing/parser_callbacks.md) to `parse`. This still builds a DOM, but the callback can
discard finished elements as soon as they are handled, so memory usage stays bounded by one element (plus the
unparsed remainder of the input) instead of the whole document -- see the [recipe for streaming a large homogeneous
array](parsing/parser_callbacks.md#recipe-streaming-a-large-homogeneous-array).
If the data is naturally record-oriented, consider [JSON Lines](parsing/json_lines.md) instead of one large JSON
document: reading and parsing it line by line with `#!cpp std::getline` means only one line's value is ever in memory
at a time, and a malformed line does not invalidate lines already processed.
## Binary formats
JSON text is not a compact format. If the data is only exchanged between programs (not read by humans), the
[binary formats](binary_formats/index.md) -- BJData, BON8, BSON, CBOR, MessagePack, and UBJSON -- encode the same
values more compactly, which reduces both the bytes transferred and, for most of them, the work needed to parse them
back. The [size comparison](binary_formats/index.md#sizes) on that page, measured against minified JSON for four
reference documents, shows the effect varies a lot by document shape: CBOR and MessagePack come out at 50.5% of the
minified JSON size for the numeric-array-heavy `canada.json`, but only around 87-88% for the string-heavy
`jeopardy.json`, where there is less numeric data to encode more compactly. BON8 is the most compact option in that
comparison for text-heavy documents (63.5%-87.5%), at the cost of an
[incomplete serializer](binary_formats/index.md#completeness) (no unsigned integers above int64). Which format -- and
whether it is worth the loss of human readability at all -- depends on the actual data; see the
[comparison tables](binary_formats/index.md#comparison) before choosing one.
## Object type: `json` vs. `ordered_json`
The default [`json`](../api/json.md) type stores object keys in a `#!cpp std::map`, giving logarithmic-time lookup,
insertion, and erasure, at the cost of sorting keys alphabetically rather than preserving insertion order (see
[Object Order](object_order.md)). [`ordered_json`](../api/ordered_json.md) uses
[`nlohmann::ordered_map`](../api/ordered_map.md) instead, a `#!cpp std::vector`-backed container with no lookup index:
every key-based operation is a **linear scan**, so building an object of `n` distinct keys costs **O(n²)** in total --
this applies equally to inserting keys one by one and to parsing an object, since the parser inserts each key as it is
read. The [measurements on the `ordered_map` page](../api/ordered_map.md#complexity) show this is
negligible at typical object sizes (2000 keys: 0.7 ms for `json` vs. 3.6 ms for `ordered_json`, a 5x factor) but grows
steeply for large, machine-generated objects (16 000 keys: 3.3 ms vs. 181.6 ms, a 54x factor).
If insertion order matters *and* an object routinely has many thousands of keys, `ordered_json`'s quadratic build cost
may not be acceptable. The library's [`ObjectType` template parameter](types/template_parameters.md#objecttype) can be
set to a different container instead: `#!cpp nlohmann::fifo_map` keeps insertion order with a real lookup index
(avoiding the quadratic cost), while `#!cpp std::unordered_map`, `#!cpp boost::unordered_flat_map`,
`#!cpp absl::flat_hash_map`, and similar hash maps trade insertion order for average-case constant-time lookup (through
an adapter, since their template argument order does not match what `basic_json` expects) -- see
[Object Order](object_order.md#alternative-behavior-preserve-insertion-order) for the full list.
## Avoiding copies
- **Move instead of copy.** Constructing a `basic_json` from an existing one is
[linear in its size](../api/basic_json/basic_json.md#complexity) for the copy constructor but
[constant](../api/basic_json/basic_json.md#complexity) for the move constructor. The same applies to assigning a
large `#!cpp std::string`, `#!cpp std::vector`, or other container into a value: pass it as `#!cpp std::move(x)`
rather than `x` whenever `x` is no longer needed afterwards.
- **Access without copying.** [`get<T>()`](../api/basic_json/get.md) returns a copy of the stored value converted to
`T`. When a reference or pointer to the value already stored inside the `basic_json` is enough,
[`get_ref()`](../api/basic_json/get_ref.md) and [`get_ptr()`](../api/basic_json/get_ptr.md) access it directly:
both pages state, word for word, "No copies are made." -- at the cost of that reference or pointer becoming invalid
once the underlying value changes.
- **Iterate by reference.** `#!cpp basic_json::iterator::operator*()` returns a `reference` (an alias for
`#!cpp basic_json&`), but a range-based for loop with a by-value loop variable (`#!cpp for (auto el : j)`) still
copies each element, because plain `#!cpp auto` drops the reference. Write `#!cpp for (const auto& el : j)` (or
`#!cpp auto&` for a mutable loop), and use [`items()`](../api/basic_json/items.md) the same way when the key is
needed too -- its own examples use `#!cpp for (auto& el : j.items())`.
- **Construct in place.** [`emplace_back()`](../api/basic_json/emplace_back.md) (arrays, amortized constant time) and
[`emplace()`](../api/basic_json/emplace.md) (objects, logarithmic in the size of the container for `json`) forward
their arguments directly to a `basic_json` constructor, rather than requiring a temporary value to be
constructed and then copied or moved in. [`push_back()`](../api/basic_json/push_back.md) has an rvalue overload
(`#!cpp push_back(basic_json&&)`) for a value that already exists: `#!cpp j.push_back(std::move(value))` moves it
in instead of copying it.
- **Skip the bounds check when it is redundant.** [`at()`](../api/basic_json/at.md) and
[`operator[]`](../api/basic_json/operator%5B%5D.md) have the same complexity (constant for a valid array index,
logarithmic for an object key in `json`) -- the difference is that `at()` additionally checks the key or index and
throws if it is invalid, while `operator[]` does not (see [unchecked access](element_access/unchecked_access.md) and
[checked access](element_access/checked_access.md)). Prefer `operator[]` when the surrounding code has already
established that the access is valid.
- **Reserve array capacity.** `basic_json` has no public `reserve()`, but when building a large array
incrementally with a known final size, [`get_ref()`](../api/basic_json/get_ref.md) exposes the underlying
`#!cpp array_t` so it can be reserved directly -- see
["reserving array capacity"](element_access/unchecked_access.md#performance-reserving-array-capacity) for the
one-line recipe.
## Serialization
[`dump()`](../api/basic_json/dump.md) with the default `#!cpp indent = -1` selects "the most compact representation"
(word for word from the page); any non-negative `indent` pretty-prints instead, which is more readable but produces
more bytes and more work. `dump()` builds and returns a complete `#!cpp string_t` containing the whole serialization.
[`operator<<`](../api/operator_ltlt.md) writes directly to a `#!cpp std::ostream` instead, through the same
serializer, but without ever materializing that intermediate string -- so if the destination is a stream (a file, or
`#!cpp std::cout`), `#!cpp os << j;` avoids the allocation and copy that `#!cpp os << j.dump();` would incur for large
values.
## Diagnostics overhead
Two opt-in macros add diagnostic information to exceptions and to every value, at a cost that is only worth paying
while it is in use:
- [`JSON_DIAGNOSTICS`](../api/macros/json_diagnostics.md) adds a JSON Pointer to exception messages, pointing at the
value that triggered the exception. Quoting the page directly: "enabling this macro increases the size of every
JSON value by one pointer and adds some runtime overhead" -- every value gains a parent pointer that has to be kept
up to date as the document is built and modified.
- [`JSON_DIAGNOSTIC_POSITIONS`](../api/macros/json_diagnostic_positions.md) adds
[`start_pos()`](../api/basic_json/start_pos.md) and [`end_pos()`](../api/basic_json/end_pos.md), the byte offsets a
value occupied in its parsed input. Quoting the page: "enabling this macro increases the size of every JSON value by
two `std::size_t` fields and adds slight runtime overhead to parsing, copying JSON value objects, and the generation
of error messages for exceptions."
Both default to off. Enable them where better diagnostics are worth the overhead (for example, while validating
untrusted input, or in a debug build), and keep them off in a release build that does not need them.
## Compile time
[`<nlohmann/json_fwd.hpp>`](../home/architecture.md#source-layout) forward-declares
[`basic_json`](../api/basic_json/index.md), [`json`](../api/json.md), [`ordered_json`](../api/ordered_json.md),
[`json_pointer`](../api/json_pointer/index.md), and [`adl_serializer`](../api/adl_serializer/index.md), pulling in only
a handful of lightweight standard headers instead of the full `json.hpp`. A header that only needs to *name*
`nlohmann::json` -- in a function signature or a class member declaration, for instance -- can include `json_fwd.hpp`
and leave `#!cpp #include <nlohmann/json.hpp>` to the source files that actually parse, build, or serialize values,
the same way a project would forward-declare any other heavy class to keep it out of widely-included headers:
```cpp
// my_type.hpp
#include <nlohmann/json_fwd.hpp>
class my_type
{
nlohmann::json config() const;
};
// my_type.cpp
#include <nlohmann/json.hpp>
#include "my_type.hpp"
nlohmann::json my_type::config() const { /* ... */ }
```
One caveat: ABI-affecting macros such as `JSON_DIAGNOSTICS` and `JSON_DIAGNOSTIC_POSITIONS` are encoded into the
library's [inline namespace name](namespace.md#limitations). Every translation unit -- whether it includes
`json_fwd.hpp` or the full header -- must define them the same way, or linking fails with undefined references
instead of a compile error.
If I/O support is not needed at all, [`JSON_NO_IO`](../api/macros/json_no_io.md) excludes `<cstdio>`, `<ios>`,
`<iosfwd>`, `<istream>`, and `<ostream>` outright and drops the `#!cpp std::istream`/`#!cpp FILE*` `parse` overloads and
[`operator<<`](../api/operator_ltlt.md) that depend on them (`dump()` itself is unaffected, since it only returns a
string); it exists for environments where those headers are unavailable (such as Intel SGX), and as a side effect
those headers are then never processed by the compiler at all.
## See also
- [Design goals](../home/design_goals.md) - why this library does not optimize for speed first
- [Architecture](../home/architecture.md) - how input adapters, the lexer, and the serializer fit together
- [Parsing](parsing/index.md) - the available parsing functions and inputs
- [SAX interface](parsing/sax_interface.md) - parse without building a DOM
- [Binary formats](binary_formats/index.md) - compact alternatives to JSON text
- [Object Order](object_order.md) - `json` vs. `ordered_json` and other `ObjectType` choices
- [Template Parameter Requirements](types/template_parameters.md) - custom container and allocator types
- [Supported macros](macros.md) - overview of all configuration macros, including the diagnostics ones above
+2 -2
View File
@@ -28,7 +28,7 @@ std::cout << j << std::endl;
By default, `dump` produces the most compact representation without any superfluous whitespace. Passing a non-negative
`indent` argument pretty-prints the output with the given number of spaces per level:
??? example
??? example "Example: pretty-print JSON values with `dump()`"
```cpp
--8<-- "examples/dump.cpp"
@@ -65,7 +65,7 @@ serialization fails by default. The fourth argument of `dump` selects an
- `replace` — replace invalid bytes with the Unicode replacement character U+FFFD (`�`).
- `ignore` — silently drop invalid bytes.
??? example
??? example "Example: serialize invalid UTF-8 with different error handlers"
```cpp
--8<-- "examples/error_handler_t.cpp"
@@ -67,14 +67,29 @@ Positive integers are stored as `#!c std::uint64_t`, while negative integers are
distinction is determined at parse time: if the JSON number has a leading minus sign, it uses signed integer storage;
otherwise, it uses unsigned integer storage.
```mermaid
flowchart TD
A["number literal"] --> B{"has a fraction (.) or exponent (e/E)?"}
B -->|"yes"| F["number_float_t"]
B -->|"no"| C{"has a leading minus sign?"}
C -->|"yes"| D["try number_integer_t"]
C -->|"no"| E["try number_unsigned_t"]
D -->|"overflow"| F
E -->|"overflow"| F
```
!!! info "Notes"
- Numbers with a decimal digit or scientific notation are always stored as `#!c double`.
- The number types can be changed, see [Template number types](#template-number-types).
- As of version 3.9.1, the conversion is realized by
- Integers are converted by the library's own digit parser. Floating-point numbers are converted with
[`std::from_chars`](https://en.cppreference.com/w/cpp/utility/from_chars) if the library is compiled with C++17
and the standard library supports it, then with an exact fast path for `#!c double` values with few significant
digits, and otherwise with the locale-aware
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) (`std::strtof`/`std::strtold` for the
other floating-point types). Before version 3.13.0, the conversion was realized by
[`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul),
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof), respectively.
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and `std::strtod`, respectively.
!!! example "Examples"
@@ -26,8 +26,9 @@ Requirements are split into two groups:
diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside
the library. Four violations are not caught at compile time at all:
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers,
because the lexer hands the buffer to `#!cpp std::strtoull`/`#!cpp std::strtoll`/`#!cpp std::strtod`.
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and can silently misparse
floating-point numbers, because the lexer may hand the buffer to `#!cpp std::strtod`, which reads up to the
terminating null character.
- A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation,
and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance.
- The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion
@@ -213,7 +214,7 @@ The library does not sort or de-duplicate keys itself; the behavior described in
--8<-- "examples/custom_object_type.hpp"
```
??? example "Compiling and using it"
??? example "Example: use the custom `ObjectType`"
```cpp
--8<-- "examples/custom_object_type.cpp"
@@ -306,7 +307,7 @@ using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
--8<-- "examples/custom_array_type.hpp"
```
??? example "Compiling and using it"
??? example "Example: use the custom `ArrayType`"
```cpp
--8<-- "examples/custom_array_type.cpp"
@@ -348,16 +349,16 @@ using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
### Always required
- A member type `value_type` that is one byte wide and `char`-compatible. The library stores and processes UTF-8
encoded `char` data and hands `data()` to `#!cpp std::strtoull`/`#!cpp std::strtoll`.
encoded `char` data and passes `data()` to functions that take a `#!cpp const char*`, such as `#!cpp std::strtod`.
`#!cpp std::wstring`, `#!cpp std::u16string`, and `#!cpp std::u32string` are **not** valid choices; see the FAQ on
[wide string handling](../../home/faq.md#wide-string-handling).
- Constructors: default, copy, move, from `#!cpp const char*` (which must not be `#!cpp explicit`), from
`#!cpp (const char*, size_type)`, and from `#!cpp (size_type, char)`; and copy or move assignment.
- Member functions `size()`, `clear()`, `resize(n, c)`, `data()`, `push_back(char)`, and `operator[]`
(const and non-const, returning references). `c_str()` and `back()` are **not** required.
- `data()` must return a pointer to a contiguous, **null-terminated** buffer -- the parser hands it to
`#!cpp std::strtoull`. A type whose `data()` is not null-terminated does not fail to compile; it silently
misparses numbers.
- `data()` must return a pointer to a contiguous, **null-terminated** buffer -- the parser may hand it to
`#!cpp std::strtod`, which reads up to the null character. A type whose `data()` is not null-terminated does not
fail to compile; it can silently misparse floating-point numbers.
- `append(const char*, size_type)`, used by [`dump`](../../api/basic_json/dump.md), and `append(const StringType&)`,
used by the CBOR reader for indefinite-length strings. The library's internal string concatenation additionally has
to append a `#!cpp char` and a `#!cpp const char*`; for each it selects between `append(arg)`, `#!cpp operator+=`,
@@ -394,6 +395,7 @@ using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
| [`to_bson`](../../api/basic_json/to_bson.md) | `find(value_type)` and `npos` |
| [`parse`](../../api/basic_json/parse.md) from a `string_t` | the input adapters must accept it; otherwise pass a character range |
| `#!cpp operator<<(std::ostream&, const json_pointer&)` | streamability to `#!cpp std::ostream` |
| [`to_string`](../../api/basic_json/to_string.md) | conversion of `StringType` to `#!cpp std::string` (the function returns a `#!cpp std::string`) |
| exception messages | `data()` and `size()`, or `begin()` and `end()` |
### Compatible types
@@ -448,7 +450,7 @@ using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
--8<-- "examples/custom_string_type.hpp"
```
??? example "Compiling and using it"
??? example "Example: use the custom `StringType`"
```cpp
--8<-- "examples/custom_string_type.cpp"
@@ -535,8 +537,9 @@ therefore silently changes parse results rather than raising an error. See
`NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`:
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::strtof`, `#!cpp std::strtod`, or
`#!cpp std::strtold`; the library provides overloads for exactly these three types.
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::from_chars` or, as a fallback, with
`#!cpp std::strtof`, `#!cpp std::strtod`, or `#!cpp std::strtold`; the library provides overloads for exactly these
three types.
- [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`).
@@ -669,7 +672,7 @@ such a container to a `basic_json` value.
--8<-- "examples/custom_binary_type.hpp"
```
??? example "Compiling and using it"
??? example "Example: use the custom `BinaryType`"
```cpp
--8<-- "examples/custom_binary_type.cpp"