* Review and extend the documentation, and check it in CI
A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.
Fixes:
- wrong signatures and version histories (operator!= C++20 member,
binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed
Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
Mermaid diagrams, new API pages for json_pointer::operator<=> and
byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release
Checks:
- strict documentation build (broken links/anchors fail it); CI and
the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
header links, docset index coverage; its unused-example check works
again
- all examples produce the same output on every platform
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the customer links that could not be fixed
A dead link on the customers page is still the evidence of where the
use of the library was documented. Keep the original URLs of the entries
without a working replacement (Marne, Cisco Webex Desk Camera, Philips
Hue, CyberArk) and exclude exactly these URLs from the link check.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Correct the duplicate-key recipe's claim about SAX positions
The SAX interface's key() receives no position either; only parse_error()
does. Also note that the recipe does not report the path to the repeated
key (see discussion #5085).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Say the library is available as a single header and mention json_fwd.hpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Correct documentation errors found while hunting for bugs
- patch/patch_inplace: list the JSON pointer errors parse_error.106-109
and out_of_range.402/404, and quote the actual parse_error.105 message.
- unflatten: list parse_error.106/107/108 and out_of_range.404.
- to_bson: list out_of_range.415 (binary subtype above 255) and note
that 412 and 415 are new in 3.13.0.
- to_string: state that string_t must be convertible to std::string, also
in the StringType requirements table.
- JSON Lines: a `while (input >> j)` loop also throws after the last value
for concatenated JSON values; show a loop that works for both.
- BON8: a string gets 0xFF only if nothing follows it in the message; a
string at the end of an array or object is ended by 0xFE.
- custom_string_type.hpp: add operator+=(char), which the "Always
required" list asks for (json_pointer::to_string, flatten, unflatten,
and diff did not compile), and an ADL int_to_string for diff and items.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Cache the release headers with functools.lru_cache
Codacy (Pylint) flagged the mutable default argument that header() used
as its cache. functools.lru_cache keeps the same memoization without it.
The script's output is unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)
from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).
Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.
There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* ⚡ validate only newly read bytes of binary-format strings
get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
- warn about BSON marker 0x11 interoperability in both directions
- explain subtype-less binary normalization to subtype 0x00
- add a round-trip test for binary values without a subtype
Signed-off-by: YingqiDuan <141370165+YingqiDuan@users.noreply.github.com>