Review and extend the documentation, and check it in CI (#5638)

* Review and extend the documentation, and check it in CI

A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.

Fixes:
- wrong signatures and version histories (operator!= C++20 member,
  binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
  example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
  suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed

Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
  Mermaid diagrams, new API pages for json_pointer::operator<=> and
  byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release

Checks:
- strict documentation build (broken links/anchors fail it); CI and
  the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
  header links, docset index coverage; its unused-example check works
  again
- all examples produce the same output on every platform

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep the customer links that could not be fixed

A dead link on the customers page is still the evidence of where the
use of the library was documented. Keep the original URLs of the entries
without a working replacement (Marne, Cisco Webex Desk Camera, Philips
Hue, CyberArk) and exclude exactly these URLs from the link check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the duplicate-key recipe's claim about SAX positions

The SAX interface's key() receives no position either; only parse_error()
does. Also note that the recipe does not report the path to the repeated
key (see discussion #5085).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Say the library is available as a single header and mention json_fwd.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct documentation errors found while hunting for bugs

- patch/patch_inplace: list the JSON pointer errors parse_error.106-109
  and out_of_range.402/404, and quote the actual parse_error.105 message.
- unflatten: list parse_error.106/107/108 and out_of_range.404.
- to_bson: list out_of_range.415 (binary subtype above 255) and note
  that 412 and 415 are new in 3.13.0.
- to_string: state that string_t must be convertible to std::string, also
  in the StringType requirements table.
- JSON Lines: a `while (input >> j)` loop also throws after the last value
  for concatenated JSON values; show a loop that works for both.
- BON8: a string gets 0xFF only if nothing follows it in the message; a
  string at the end of an array or object is ended by 0xFE.
- custom_string_type.hpp: add operator+=(char), which the "Always
  required" list asks for (json_pointer::to_string, flatten, unflatten,
  and diff did not compile), and an ADL int_to_string for diff and items.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cache the release headers with functools.lru_cache

Codacy (Pylint) flagged the mutable default argument that header() used
as its cache. functools.lru_cache keeps the same memoization without it.
The script's output is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-10-02 11:32:15 +02:00
committed by GitHub
parent 48ff79647f
commit 63c10a51fc
276 changed files with 5318 additions and 624 deletions
+113 -1
View File
@@ -4,6 +4,7 @@ import glob
import os.path
import re
import sys
import urllib.parse
import yaml
@@ -152,7 +153,7 @@ def check_structure() -> None:
def check_examples() -> None:
example_files = sorted(glob.glob("../../examples/*.cpp"))
example_files = sorted(glob.glob("examples/*.cpp"))
markdown_files = sorted(glob.glob("**/*.md", recursive=True))
# check if every example file is used in at least one markdown file
@@ -211,11 +212,122 @@ def check_links() -> None:
report("nav/duplicate_files", "mkdocs.yml", f'file "{duplicate_file}" is linked with multiple keys in "nav": {file_list_str}; only one is rendered properly, see #4564')
FENCE_RE = re.compile(r"^\s*(`{3,}|~{3,})")
INLINE_CODE_RE = re.compile(r"(`+).+?\1")
def markdown_lines(file):
"""Yield (lineno, line) for all lines outside fenced code blocks."""
fence = None
with open(file, encoding="utf-8") as content:
for lineno, line in enumerate(content, 1):
line = line.rstrip("\n")
match = FENCE_RE.match(line)
if fence is None:
if match:
fence = match.group(1)
else:
yield lineno, line
elif match and line.strip() == match.group(1) and match.group(1)[0] == fence[0] \
and len(match.group(1)) >= len(fence):
fence = None
def check_example_titles() -> None:
"""On API pages with more than one example, every example needs a title of the form "Example: ..."."""
example_re = re.compile(r'^\s*(?:\?\?\?\+?|!!!) example(?: "(.*)")?\s*$')
for file in sorted(glob.glob("api/**/*.md", recursive=True)):
examples = [(lineno, m.group(1)) for lineno, line in markdown_lines(file) if (m := example_re.match(line))]
if len(examples) < 2:
continue
for lineno, title in examples:
if title is None or not title.startswith("Example: "):
report("style/example_title", f"{file}:{lineno}",
f'pages with several examples need titles like "Example: ..." (found: {title!r})')
def check_heading_levels() -> None:
"""Headings start at level 1 and never skip a level."""
heading_re = re.compile(r"^(#{1,6})\s|^<h([1-6])[\s>]")
for file in sorted(glob.glob("**/*.md", recursive=True)):
previous = 0
for lineno, line in markdown_lines(file):
match = heading_re.match(line)
if not match:
continue
level = len(match.group(1)) if match.group(1) else int(match.group(2))
if previous == 0 and level != 1:
report("structure/heading_level", f"{file}:{lineno}", f"first heading should have level 1, not {level}")
elif level > previous + 1 and previous != 0:
report("structure/heading_level", f"{file}:{lineno}", f"heading level jumps from {previous} to {level}")
previous = level
def check_image_alt_text() -> None:
"""Images need an alternative text."""
empty_alt_re = re.compile(r"!\[\s*\][(\[]")
img_re = re.compile(r"<img\b[^>]*>", re.IGNORECASE)
alt_re = re.compile(r'\balt\s*=\s*"[^"]*\S[^"]*"', re.IGNORECASE)
for file in sorted(glob.glob("**/*.md", recursive=True)):
for lineno, line in markdown_lines(file):
line = INLINE_CODE_RE.sub("", line)
if empty_alt_re.search(line) or any(not alt_re.search(tag) for tag in img_re.findall(line)):
report("style/image_alt_text", f"{file}:{lineno}", "image without alternative text")
def check_header_links() -> None:
"""Links to the documentation in the library's headers point to existing pages."""
url_re = re.compile(r"https://json\.nlohmann\.me/([^\s#)>\"']*)")
for header in sorted(glob.glob("../../../include/nlohmann/**/*.hpp", recursive=True)):
with open(header, encoding="utf-8") as content:
for lineno, line in enumerate(content, 1):
for match in url_re.finditer(line):
path = urllib.parse.unquote(match.group(1)).strip("/")
if path and not (os.path.isfile(f"{path}.md") or os.path.isfile(f"{path}/index.md")):
report("links/header_link", f"{os.path.relpath(header, '../../..')}:{lineno}",
f'link to "{match.group(0)}" does not point to a documentation page')
def check_docset() -> None:
"""Every API page and every macro has an entry in the docset index; no entry points to a missing page."""
entry_re = re.compile(r"VALUES \('((?:[^']|'')*)', '(\w+)', '([^']*)'\);")
names_by_path = {}
with open("../../docset/docSet.sql", encoding="utf-8") as sql:
for name, _, path in entry_re.findall(sql.read()):
names_by_path.setdefault(path, set()).add(name.replace("''", "'"))
def to_path(page):
if os.path.basename(page) == "index.md":
return page[:-len("index.md")] + "index.html"
return page[:-len(".md")] + "/index.html"
pages = sorted(glob.glob("**/*.md", recursive=True))
for path in sorted(set(names_by_path) - {to_path(p) for p in pages}):
report("docset/stale_entry", "../../docset/docSet.sql", f'entry "{path}" has no documentation page')
for page in (p for p in pages if p.startswith("api/")):
names = names_by_path.get(to_path(page))
if not names:
report("docset/missing_entry", page, "page has no entry in docs/docset/docSet.sql")
elif page.startswith("api/macros/") and os.path.basename(page) != "index.md":
with open(page, encoding="utf-8") as content:
text = content.read()
match = re.search(r"^# (.+)$", text, re.MULTILINE) or re.search(r"<h1>(.*?)</h1>", text, re.DOTALL)
title = re.sub(r"<[^>]+>|\s+", " ", match.group(1))
for macro in filter(None, (x.strip() for x in re.split(r"[,/]", title))):
if macro not in names:
report("docset/missing_macro", page, f'macro "{macro}" has no entry in docs/docset/docSet.sql')
if __name__ == "__main__":
print(120 * "-")
check_structure()
check_examples()
check_links()
check_example_titles()
check_heading_levels()
check_image_alt_text()
check_header_links()
check_docset()
print(120 * "-")
if warnings > 0: