Files
json/docs/mkdocs/docs/features/parsing/index.md
T
78821cd9c2 docs: document standards compliance and parse() vs operator>> strictness (#5326)
* docs: document RFC 8259 / JSONTestSuite compliance and parse() vs operator>> strictness

The compliance story lived only in tests/src/unit-testsuites.cpp, so
drive-by comparisons kept claiming the library "does not fully pass
JSONTestSuite". Make it discoverable:

- README: add a "Standards compliance" note stating that both nst
  JSONTestSuite revisions run in CI, that all mandatory y_/n_ cases pass
  through the strict parse() entry point, and listing the deliberate
  implementation-defined i_ choices (unbounded nesting, silent BOM
  stripping, noncharacters forwarded, strict rejection of invalid UTF-8
  and lone surrogates, out_of_range.406 on numeric overflow).
- features/parsing: add a "Strictness and trailing data" section
  documenting that parse() is strict and rejects trailing data while
  operator>> follows relaxed iostream semantics (parses one value and
  leaves the stream positioned after it) -- the single place a naive
  test yields a "non-compliant" result.

Documentation only; no parser behavior change. Closes #5290.

Signed-off-by: manon <youdie006@users.noreply.github.com>

* docs: correct test-data vendoring and parse()/operator>> claims per review

- README: the JSONTestSuite data is downloaded from nlohmann/json_test_data at
  configure time, not vendored/committed; say so.
- README: only the updated suite runs y_ and n_ cases through strict parse();
  the original suite's y_ cases go through operator>>. Narrow the claim.
- parsing/index.md and operator_gtgt.md: note that operator>> consumes a number's
  terminating byte, so concatenated numbers must be whitespace-separated (1 2
  works, 1true does not); structural and literal values are unaffected.

Signed-off-by: manon <youdie006@users.noreply.github.com>

---------

Signed-off-by: manon <youdie006@users.noreply.github.com>
Co-authored-by: manon <youdie006@users.noreply.github.com>
2026-08-03 19:15:53 +02:00

4.1 KiB

Parsing

This library can create a JSON value from a wide range of inputs. This page gives an overview of the available parsing functions and how they behave; the linked pages go into more detail.

Input

The parse function reads a JSON value from an input. The input can be

  • a string (#!cpp std::string, C string, or string literal),
  • a #!cpp std::istream (e.g., an #!cpp std::ifstream reading from a file),
  • a #!cpp FILE* pointer,
  • a pair of iterators over a contiguous range (e.g., a #!cpp std::vector<std::uint8_t>), or
  • a contiguous container.
// parse from a string
json j = json::parse(R"({"happy": true, "pi": 3.141})");

// parse from a file
std::ifstream f("example.json");
json data = json::parse(f);

The input must be encoded in UTF-8; other encodings are not supported. A single input may contain only one JSON value. Inputs consisting of multiple values separated by newlines are handled by the JSON Lines format.

By default, the library rejects comments and trailing commas. Both can be enabled with parameters of the parse function — see comments and trailing commas.

Strictness and trailing data

parse reads a single JSON value and requires the whole input to be consumed: any non-whitespace data after the value is reported as a parse error. Use it when you want to guarantee that an input is exactly one complete JSON document.

operator>> follows relaxed #!cpp std::istream semantics instead: it parses one JSON value and leaves the stream positioned right after it, without requiring the rest of the stream to be consumed. This is what makes it possible to read several concatenated values from the same stream, but it also means that "a valid document followed by trailing bytes" is accepted rather than rejected. If you are validating conformance, or need to reject any input that is not exactly one JSON document, prefer parse.

When using operator>> to read several concatenated values this way, a value that is a number must be followed by whitespace, because operator>> consumes the character that terminates a number — see the operator>> notes for details and examples.

SAX vs. DOM parsing

The library offers two parsing models:

  • DOM parsing (the default): the complete input is read and stored as an in-memory basic_json value that can be traversed and modified freely. This is what parse does, and it is the right choice for most use cases.
  • SAX parsing: instead of building a value, the parser reports events (such as "a string was read" or "an object started") to a handler that you implement. This avoids building the full value in memory and is useful for very large inputs or when you only need to extract parts of the input. See the SAX interface for details and sax_parse for the API.

You can influence a DOM parse without switching to the SAX interface by passing a parser callback, which is called during parsing and can, for example, discard parts of the input.

Exceptions

When the input is not valid JSON, the parse function throws an exception by default. If exceptions are undesired or unavailable, the parser can instead return a discarded value, or accept can be used to only check whether an input is valid JSON. See parsing and exceptions for the available options.

See also