mirror of
https://github.com/nlohmann/json.git
synced 2026-10-04 05:30:31 +00:00
* docs: qualify the operator>> stream positioning guarantee operator>>'s notes state that it leaves the stream positioned right after the parsed value, so that concatenated JSON values can be read back to back. That does not hold when the value is a number: a number is only terminated by the character that follows it, and the lexer's unget() is simulated (it rewinds only the lexer's own bookkeeping), so that character stays consumed from the stream. Document the actual behaviour: the guarantee holds for all value types except numbers, which must be followed by whitespace. Also qualify the cross-reference on the JSON Lines page, which repeated the unqualified claim. Documentation only; the behaviour itself is tracked in #5340. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * fix: restore the character that terminates a number (#5340) operator>> is documented to leave the stream positioned right after the parsed value, so that concatenated JSON values can be read back to back. That did not hold for numbers: a number is only terminated by the character following it, and lexer::scan_number() reads that character and calls unget() -- which is simulated and rewinds only the lexer's own bookkeeping. input_stream_adapter consumes via sbumpc() with no matching sungetc(), so the terminating character stayed consumed and the next extraction started one byte too late ('1true' left the stream at 'rue'). Propagating unget() to the adapter directly does not work: next_unget makes the following get() replay the cached character, so the terminator would be delivered twice. Instead, restore the still-pending character once at the end of a non-strict parse, where the input is handed back to the caller: - input_stream_adapter gains unget_character() (sungetc()) and advertises it via supports_unget, detected the same way as supports_seek. - lexer::restore_pending_unget() turns a pending simulated unget of a real (non-EOF) character into a real one and clears next_unget so the character is not also replayed. It is a no-op for adapters that cannot unget, and reports failure when sungetc() fails, in which case the input is left as it was before. - parser calls it on the three non-strict paths, i.e. for operator>> and sax_parse(strict = false). Strict parse()/accept() are unaffected: they require the input to end after the value, so the character is consumed by the end-of-input check anyway. Parse error messages and reported positions are unchanged. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * tests: fix CI failures in the #5340 test helpers Four CI failures, all in the new test code: - GCC (-Werror=useless-cast): drop the `json(...)` wrapper around `json::parse(...)`, which already returns a `json`. - GCC (-Werror=unused-result): assign the discarded `json::parse()` result to a dummy, the idiom used elsewhere in the test suite, and catch `json::parse_error&` for consistency. - clang-tidy (google-default-arguments): remove the default argument from the `pbackfail()` override; `sungetc()` supplies the base declaration's default. - MSVC (bad allocation): `no_putback_streambuf::underflow()` set a one-character get area without advancing `m_pos`, so an implementation whose `istream::get` peeks before it bumps re-read the same character forever. Keep no get area at all: `underflow()` peeks, `uflow()` consumes, and `sungetc()` still always lands in `pbackfail()`, which is what the test needs. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * fix: leave the character that terminates a number in the input Read the character following a number without consuming it, instead of consuming it and putting it back. input_stream_adapter now peeks with sgetc() and only steps over the character when the next one is requested or when the adapter is destroyed, so releasing it cannot fail - no putback position is required from the streambuf. Suggested by gregmarr in #5344. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * docs: match the version history wording to the peek-based fix Signed-off-by: Niels Lohmann <mail@nlohmann.me> * docs: drop the whitespace-separator caveat from the parsing pages The caveat added in #5343 describes the behavior this branch fixes: a number no longer consumes the character that terminates it, so concatenated values need no separator. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * refactor: split the strict and non-strict paths in parser Folding the release_lookahead() call into the existing strict check left the "in strict mode" comment on an else-if branch, and made the strict condition in sax_parse() redundant with the branch it followed. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Put the stream position fix behind JSON_PRECISE_STREAM_POSITION Leaving the character that terminates a number in the stream is observable: reading "1,2,3" with repeated operator>> works today only because the comma after each number is swallowed, and std::getline after a number skips the line break. Both break with the fix, so make it opt-in for 3.x, as suggested by @gregmarr in the review. - JSON_PRECISE_STREAM_POSITION (default 0) selects the peek-based input_stream_adapter. Without it, the adapter is the consuming one from develop and has no supports_lookahead, so lexer::release_lookahead() and the parser's calls to it compile to nothing. - The macro changes input_stream_adapter's layout and member functions, so it gets the ABI tag _psp, after _bics. The ABI config tests, the natvis generator, and nlohmann_json.natvis (regenerated) know the tag. - The tests for the fix move to unit-precise-stream-position.cpp, which defines the macro itself and runs in every build, and gain the two cases above. unit-deserialization.cpp pins the default behavior instead. - The docs describe the default behavior again and point to the new macro page; version history says "added in 3.13.0, planned default in 4.0.0". Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
79 lines
4.2 KiB
Markdown
79 lines
4.2 KiB
Markdown
# Parsing
|
|
|
|
This library can create a JSON value from a wide range of inputs. This page gives an overview of the available parsing
|
|
functions and how they behave; the linked pages go into more detail.
|
|
|
|
## Input
|
|
|
|
The [`parse`](../../api/basic_json/parse.md) function reads a JSON value from an input. The input can be
|
|
|
|
- a string (`#!cpp std::string`, C string, or string literal),
|
|
- a `#!cpp std::istream` (e.g., an `#!cpp std::ifstream` reading from a file),
|
|
- a `#!cpp FILE*` pointer,
|
|
- a pair of iterators over a contiguous range (e.g., a `#!cpp std::vector<std::uint8_t>`), or
|
|
- a contiguous container.
|
|
|
|
```cpp
|
|
// parse from a string
|
|
json j = json::parse(R"({"happy": true, "pi": 3.141})");
|
|
|
|
// parse from a file
|
|
std::ifstream f("example.json");
|
|
json data = json::parse(f);
|
|
```
|
|
|
|
The input must be encoded in UTF-8; other encodings are not supported. A single input may contain only one JSON value.
|
|
Inputs consisting of multiple values separated by newlines are handled by the [JSON Lines](json_lines.md) format.
|
|
|
|
By default, the library rejects comments and trailing commas. Both can be enabled with parameters of the `parse`
|
|
function — see [comments](../comments.md) and [trailing commas](../trailing_commas.md).
|
|
|
|
## Strictness and trailing data
|
|
|
|
[`parse`](../../api/basic_json/parse.md) reads a single JSON value and requires the whole input to be consumed: any
|
|
non-whitespace data after the value is reported as a parse error. Use it when you want to guarantee that an input is
|
|
exactly one complete JSON document.
|
|
|
|
[`operator>>`](../../api/operator_gtgt.md) follows relaxed `#!cpp std::istream` semantics instead: it parses one JSON
|
|
value and leaves the stream positioned right after it, without requiring the rest of the stream to be consumed. This is
|
|
what makes it possible to read several concatenated values from the same stream, but it also means that "a valid
|
|
document followed by trailing bytes" is accepted rather than rejected. If you are validating conformance, or need to
|
|
reject any input that is not exactly one JSON document, prefer `parse`.
|
|
|
|
When using `operator>>` to read several concatenated values this way, a value that is a number must be followed by
|
|
whitespace, because `operator>>` consumes the character that terminates a number, unless
|
|
[`JSON_PRECISE_STREAM_POSITION`](../../api/macros/json_precise_stream_position.md) is defined to `1` — see the
|
|
[`operator>>` notes](../../api/operator_gtgt.md#notes) for details and examples.
|
|
|
|
## SAX vs. DOM parsing
|
|
|
|
The library offers two parsing models:
|
|
|
|
- **DOM parsing** (the default): the complete input is read and stored as an in-memory `basic_json` value that can be
|
|
traversed and modified freely. This is what [`parse`](../../api/basic_json/parse.md) does, and it is the right choice
|
|
for most use cases.
|
|
- **SAX parsing**: instead of building a value, the parser reports events (such as "a string was read" or "an object
|
|
started") to a handler that you implement. This avoids building the full value in memory and is useful for very large
|
|
inputs or when you only need to extract parts of the input. See the [SAX interface](sax_interface.md) for details and
|
|
[`sax_parse`](../../api/basic_json/sax_parse.md) for the API.
|
|
|
|
You can influence a DOM parse without switching to the SAX interface by passing a
|
|
[parser callback](parser_callbacks.md), which is called during parsing and can, for example, discard parts of the input.
|
|
|
|
## Exceptions
|
|
|
|
When the input is not valid JSON, the `parse` function throws an exception by default. If exceptions are undesired or
|
|
unavailable, the parser can instead return a discarded value, or [`accept`](../../api/basic_json/accept.md) can be used
|
|
to only check whether an input is valid JSON. See [parsing and exceptions](parse_exceptions.md) for the available
|
|
options.
|
|
|
|
## See also
|
|
|
|
- [`parse`](../../api/basic_json/parse.md) - deserialize from a compatible input
|
|
- [`accept`](../../api/basic_json/accept.md) - check if the input is valid JSON
|
|
- [`sax_parse`](../../api/basic_json/sax_parse.md) - generate SAX events
|
|
- [JSON Lines](json_lines.md) - parse newline-delimited JSON
|
|
- [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function
|
|
- [SAX interface](sax_interface.md) - implement a custom SAX handler
|
|
- [parsing and exceptions](parse_exceptions.md) - control error handling
|