Put the stream position fix behind JSON_PRECISE_STREAM_POSITION

Leaving the character that terminates a number in the stream is observable:
reading "1,2,3" with repeated operator>> works today only because the comma
after each number is swallowed, and std::getline after a number skips the
line break. Both break with the fix, so make it opt-in for 3.x, as suggested
by @gregmarr in the review.

- JSON_PRECISE_STREAM_POSITION (default 0) selects the peek-based
  input_stream_adapter. Without it, the adapter is the consuming one from
  develop and has no supports_lookahead, so lexer::release_lookahead() and
  the parser's calls to it compile to nothing.
- The macro changes input_stream_adapter's layout and member functions, so
  it gets the ABI tag _psp, after _bics. The ABI config tests, the natvis
  generator, and nlohmann_json.natvis (regenerated) know the tag.
- The tests for the fix move to unit-precise-stream-position.cpp, which
  defines the macro itself and runs in every build, and gain the two cases
  above. unit-deserialization.cpp pins the default behavior instead.
- The docs describe the default behavior again and point to the new macro
  page; version history says "added in 3.13.0, planned default in 4.0.0".

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-25 08:41:32 +02:00
parent ddffd920d0
commit 87a33734c4
21 changed files with 1555 additions and 196 deletions
+4 -3
View File
@@ -70,7 +70,8 @@ The SAX event lister must follow the interface of [`json_sax`](../json_sax/index
`strict` (in)
: whether the input has to be consumed completely (optional, `#!cpp true` by default); when `#!cpp false` and the
input is a `#!cpp std::istream`, the stream is left positioned right after the parsed value
input is a `#!cpp std::istream`, the character that terminates a number is consumed unless
[`JSON_PRECISE_STREAM_POSITION`](../macros/json_precise_stream_position.md) is defined to `1`; see [`operator>>`](../operator_gtgt.md#notes)
`ignore_comments` (in)
: whether comments should be ignored and treated like whitespace (`#!cpp true`) or yield a parse error
@@ -137,8 +138,8 @@ A UTF-8 byte order mark is silently ignored.
- Added `ignore_trailing_commas` in version 3.13.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Changed in version 4.0.0 to leave a `#!cpp std::istream` positioned right after the parsed value when `strict` is
`#!cpp false`; see [`operator>>`](../operator_gtgt.md#notes).
- `JSON_PRECISE_STREAM_POSITION` added in version 3.13.0 to optionally leave a `#!cpp std::istream` positioned right
after the parsed value when `strict` is `#!cpp false`.
!!! warning "Deprecation"
+2
View File
@@ -16,6 +16,8 @@ header. See also the [macro overview page](../../features/macros.md).
## Parsing
- [**JSON_PRECISE_STREAM_POSITION**](json_precise_stream_position.md) - opt in to leaving an input stream positioned
right after a parsed number
- [**JSON_STRICT_NUL_HANDLING**](json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input instead of
treating it as end of input
@@ -0,0 +1,131 @@
# JSON_PRECISE_STREAM_POSITION
```cpp
#define JSON_PRECISE_STREAM_POSITION /* value */
```
When defined to `1`, [`operator>>`](../operator_gtgt.md) and [`sax_parse`](../basic_json/sax_parse.md) with
`strict = false` leave a `#!cpp std::istream` positioned right after the parsed value for every value type. By default,
the character that terminates a number is consumed as well.
The macro only affects reading from a `#!cpp std::istream` when the rest of the stream is not required to be consumed.
[`parse`](../basic_json/parse.md), [`accept`](../basic_json/accept.md), and all other inputs (strings, iterators,
containers, `#!cpp FILE*`) are never affected.
## Default definition
The default value is `0` (disabled — existing behavior is preserved).
```cpp
#define JSON_PRECISE_STREAM_POSITION 0
```
## Notes
!!! note "Background"
A number is the only JSON value whose end can be detected solely by reading the character that follows it. By
default, that character is consumed and not put back, so the stream is left one byte too far after a number, and
only after a number:
```cpp
std::istringstream input("1true");
json j;
input >> j; // j == 1, but the stream now starts at "rue"
```
With this macro, the character is only looked at and left in the stream, so the stream starts at `true`. This
does not require the stream buffer to support putting a character back.
This was not changed unconditionally, because code can depend on the consumed character, even unknowingly (see
[#5340](https://github.com/nlohmann/json/issues/5340)). Both of the following work by default only because the
character after each number is swallowed, and behave differently with this macro:
```cpp
std::istringstream input("1,2,3");
json j1, j2, j3;
input >> j1 >> j2 >> j3; // default: 1, 2, 3
// with the macro: throws parse_error.101 at the ','
```
```cpp
std::istringstream input("42\nfoo");
json j;
std::string line;
input >> j;
std::getline(input, line); // default: "foo"
// with the macro: "" (like after reading an int with >>)
```
In both cases, the behavior with the macro is what you already get today when the value is not a number: `"a","b"`
fails at the `,`, and `std::getline` after `{}` returns an empty string. This macro offers an opt-in path to
the consistent behavior ahead of version 4.0.0, where it is planned to become the default.
!!! warning "Opt-in only"
This macro must be defined **before** including `<nlohmann/json.hpp>`. Defining it after the include has no
effect.
!!! note "ABI compatibility"
The value of this macro is encoded in the [namespace](../../features/namespace.md) (tag `_psp`), resulting in
distinct symbol names. Translation units compiled with and without it can therefore be linked into the same program
without One Definition Rule (ODR) violations, but they cannot exchange instances of library types.
!!! tip "Workaround without the macro"
Separate the values in the stream with whitespace. The character consumed after a number is then the separator,
and whitespace before the next value is skipped anyway.
## Examples
??? example "Default behavior (macro not defined)"
Without the macro, the character after a number is consumed:
```cpp
#include <iostream>
#include <sstream>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
std::istringstream input("1true");
json j1, j2;
input >> j1; // j1 == 1
input >> j2; // throws parse_error.101: the stream now starts at "rue"
}
```
??? example "Opt-in precise stream position (macro defined to 1)"
With the macro, the stream is positioned right after the number:
```cpp
#define JSON_PRECISE_STREAM_POSITION 1
#include <iostream>
#include <sstream>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
std::istringstream input("1true");
json j1, j2;
input >> j1; // j1 == 1
input >> j2; // j2 == true
}
```
## See also
- [**operator>>**](../operator_gtgt.md) - deserialize from stream
- [**sax_parse**](../basic_json/sax_parse.md) - generate SAX events
## Version history
- Added in version 3.13.0.
- Planned to become the default (with the macro removed) in version 4.0.0.
+34 -16
View File
@@ -33,26 +33,43 @@ A UTF-8 byte order mark is silently ignored.
Invalid Unicode escapes and unpaired surrogates in the input are reported as
[`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) with a detailed message.
`operator>>` parses exactly one JSON value and leaves the stream positioned right after it, so it can be called
repeatedly to read a sequence of concatenated JSON values from the same stream:
`operator>>` parses exactly one JSON value, so it can be called repeatedly to read a sequence of concatenated JSON
values from the same stream:
```cpp
std::istringstream input("1true[2]");
json j1, j2, j3;
input >> j1; // j1 == 1, stream now positioned right after it
input >> j2; // j2 == true
input >> j3; // j3 == [2]
json j1, j2;
input >> j1; // parses the first value
input >> j2; // parses the next value
```
!!! note "Changed behavior for numbers"
!!! warning "A number must be followed by whitespace"
A number is the only value whose end can be detected solely by reading the character that follows it. Up to
version 3.13.0 that character was consumed and not put back, so the stream was left one byte too far whenever a
number was immediately followed by another value: reading `1true` yielded `1` and left the stream at `rue`.
Values had to be separated by whitespace to work around this.
A number is only terminated by the character that follows it. That character is read from the stream to detect the
end of the number, and it is **not** put back. When a value that is a number is immediately followed by the next
value, the first character of that next value is lost:
The terminating character is now only looked at and left in the stream, so no separator is required. Code that
relied on the extra byte being swallowed will observe it again.
```cpp
std::istringstream input("1true");
json j1, j2;
input >> j1; // j1 == 1
input >> j2; // throws parse_error.101: the stream now starts at "rue"
```
Separating the values with whitespace avoids this, because the character that is eaten is then the separator:
```cpp
std::istringstream input("1 true");
json j1, j2;
input >> j1; // j1 == 1
input >> j2; // j2 == true
```
Only numbers are affected. Values ending in a self-delimiting character do not read past themselves, so
`truefalse`, `[1][2]`, `{"a":1}{"b":2}`, and `"a""b"` can be read back to back without a separator.
Define [`JSON_PRECISE_STREAM_POSITION`](macros/json_precise_stream_position.md) to `1` to leave the terminating character in the stream
instead, so that the stream is positioned right after the value for every value type and no separator is
needed. This is tracked in [#5340](https://github.com/nlohmann/json/issues/5340).
Note that reading concatenated values does **not** work for [JSON Lines](../features/parsing/json_lines.md)
(newline-delimited JSON) input -- see that page for why and for the recommended alternative.
@@ -92,11 +109,12 @@ being read.
- [parse](basic_json/parse.md) - deserialize from a compatible input
- [`JSON_STRICT_NUL_HANDLING`](macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input
instead of treating it as end of input
- [`JSON_PRECISE_STREAM_POSITION`](macros/json_precise_stream_position.md) - opt in to leaving the stream positioned right after a number
## Version history
- Added in version 1.0.0.
- Changed in version 4.0.0 to leave the character that terminates a number in the stream, so that the stream is
positioned right after the parsed value for every value type.
- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating
it as end of input; planned to become the default in version 4.0.0.
- `JSON_PRECISE_STREAM_POSITION` added in version 3.13.0 to optionally leave the character that terminates a number in
the stream; planned to become the default in version 4.0.0.
+9
View File
@@ -98,6 +98,15 @@ rather than descending into a bounded number of levels first, which is slower bu
See [full documentation of `JSON_NO_THREAD_LOCAL`](../api/macros/json_no_thread_local.md).
## `JSON_PRECISE_STREAM_POSITION`
When defined to `1`, [`operator>>`](../api/operator_gtgt.md) and non-strict
[`sax_parse`](../api/basic_json/sax_parse.md) leave an input stream positioned right after the parsed value, instead of
also consuming the character that terminates a number. The default value is `0`, which preserves the existing behavior;
this is planned to become the default in version 4.0.0.
See [full documentation of `JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md).
## `JSON_SKIP_LIBRARY_VERSION_CHECK`
When defined, the library will not create a compiler warning when a different version of the library was already
+1
View File
@@ -18,6 +18,7 @@ The complete default namespace name is derived as follows:
- [`JSON_DIAGNOSTIC_POSITIONS`](../api/macros/json_diagnostic_positions.md) defined non-zero appends `_dp`.
- [`JSON_BRACE_INIT_COPY_SEMANTICS`](../api/macros/json_brace_init_copy_semantics.md) defined non-zero appends
`_bics`.
- [`JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md) defined non-zero appends `_psp`.
- The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by
underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component)
below.
+3 -1
View File
@@ -40,7 +40,9 @@ what makes it possible to read several concatenated values from the same stream,
document followed by trailing bytes" is accepted rather than rejected. If you are validating conformance, or need to
reject any input that is not exactly one JSON document, prefer `parse`.
Values read this way do not need to be separated by whitespace; see the
When using `operator>>` to read several concatenated values this way, a value that is a number must be followed by
whitespace, because `operator>>` consumes the character that terminates a number, unless
[`JSON_PRECISE_STREAM_POSITION`](../../api/macros/json_precise_stream_position.md) is defined to `1` — see the
[`operator>>` notes](../../api/operator_gtgt.md#notes) for details and examples.
## SAX vs. DOM parsing
@@ -49,4 +49,5 @@ JSON Lines input with more than one value is treated as invalid JSON by the [`pa
with a JSON Lines input does not work, because the parser will try to parse one value after the last one.
This is different from parsing a stream of *concatenated* (non-newline-delimited) JSON values, for which
`operator>>` does work -- see its [notes](../../api/operator_gtgt.md#notes) for details.
`operator>>` does work, provided that a value that is a number is followed by whitespace -- see its
[notes](../../api/operator_gtgt.md#notes) for details.
+1
View File
@@ -293,6 +293,7 @@ nav:
- 'JSON_NOEXCEPTION': api/macros/json_noexception.md
- 'JSON_NO_IO': api/macros/json_no_io.md
- 'JSON_NO_THREAD_LOCAL': api/macros/json_no_thread_local.md
- 'JSON_PRECISE_STREAM_POSITION': api/macros/json_precise_stream_position.md
- 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md
- 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md
- 'JSON_STRICT_NUL_HANDLING': api/macros/json_strict_nul_handling.md
+15 -4
View File
@@ -38,6 +38,10 @@
#define JSON_BRACE_INIT_COPY_SEMANTICS 0
#endif
#ifndef JSON_PRECISE_STREAM_POSITION
#define JSON_PRECISE_STREAM_POSITION 0
#endif
#if JSON_DIAGNOSTICS
#define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag
#else
@@ -62,21 +66,28 @@
#define NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS
#endif
#if JSON_PRECISE_STREAM_POSITION
#define NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION _psp
#else
#define NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION
#endif
#ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION
#define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0
#endif
// Construct the namespace ABI tags component
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d) json_abi ## a ## b ## c ## d
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d)
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e) json_abi ## a ## b ## c ## d ## e
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e)
#define NLOHMANN_JSON_ABI_TAGS \
NLOHMANN_JSON_ABI_TAGS_CONCAT( \
NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS, \
NLOHMANN_JSON_ABI_TAG_LEGACY_DISCARDED_VALUE_COMPARISON, \
NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS)
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \
NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION)
// Construct the namespace version component
#define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \
@@ -101,9 +101,11 @@ class input_stream_adapter
// maintain ifstream flags, except eof
if (is != nullptr)
{
#if JSON_PRECISE_STREAM_POSITION
// consume the character last returned by get_character() unless it
// was given back with release_lookahead()
commit_lookahead();
#endif
is->clear(is->rdstate() & std::ios::eofbit);
}
}
@@ -117,6 +119,7 @@ class input_stream_adapter
input_stream_adapter& operator=(input_stream_adapter&) = delete;
input_stream_adapter& operator=(input_stream_adapter&&) = delete;
#if JSON_PRECISE_STREAM_POSITION
input_stream_adapter(input_stream_adapter&& rhs) noexcept
: is(rhs.is), sb(rhs.sb), lookahead(rhs.lookahead)
{
@@ -167,11 +170,38 @@ class input_stream_adapter
{
lookahead = false;
}
#else
input_stream_adapter(input_stream_adapter&& rhs) noexcept
: is(rhs.is), sb(rhs.sb)
{
rhs.is = nullptr;
rhs.sb = nullptr;
}
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
// end up as the same value, e.g., 0xFFFFFFFF.
//
// The character is consumed, so the character that terminates a number
// stays consumed after parsing; see JSON_PRECISE_STREAM_POSITION.
std::char_traits<char>::int_type get_character()
{
auto res = sb->sbumpc();
// set eof manually, as we don't use the istream interface.
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
{
is->clear(is->rdstate() | std::ios::eofbit);
}
return res;
}
#endif
template<class T>
std::size_t get_elements(T* dest, std::size_t count = 1)
{
#if JSON_PRECISE_STREAM_POSITION
commit_lookahead();
#endif
auto res = static_cast<std::size_t>(sb->sgetn(reinterpret_cast<char*>(dest), static_cast<std::streamsize>(count * sizeof(T))));
if (JSON_HEDLEY_UNLIKELY(res < count * sizeof(T)))
{
@@ -181,6 +211,7 @@ class input_stream_adapter
}
private:
#if JSON_PRECISE_STREAM_POSITION
// Step over the character last returned by get_character(). The character
// has already been peeked successfully, so for every streambuf with a get
// area this is a pointer increment that cannot fail.
@@ -192,12 +223,15 @@ class input_stream_adapter
sb->sbumpc();
}
}
#endif
/// the associated input stream
std::istream* is = nullptr;
std::streambuf* sb = nullptr;
#if JSON_PRECISE_STREAM_POSITION
/// whether get_character() peeked a character that is not consumed yet
bool lookahead = false;
#endif
};
#endif // JSON_NO_IO
+6 -3
View File
@@ -128,8 +128,9 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
}
// Detect whether an input adapter reads with one character of lookahead that
// can be left in the input (see input_stream_adapter::supports_lookahead),
// detected like supports_seek above.
// can be left in the input (see input_stream_adapter::supports_lookahead,
// which is only defined with JSON_PRECISE_STREAM_POSITION), detected like
// supports_seek above.
template<typename InputAdapterType>
using detect_supports_lookahead = decltype(InputAdapterType::supports_lookahead);
@@ -2011,7 +2012,9 @@ scan_number_done:
right after the value.
Adapters without lookahead (see input_adapter_supports_lookahead) are not
handed back to the user, so this is a no-op for them.
handed back to the user, so this is a no-op for them. Without
JSON_PRECISE_STREAM_POSITION, no adapter has lookahead, so this is always a
no-op and the terminating character stays consumed.
Scanning may continue after this call: @a next_unget is cleared, and the
character is read from the input again instead of being replayed from
@@ -45,6 +45,7 @@
#undef JSON_HAS_STATIC_RTTI
#undef JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
#undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_PRECISE_STREAM_POSITION
#endif
#include <nlohmann/thirdparty/hedley/hedley_undef.hpp>
File diff suppressed because it is too large Load Diff
+56 -7
View File
@@ -95,6 +95,10 @@
#define JSON_BRACE_INIT_COPY_SEMANTICS 0
#endif
#ifndef JSON_PRECISE_STREAM_POSITION
#define JSON_PRECISE_STREAM_POSITION 0
#endif
#if JSON_DIAGNOSTICS
#define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag
#else
@@ -119,21 +123,28 @@
#define NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS
#endif
#if JSON_PRECISE_STREAM_POSITION
#define NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION _psp
#else
#define NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION
#endif
#ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION
#define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0
#endif
// Construct the namespace ABI tags component
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d) json_abi ## a ## b ## c ## d
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d)
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e) json_abi ## a ## b ## c ## d ## e
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e)
#define NLOHMANN_JSON_ABI_TAGS \
NLOHMANN_JSON_ABI_TAGS_CONCAT( \
NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS, \
NLOHMANN_JSON_ABI_TAG_LEGACY_DISCARDED_VALUE_COMPARISON, \
NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS)
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \
NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION)
// Construct the namespace version component
#define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \
@@ -7419,9 +7430,11 @@ class input_stream_adapter
// maintain ifstream flags, except eof
if (is != nullptr)
{
#if JSON_PRECISE_STREAM_POSITION
// consume the character last returned by get_character() unless it
// was given back with release_lookahead()
commit_lookahead();
#endif
is->clear(is->rdstate() & std::ios::eofbit);
}
}
@@ -7435,6 +7448,7 @@ class input_stream_adapter
input_stream_adapter& operator=(input_stream_adapter&) = delete;
input_stream_adapter& operator=(input_stream_adapter&&) = delete;
#if JSON_PRECISE_STREAM_POSITION
input_stream_adapter(input_stream_adapter&& rhs) noexcept
: is(rhs.is), sb(rhs.sb), lookahead(rhs.lookahead)
{
@@ -7485,11 +7499,38 @@ class input_stream_adapter
{
lookahead = false;
}
#else
input_stream_adapter(input_stream_adapter&& rhs) noexcept
: is(rhs.is), sb(rhs.sb)
{
rhs.is = nullptr;
rhs.sb = nullptr;
}
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
// end up as the same value, e.g., 0xFFFFFFFF.
//
// The character is consumed, so the character that terminates a number
// stays consumed after parsing; see JSON_PRECISE_STREAM_POSITION.
std::char_traits<char>::int_type get_character()
{
auto res = sb->sbumpc();
// set eof manually, as we don't use the istream interface.
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
{
is->clear(is->rdstate() | std::ios::eofbit);
}
return res;
}
#endif
template<class T>
std::size_t get_elements(T* dest, std::size_t count = 1)
{
#if JSON_PRECISE_STREAM_POSITION
commit_lookahead();
#endif
auto res = static_cast<std::size_t>(sb->sgetn(reinterpret_cast<char*>(dest), static_cast<std::streamsize>(count * sizeof(T))));
if (JSON_HEDLEY_UNLIKELY(res < count * sizeof(T)))
{
@@ -7499,6 +7540,7 @@ class input_stream_adapter
}
private:
#if JSON_PRECISE_STREAM_POSITION
// Step over the character last returned by get_character(). The character
// has already been peeked successfully, so for every streambuf with a get
// area this is a pointer increment that cannot fail.
@@ -7510,12 +7552,15 @@ class input_stream_adapter
sb->sbumpc();
}
}
#endif
/// the associated input stream
std::istream* is = nullptr;
std::streambuf* sb = nullptr;
#if JSON_PRECISE_STREAM_POSITION
/// whether get_character() peeked a character that is not consumed yet
bool lookahead = false;
#endif
};
#endif // JSON_NO_IO
@@ -8928,8 +8973,9 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
}
// Detect whether an input adapter reads with one character of lookahead that
// can be left in the input (see input_stream_adapter::supports_lookahead),
// detected like supports_seek above.
// can be left in the input (see input_stream_adapter::supports_lookahead,
// which is only defined with JSON_PRECISE_STREAM_POSITION), detected like
// supports_seek above.
template<typename InputAdapterType>
using detect_supports_lookahead = decltype(InputAdapterType::supports_lookahead);
@@ -10811,7 +10857,9 @@ scan_number_done:
right after the value.
Adapters without lookahead (see input_adapter_supports_lookahead) are not
handed back to the user, so this is a no-op for them.
handed back to the user, so this is a no-op for them. Without
JSON_PRECISE_STREAM_POSITION, no adapter has lookahead, so this is always a
no-op and the terminating character stays consumed.
Scanning may continue after this call: @a next_unget is cleared, and the
character is read from the input again instead of being replayed from
@@ -30904,6 +30952,7 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#undef JSON_HAS_STATIC_RTTI
#undef JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
#undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_PRECISE_STREAM_POSITION
#endif
// #include <nlohmann/thirdparty/hedley/hedley_undef.hpp>
+15 -4
View File
@@ -56,6 +56,10 @@
#define JSON_BRACE_INIT_COPY_SEMANTICS 0
#endif
#ifndef JSON_PRECISE_STREAM_POSITION
#define JSON_PRECISE_STREAM_POSITION 0
#endif
#if JSON_DIAGNOSTICS
#define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag
#else
@@ -80,21 +84,28 @@
#define NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS
#endif
#if JSON_PRECISE_STREAM_POSITION
#define NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION _psp
#else
#define NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION
#endif
#ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION
#define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0
#endif
// Construct the namespace ABI tags component
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d) json_abi ## a ## b ## c ## d
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d)
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e) json_abi ## a ## b ## c ## d ## e
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e)
#define NLOHMANN_JSON_ABI_TAGS \
NLOHMANN_JSON_ABI_TAGS_CONCAT( \
NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS, \
NLOHMANN_JSON_ABI_TAG_LEGACY_DISCARDED_VALUE_COMPARISON, \
NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS)
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \
NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION)
// Construct the namespace version component
#define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \
+4
View File
@@ -36,6 +36,10 @@ TEST_CASE("default namespace")
expected += "_bics";
#endif
#if JSON_PRECISE_STREAM_POSITION
expected += "_psp";
#endif
expected += "_v" STRINGIZE(NLOHMANN_JSON_VERSION_MAJOR);
expected += "_" STRINGIZE(NLOHMANN_JSON_VERSION_MINOR);
expected += "_" STRINGIZE(NLOHMANN_JSON_VERSION_PATCH) "::basic_json";
+4
View File
@@ -37,6 +37,10 @@ TEST_CASE("default namespace without version component")
expected += "_bics";
#endif
#if JSON_PRECISE_STREAM_POSITION
expected += "_psp";
#endif
expected += "::basic_json";
// fallback for Clang
+35 -156
View File
@@ -22,15 +22,11 @@ using nlohmann::json;
using namespace nlohmann::literals; // NOLINT(google-build-using-namespace)
#endif
#include <cstddef>
#include <iostream>
#include <iterator>
#include <sstream>
#include <streambuf>
#include <string>
#include <utility>
#include <valarray>
#include <vector>
#if defined(_WIN32)
#define NOMINMAX
@@ -232,58 +228,6 @@ class proxy_iterator
iterator* m_it = nullptr;
};
// A streambuf that keeps no get area at all and refuses every putback: with an
// empty get area, sungetc() always ends up in pbackfail(). Used to check that
// the character terminating a number is left in the input without relying on
// the streambuf being able to put a consumed character back.
class no_putback_streambuf : public std::streambuf
{
public:
explicit no_putback_streambuf(std::string s) : m_data(std::move(s)) {}
protected:
// peek at the next character without consuming it
int_type underflow() override
{
if (m_pos >= m_data.size())
{
return traits_type::eof();
}
return traits_type::to_int_type(m_data[m_pos]);
}
// consume the next character
int_type uflow() override
{
if (m_pos >= m_data.size())
{
return traits_type::eof();
}
return traits_type::to_int_type(m_data[m_pos++]);
}
int_type pbackfail(int_type /*c*/) override
{
return traits_type::eof();
}
private:
std::string m_data;
std::size_t m_pos = 0;
};
// read the characters that are left in a stream
std::string remaining(std::istream& is)
{
std::string result;
char c = 0;
while (is.get(c))
{
result += c;
}
return result;
}
// JSON_HAS_CPP_20
#if defined(__cpp_char8_t)
bool check_utf8()
@@ -1290,119 +1234,54 @@ TEST_CASE("deserialization")
}
}
SECTION("stream position after extraction (#5340)")
SECTION("stream position after extraction without JSON_PRECISE_STREAM_POSITION (#5340)")
{
SECTION("a number does not consume the character that terminates it")
// By default, the character that terminates a number is consumed, so
// the stream is left one byte too far after a number (and only after a
// number). JSON_PRECISE_STREAM_POSITION changes this; see
// unit-precise-stream-position.cpp. These checks pin the default.
const auto remaining = [](std::istream & is)
{
// a number is only terminated by the character following it; that
// character must be given back so the stream is positioned right
// after the value
const std::vector<std::pair<std::string, std::string>> tests =
{
{"1true", "true"},
{"1[2]", "[2]"},
{"1{}", "{}"},
{R"(1"a")", R"("a")"},
{"1 true", " true"},
{"12,", ","},
{"-0.5e3x", "x"},
{"1null", "null"}
};
return std::string(std::istreambuf_iterator<char>(is), std::istreambuf_iterator<char>());
};
for (const auto& test : tests)
{
CAPTURE(test.first);
std::istringstream ss(test.first);
json j;
ss >> j;
CHECK(j == json::parse(test.first.substr(0, test.first.size() - test.second.size())));
CHECK(remaining(ss) == test.second);
}
}
SECTION("values that are self-delimiting are unaffected")
{
const std::vector<std::pair<std::string, std::string>> tests =
{
{"truefalse", "false"},
{"[1][2]", "[2]"},
{R"({"a":1}{"b":2})", R"({"b":2})"},
{R"("a""b")", R"("b")"},
{"null null", " null"}
};
for (const auto& test : tests)
{
CAPTURE(test.first);
std::istringstream ss(test.first);
json j;
ss >> j;
CHECK(remaining(ss) == test.second);
}
}
SECTION("a number at the end of the input leaves nothing behind")
{
for (const std::string s :
{"1", "12", "-3.5e2", " 7 "
})
{
CAPTURE(s);
std::istringstream ss(s);
json j;
ss >> j;
CHECK(remaining(ss).find_first_not_of(" \t\n\r") == std::string::npos);
}
}
SECTION("repeated extraction of concatenated values")
{
std::istringstream ss(R"(1true[2]3"x"{"a":4}5)");
const std::vector<json> expected =
{
json(1), json(true), json::parse("[2]"), json(3),
json("x"), json::parse(R"({"a":4})"), json(5)
};
for (const auto& e : expected)
{
json j;
ss >> j;
CHECK(j == e);
}
}
SECTION("sax_parse with strict == false")
SECTION("the character after a number is consumed")
{
std::istringstream ss("1true");
SaxEventLogger l;
CHECK(json::sax_parse(ss, &l, nlohmann::detail::input_format_t::json, false));
CHECK(l.events.size() == 1);
CHECK(l.events[0] == "number_unsigned(1)");
json j;
ss >> j;
CHECK(j == 1);
CHECK(remaining(ss) == "rue");
}
SECTION("the character after other values is not consumed")
{
std::istringstream ss("[1]true");
json j;
ss >> j;
CHECK(j == json::parse("[1]"));
CHECK(remaining(ss) == "true");
}
SECTION("strict parsing still rejects trailing data")
SECTION("comma-separated numbers can be read one by one")
{
std::istringstream ss("1true");
json _;
CHECK_THROWS_WITH_AS(_ = json::parse(ss),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - unexpected true literal; expected end of input", json::parse_error&);
std::istringstream ss2("1true");
CHECK_FALSE(json::accept(ss2));
std::istringstream ss("1,2,3");
json j1, j2, j3;
ss >> j1 >> j2 >> j3;
CHECK(j1 == 1);
CHECK(j2 == 2);
CHECK(j3 == 3);
}
SECTION("a streambuf that cannot put back is not needed")
SECTION("std::getline after a number skips the line break")
{
// the terminating character is never consumed, so no putback
// position is required
no_putback_streambuf buf("1true");
std::istream is(&buf);
std::istringstream ss("42\nfoo");
json j;
is >> j;
CHECK(j == json(1));
CHECK(remaining(is) == "true");
std::string line;
ss >> j;
std::getline(ss, line);
CHECK(j == 42);
CHECK(line == "foo");
}
}
+237
View File
@@ -0,0 +1,237 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// This file tests the opt-in JSON_PRECISE_STREAM_POSITION, so it defines the
// macro itself rather than relying on a -D flag, and runs in every build. The
// default behavior is pinned in unit-deserialization.cpp.
#ifdef JSON_PRECISE_STREAM_POSITION
#undef JSON_PRECISE_STREAM_POSITION
#endif
#define JSON_PRECISE_STREAM_POSITION 1
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <cstddef>
#include <sstream>
#include <streambuf>
#include <string>
#include <utility>
#include <vector>
#define STRINGIZE_EX(x) #x
#define STRINGIZE(x) STRINGIZE_EX(x)
namespace
{
// A streambuf that keeps no get area at all and refuses every putback: with an
// empty get area, sungetc() always ends up in pbackfail(). Used to check that
// the character terminating a number is left in the input without relying on
// the streambuf being able to put a consumed character back.
class no_putback_streambuf : public std::streambuf
{
public:
explicit no_putback_streambuf(std::string s) : m_data(std::move(s)) {}
protected:
// peek at the next character without consuming it
int_type underflow() override
{
if (m_pos >= m_data.size())
{
return traits_type::eof();
}
return traits_type::to_int_type(m_data[m_pos]);
}
// consume the next character
int_type uflow() override
{
if (m_pos >= m_data.size())
{
return traits_type::eof();
}
return traits_type::to_int_type(m_data[m_pos++]);
}
int_type pbackfail(int_type /*c*/) override
{
return traits_type::eof();
}
private:
std::string m_data;
std::size_t m_pos = 0;
};
// read the characters that are left in a stream
std::string remaining(std::istream& is)
{
std::string result;
char c = 0;
while (is.get(c))
{
result += c;
}
return result;
}
} // namespace
TEST_CASE("JSON_PRECISE_STREAM_POSITION")
{
SECTION("the macro is part of the ABI tag")
{
const std::string ns = STRINGIZE(NLOHMANN_JSON_NAMESPACE);
// other tags may come before it, e.g. json_abi_diag_psp
CHECK(ns.find("_psp") != std::string::npos);
}
SECTION("a number does not consume the character that terminates it")
{
// a number is only terminated by the character following it; that
// character must be given back so the stream is positioned right
// after the value
const std::vector<std::pair<std::string, std::string>> tests =
{
{"1true", "true"},
{"1[2]", "[2]"},
{"1{}", "{}"},
{R"(1"a")", R"("a")"},
{"1 true", " true"},
{"12,", ","},
{"-0.5e3x", "x"},
{"1null", "null"}
};
for (const auto& test : tests)
{
CAPTURE(test.first);
std::istringstream ss(test.first);
json j;
ss >> j;
CHECK(j == json::parse(test.first.substr(0, test.first.size() - test.second.size())));
CHECK(remaining(ss) == test.second);
}
}
SECTION("values that are self-delimiting are unaffected")
{
const std::vector<std::pair<std::string, std::string>> tests =
{
{"truefalse", "false"},
{"[1][2]", "[2]"},
{R"({"a":1}{"b":2})", R"({"b":2})"},
{R"("a""b")", R"("b")"},
{"null null", " null"}
};
for (const auto& test : tests)
{
CAPTURE(test.first);
std::istringstream ss(test.first);
json j;
ss >> j;
CHECK(remaining(ss) == test.second);
}
}
SECTION("a number at the end of the input leaves nothing behind")
{
for (const std::string s :
{"1", "12", "-3.5e2", " 7 "
})
{
CAPTURE(s);
std::istringstream ss(s);
json j;
ss >> j;
CHECK(remaining(ss).find_first_not_of(" \t\n\r") == std::string::npos);
}
}
SECTION("repeated extraction of concatenated values")
{
std::istringstream ss(R"(1true[2]3"x"{"a":4}5)");
const std::vector<json> expected =
{
json(1), json(true), json::parse("[2]"), json(3),
json("x"), json::parse(R"({"a":4})"), json(5)
};
for (const auto& e : expected)
{
json j;
ss >> j;
CHECK(j == e);
}
}
SECTION("differences to the default behavior")
{
// both of these work by accident without the macro, because the
// character after a number is swallowed; see unit-deserialization.cpp
SECTION("a separator after a number is not skipped")
{
std::istringstream ss("1,2");
json j;
ss >> j;
CHECK(j == 1);
CHECK_THROWS_AS(ss >> j, json::parse_error&);
}
SECTION("std::getline after a number sees the line break")
{
std::istringstream ss("42\nfoo");
json j;
std::string line;
ss >> j;
std::getline(ss, line);
CHECK(j == 42);
CHECK(line.empty());
std::getline(ss, line);
CHECK(line == "foo");
}
}
SECTION("sax_parse with strict == false")
{
std::istringstream ss("1true");
json j;
nlohmann::detail::json_sax_dom_parser<json, nlohmann::detail::input_stream_adapter> sdp(j, true);
CHECK(json::sax_parse(ss, &sdp, nlohmann::detail::input_format_t::json, false));
CHECK(j == 1);
CHECK(remaining(ss) == "true");
}
SECTION("strict parsing still rejects trailing data")
{
std::istringstream ss("1true");
json _;
CHECK_THROWS_WITH_AS(_ = json::parse(ss),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - unexpected true literal; expected end of input", json::parse_error&);
std::istringstream ss2("1true");
CHECK_FALSE(json::accept(ss2));
}
SECTION("a streambuf that cannot put back is not needed")
{
// the terminating character is never consumed, so no putback
// position is required
no_putback_streambuf buf("1true");
std::istream is(&buf);
json j;
is >> j;
CHECK(j == json(1));
CHECK(remaining(is) == "true");
}
}
+1 -1
View File
@@ -20,7 +20,7 @@ if __name__ == '__main__':
namespaces = ['nlohmann']
abi_prefix = 'json_abi'
abi_tags = ['_diag', '_ldvcmp', '_dp', '_bics']
abi_tags = ['_diag', '_ldvcmp', '_dp', '_bics', '_psp']
version = '_v' + args.version.replace('.', '_')
inline_namespaces = []