`std::from_chars` is locale-independent and thus avoids the TOCTOU
issue #5198 when it's available.
This commit introduces a `from_chars` utility in the `detail` namespace
which merely wraps `std::from_chars` if it is available and emulates its
API and behavior using the `strto*` family of functions when it is not.
As of this commit, the `strto*`-based fallback is still locale-dependent
and only matches `std::from_chars` when the "C" locale is used.
Availability of `std::from_chars` for floating point numbers is varied.
For instance, as of libc++ 22, `std::from_chars` for `long double` isn't
supported and, consequently, the feature-test macro `__cpp_lib_to_chars`
is not set, while the overloads for `float` and `double` are available.
To avoid complicating matters further, the feature-test macro is taken
as the authoritative indicator for (full) `std::from_chars` support and
the fallback is used otherwise.
If `std::from_chars` is available, our wrapper additionally tweaks its
output in the case of floating-point out-of-range results. The behavior
in this case is poorly (under-)specified and existing implementations do
not agree. The wrapper's tweaks are intended to ensure consistent
behavior across all toolchains in this edge case. In practice, this
should only take effect when parsing out-of-range floats on libstdc++.
The unit tests for the `from_chars` helper are intentionally compiled in
both C++11 and C++17 modes. This ensures that the test expectations
align with the actual behavior we're trying to model. The tests are not
aiming to cover all of floating-point parsing comprehensively, as
`std::from_chars` is specified to behave mostly the same as `strto*`
and instead focuses on the edge cases where the two differ:
- leading whitespace is not tolerated
- `+` sign is not allowed (outside of the exponent)
- out-parameter is left unchanged in case of `invalid_argument` and
integral `result_out_of_range` errors
Signed-off-by: Jonas Greitemann <jgreitemann@gmail.com>
`lexer::get()` copied every scanned character into `token_string` on the
whole successful-parse hot path, yet that buffer is consumed only by
`get_token_string()` when rendering the "last read" fragment of a parse
error. On well-formed input the per-byte copy (plus the `unget()` pop)
is pure overhead that is always discarded.
For seekable input adapters - random-access, single-byte iterators such
as those backing `std::string`, `const char*`, and `std::vector<char>` -
the offending token is now reconstructed on demand from the input when
an error is reported, using a saved start offset, and the eager copy is
skipped. Streaming adapters (file, istream, wide-string, and user-defined
adapters) keep the eager copy; the strategy is chosen at compile time via
`input_adapter_supports_seek`, so adapters without the capability are
unaffected.
Error messages are byte-for-byte identical across all adapters, verified
by a new parity regression test. Microbenchmark (4 MB mixed JSON, parsed
from a std::string): ~149 -> ~160 MB/s, about +8%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix: integer parsed as float when EINTR set in errno
* chore: make amalgamate
* chore: make pretty
---------
Co-authored-by: Stuart Gorman <Stuart.Gorman@kallipr.com>
* Add versioned inline namespace
Add a versioned inline namespace to prevent ABI issues when linking code
using multiple library versions.
* Add namespace macros
* Encode ABI information in inline namespace
Add _diag suffix to inline namespace if JSON_DIAGNOSTICS is enabled, and
_ldvcmp suffix if JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON is enabled.
* Move ABI-affecting macros into abi_macros.hpp
* Move std_fs namespace definition into std_fs.hpp
* Remove std_fs namespace from unit test
* Format more files in tests directory
* Add unit tests
* Update documentation
* Fix GDB pretty printer
* fixup! Add namespace macros
* Derive ABI prefix from NLOHMANN_JSON_VERSION_*
I tested following strings with invalid surrogate pair and unpaired surrogate in files:
1. `"a\uD800\uD800x"`
2. `"a\uD800x"`
The error messge was: "... invalid string: surrogate U+DC00..U+DFFF must be followed by U+DC00..U+DFFF; ..."
I think it must be: "... invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; ..."