basic_json::swap() (and the friend swap() that forwards to it) only
exchanged m_data.m_type/m_data.m_value, leaving start_position/end_position
untouched under JSON_DIAGNOSTIC_POSITIONS. This is inconsistent with
copy-assignment's operator=(basic_json), which swaps positions as part of
its copy-and-swap implementation, so after swap(a, b) each value ended up
with the other value's content but its own original position.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When a parser callback rejects a value, the placeholder stored for it has to
be removed from its parent again. remove_discarded_value() found it by
scanning the parent from the beginning, so filtering a container cost one
scan per rejected member - quadratic in the number of members of a single
container.
A rejected value can only ever be the one most recently added to its parent:
the last element of an array, or the placeholder key() stored under the
current key in an object. Record that key alongside the existing
key_keep_stack, and for a container record it again alongside ref_stack so
end_object()/end_array() can find it in the parent. Removal is then O(1) for
an array and O(log n) for an object, and finding nothing there means nothing
was stored, so there is nothing to remove.
The key for a container is read before handle_value() may consume it, so it
is also correct when the callback rejects the container at its start event
and it never reaches its parent at all.
Discarding half the members of one object, before -> after:
members value rejected container rejected at start
16 000 392 ms -> 3.7 ms 803 ms -> 8.2 ms
64 000 6238 ms -> 14.6 ms 12651 ms -> 32.2 ms
128 000 25339 ms -> 30.6 ms
Results are unchanged: 48 000 randomized documents parsed under 12 different
filtering callbacks - covering duplicate keys, empty keys, rejected keys and
containers rejected at both their start and end events - produce byte
identical output before and after.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
tests/src/unit-class_parser_diagnostic_positions.cpp and
tests/src/unit-diagnostic-positions-only.cpp were maintained as near-copies
of unit-class_parser.cpp and unit-diagnostic-positions.cpp respectively, and
had drifted: trailing-comma handling, the #5342 filter-array/filter-value
sections, and the cross-input-adapter diagnostics test were never ported to
the positions-enabled copy.
Fold the position-specific assertions into the base files, guarded by
file a second time with the relevant macro set via CMake COMPILE_DEFINITIONS
(mirroring the existing test-comparison_legacy pattern) instead of
maintaining a separate source file. This removes the duplication and, as a
side effect, closes the coverage gaps above since the full test file now
compiles under JSON_DIAGNOSTIC_POSITIONS=1 as well.
Fixes#5417
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Broaden JSON_HEDLEY_WARN_UNUSED_RESULT coverage to pure query functions
Add JSON_HEDLEY_WARN_UNUSED_RESULT to the unambiguous, const,
side-effect-free observer functions whose return value is the entire
purpose of the call:
- dump()
- type(), type_name()
- all is_* predicates (is_primitive, is_structured, is_null,
is_boolean, is_number, is_number_integer, is_number_unsigned,
is_number_float, is_object, is_array, is_string, is_binary,
is_discarded)
- empty(), size(), max_size()
- count(...) (both overloads) and contains(...) (all overloads,
including the deprecated json_pointer<BasicJsonType> overload)
This mirrors the direction the standard library has taken with
[[nodiscard]] on the analogous std::vector/std::map members, and
catches real bugs such as `j.empty();` (meant `j.clear();`) or
`j.contains(k);` with the result thrown away.
Deliberately out of scope (left for a separate, later policy
decision, per the issue): at(), value(), get*(), flatten(),
unflatten(), patch(), merge_patch(), begin()/end(), comparison
operators, erase(), and emplace().
Compiling the full test suite (tests/src/unit-*.cpp) with
-Wunused-result -Werror uncovered one real hit: a regression test in
unit-regression2.cpp called dump() purely to check it does not throw,
discarding the result. Fixed by explicitly casting to void, since the
call is intentionally result-less there.
Fixes#5410
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix discarded nodiscard results across the test suite for GCC's warn_unused_result
A plain (void) cast on a call expression suppresses the C++17 [[nodiscard]]
warning but not GCC's warning for functions annotated via the GNU
__attribute__((warn_unused_result)) form -- which is what
JSON_HEDLEY_WARN_UNUSED_RESULT expands to on GCC. Several existing tests
that call a newly-annotated function (dump(), empty()) purely to check
that it throws/does not throw, discarding the result via (void), newly
warned (and failed -Werror builds) once the annotation was broadened.
Route those discards through a small ignore_return_value() helper
instead, which actually consumes the value and suppresses the warning
on both attribute forms.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use utils::ignore_return_value() for the issue #1445 dump() discard too
Addresses review feedback from @gregmarr on PR #5477: this call site was
still using the older "capture in a variable, then (void) it" pattern
from before this PR introduced utils::ignore_return_value(), instead of
the helper now used at every other discarded-nodiscard-result call site
this PR touches.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Speed up whitespace skipping in the lexer
lexer::skip_whitespace() called get() for every whitespace byte, and
get() checks the (almost always false, once past the first character)
next_unget flag on every call. skip_whitespace() now reads its first
character with get() (needed to honor a pending unget() left over from
finishing the previous token, e.g. scan_number() always ungets the
character that terminated the number) and every further whitespace
character with a new get_ignoring_pending_unget() variant that skips
that branch, since nothing in the loop calls unget().
This is a narrower fix than the full contiguous-buffer bulk-skip
suggested in the issue (scan a run of whitespace directly in the
adapter's buffer and update position counters once per run). That
approach depends on bulk-scan adapter infrastructure
(supports_bulk_scan/bulk_data()/bulk_skip()) introduced by the open,
unmerged parser-performance PR #5283, which this change intentionally
does not depend on or replicate. Building new bulk-scan adapter
infrastructure from scratch was judged out of scope/riskier than
warranted here, so this change is limited to the safe, always-correct
improvement of removing redundant per-character bookkeeping from the
existing byte-at-a-time loop; full bulk-skipping is left as future
work once #5283 (or equivalent adapter support) lands.
Line/column/byte-offset bookkeeping is untouched and verified
bit-for-bit identical before and after this change, including for
pretty-printed (dump(4)) input with embedded newlines.
Fixes#5412
Stacked on top of the PR for #5411 (branch
issue-5411-lexer-skip-conversion).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix codegen regression in skip_whitespace() from #5490
Benchmarking found the get()/get_ignoring_pending_unget() split in
skip_whitespace() made long whitespace runs (e.g. indentation in
pretty-printed JSON) 1.75x-3.2x SLOWER instead of faster, reproducible
with both Apple Clang and GCC.
Root cause: rewriting the loop from a plain do-while into an initial
get() followed by a while-loop defeated the compiler's ability to keep
the input adapter's read/end pointers in registers across iterations;
both compilers instead reloaded them from memory on every character.
The function split itself was not the problem (it still fully
inlines); the loop's control-flow shape was.
The fix keeps the same two-function structure but restores a
do-while shape (guarded by an if for the "first char not whitespace"
case), which lets both compilers hoist the pointers back into
registers, matching or beating pre-#5490 performance.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Share the position-counter bump between get() and get_ignoring_pending_unget()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Extract current_is_whitespace() to deduplicate skip_whitespace()'s two whitespace checks
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use a raw string literal for the multi-line error-position test input
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy raw-string-literal finding and guard a new test against JSON_NOEXCEPTION
The issue #5412 whitespace-skipping test added a check_error() helper that
relies on catching json::parse_error to verify the exception message; under
JSON_NOEXCEPTION, JSON_THROW aborts instead of throwing, which crashed
ci_test_noexceptions (and cascaded into the other ci_cmake_options jobs).
Guard the whole section with #if !defined(JSON_NOEXCEPTION), matching the
existing pattern used by sibling tests in this file.
Also switch one escaped string literal to a raw string literal to satisfy
clang-tidy's modernize-raw-string-literal check.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip integer conversion in accept()/SAX validation when the value is unused
lexer::scan_number() always converted every numeric token with
strtoull()/strtoll() before returning, even though accept() (and any
consumer using json_sax_acceptor) immediately discards the converted
value. For value_unsigned/value_integer tokens whose digit count
already guarantees the value fits into 64 bits, the conversion cannot
change the accept/reject decision (such tokens are always finite and
unconditionally accepted), so scan_number() can skip strtoull()/
strtoll() entirely in that case when the caller signals it does not
need the value. Numbers with more digits keep using the exact,
unmodified conversion path, so overflow reclassification to
value_float (and the finiteness check on it) is unaffected.
parse() and value_float handling are completely unchanged.
Fixes#5411
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document why the digit-count fast path is safe regardless of number_unsigned_t/number_integer_t width
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Avoid temporary-string concatenation flagged by clang-tidy in the differential test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
`lexer::get()` copied every scanned character into `token_string` on the
whole successful-parse hot path, yet that buffer is consumed only by
`get_token_string()` when rendering the "last read" fragment of a parse
error. On well-formed input the per-byte copy (plus the `unget()` pop)
is pure overhead that is always discarded.
For seekable input adapters - random-access, single-byte iterators such
as those backing `std::string`, `const char*`, and `std::vector<char>` -
the offending token is now reconstructed on demand from the input when
an error is reported, using a saved start offset, and the eager copy is
skipped. Streaming adapters (file, istream, wide-string, and user-defined
adapters) keep the eager copy; the strategy is chosen at compile time via
`input_adapter_supports_seek`, so adapters without the capability are
unaffected.
Error messages are byte-for-byte identical across all adapters, verified
by a new parity regression test. Microbenchmark (4 MB mixed JSON, parsed
from a std::string): ~149 -> ~160 MB/s, about +8%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
PR #4873 introduced a safety check in sax_parse functions to catch
nullptr passed as SAX parser object, which had been already annotated by
JSON_HEDLEY_NON_NULL macro.
Compilers (e.g. clang) which respected the non-null annotation tended to
eliminate the safety check completely in optimized builds, while
compilers which did not, compiled the safety check in. This led to
different behaviors accross different compilers/platforms and/or build
types (debug, release).
This commit reverts PR #4873 to remove this discrepancy. Passing null to
non-null annotated parameter is considered to be undefined behavior.
Fixes#5048
Signed-off-by: Richard Musil <risa2000x@gmail.com>
Co-authored-by: Richard Musil <risa2000x@gmail.com>
* Add implementation to retrieve start and end positions of json during parse
* Add more unit tests and add start/stop parsing for arrays
* Add raw value for all types
* Add more tests and fix compiler warning
* Amalgamate
* Fix CLang GCC warnings
* Fix error in build
* Style using astyle 3.1
* Fix whitespace changes
* revert
* more whitespace reverts
* Address PR comments
* Fix failing issues
* More whitespace reverts
* Address remaining PR comments
* Address comments
* Switch to using custom base class instead of default basic_json
* Adding a basic using for a json using the new base class. Also address PR comments and fix CI failures
* Address decltype comments
* Diagnostic positions macro (#4)
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Fix missed include deletion
* Add docs and address other PR comments (#5)
* Add docs and address other PR comments
---------
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Address new PR comments and fix CI tests for documentation
* Update documentation based on feedback (#6)
---------
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Address std::size_t and other comments
* Fix new CI issues
* Fix lcov
* Improve lcov case with update to handle_diagnostic_positions call for discarded values
* Fix indentation of LCOV_EXCL_STOP comments
* fix amalgamation astyle issue
---------
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Move UDLs into nlohmann::literals::json_literals namespace
* Add 'using namespace' to unit tests
* Add 'using namespace' to examples
* Add 'using namespace' to README
* Move UDL mkdocs pages out of basic_json/
* Update documentation
* Update docset index
* Add JSON_GlobalUDLs CMake option
* Add unit test
* Build examples without global UDLs
* Add CI target
* Add C++20 3-way comparison operator and fix broken comparisons
Fixes#3207.
Fixes#3409.
* Fix iterators to meet (more) std::ranges requirements
Fixes#3130.
Related discussion: #3408
* Add note about CMake standard version selection to unit tests
Document how CMake chooses which C++ standard version to use when
building tests.
* Update documentation
* CI: add legacy discarded value comparison
* Fix internal linkage errors when building a module