get_ubjson_array() and get_ubjson_object() read their elements by calling
back into the value reader, which called them again for a nested container,
so the native call stack grew with the nesting depth of the input. '[' alone
opens a container, so half a million of them crashes the process before the
input runs out (#5104). The optimized forms reach the same path through a
size or type annotation, and in plain UBJSON '[' and '{' are permitted as the
type of an optimized container, so "[$[#i\x01" repeated nests just as deeply
at six bytes a level.
Both readers now only open their container, and parse_ubjson_internal() loops:
it closes the containers that have ended, claims the next element of the
innermost one, reads its key when it is an object, and works out the marker
of the value to read next. That last part is where the formats differ, and
the loop follows what the four element loops used to do:
- a sized, typed container gives its elements no marker of their own
- a sized, untyped container reads one for each element
- a container that ends at a marker has the byte already, from the test
against ']' or '}'; for an object it is the first byte of the key
The ND-array wrapper and the 'B' binary shortcut stay as they are. Both read
a complete value rather than opening a container, and their elements are
always scalars: BJData does not permit '[' or '{' as an optimized type, which
is also why only plain UBJSON needed the type-marker case above.
A container of no-ops keeps its behaviour of holding no elements while still
announcing its declared size to the SAX parser, by opening it and then
setting its count to zero.
unit-ubjson and unit-bjdata pass unchanged, 1.39 million assertions between
them, and a behaviour comparison against the previous commit over every
container form -- sized, unsized, typed, untyped, empty, no-op, ND-array,
binary, and the forms nested inside one another -- gives identical values,
error codes, messages and byte offsets. 500,000 levels of each vector now
report a parse error instead of crashing, and a well-formed 100,000-level
value is read to completion.
The driver costs about 3 % on parsing 60,000 small objects and one array of a
million integers, for the reason given in the previous commit; reading the
frame once per element rather than per branch halved what it cost before.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized UBJSON array of one of those has no
payload: reading an element consumes no input at all. Its declared count is
therefore the only thing that decides how much is allocated, and nothing
bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing
two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and
OSS-Fuzz has an out-of-memory and a timeout report for the same shape.
Every other type costs at least one byte per element, so the end of the input
bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not
affected either: each element is preceded by its key, which costs bytes. And
BJData already refuses these markers as an optimized type, so this is a plain
UBJSON matter.
Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it
will not honour. The check runs before the SAX start event, so no container
is opened and then abandoned.
Rejecting on the read side alone would break the guarantee that anything
to_ubjson() writes can be read back, and would trip the round-trip assertion
in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized
encoding, one byte per element, for arrays of these types above the same
limit. Its decision depends only on the array's size, which is identical for
a value and for anything parsed back from it, so the round trip is stable.
No existing test changes: the largest such count in the test suite is 65,793.
The excessive-size test that already used this shape still passes, now
rejected a little earlier than by the max_size() check it used to reach.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Throw other_error.502 when UBJSON use_type is set without use_size
Fixes#5321
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Scope UBJSON use_type check to container branches and expand tests
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Re-amalgamate single_include/json.hpp
The previous commit updated the split headers but the amalgamated
file didn't go back through astyle before I committed it, so CI's
amalgamation check caught formatting drift in json_fwd.hpp and a
few noexcept clauses in basic_json, plus one doc example. None of
it touches the UBJSON logic. Applied the patch CI generated to
bring single_include back in sync.
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
---------
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
The comment asked whether the no-op marker 'N' may be ignored when a
string is read. It may not: at that point the next byte must be a string
length type specification, and 'N' is not one. No-ops at positions where
a value may start are already consumed by the callers through
get_ignore_noop(), so nothing is lost by not skipping them here.
Replace the TODO with a comment stating that, and add regression tests
pinning both directions: a no-op is accepted at top level (also
repeated), before and after an array element, and before an object key,
between key and value, and before the closing brace of an object of
unknown size; it is rejected where a length type specification is
expected, i.e. after the 'S' marker of a string value and as the key
length of an object of known size.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This adds a std::isfinite check to the UBJSON floating-point parsing path, throwing out_of_range.406 on overflow. This makes the UBJSON parser's behavior consistent with the normal JSON parser. Fixes#5322.
Signed-off-by: AJ369ninja <abhishek.j@iitg.ac.in>
Co-authored-by: AJ369ninja <abhishek.j@iitg.ac.in>
* Add iterator+sentinel tests and docs for binary deserializers
This commit extends the C++20 ranges support (iterator+sentinel pairs) to the
binary format deserializers from_cbor, from_msgpack, from_ubjson, from_bjdata,
and from_bson, matching what was already done for parse(), accept(), and
sax_parse().
Changes:
- Add istreambuf_sentinel helper to test_utils.hpp for EOF detection in tests
- Add 5 new test cases that read binary files directly via
std::istreambuf_iterator<char> + sentinel, without pre-buffering
- Update documentation for all 5 from_* functions to document overload (3)
with SentinelType parameter
- All tests pass; verified against existing test suite data
- Fix potential buffer over-read warning in heterogeneous iterator test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge iterator+sentinel overloads and fix ambiguity/CI issues
Address PR review feedback and CI failures:
- Merge the separate same-type and sentinel-type iterator overloads of
parse(), accept(), sax_parse(), and the five from_* binary deserializers
into a single overload with SentinelType defaulted to IteratorType,
as suggested in review. Applied the same simplification to the
detail::input_adapter() free functions.
- Fix a latent ambiguity: some compilers (e.g. GCC 4.8) unreliably SFINAE
the operator!= detection for std::nullptr_t against container/string
types, making calls like parse(s, nullptr, ...) ambiguous with the
compatible-input overload. can_compare_ne now explicitly excludes
std::nullptr_t as a SentinelType.
- Use a named enable_if_t template parameter instead of an unnamed
function parameter for the SFINAE guard, fixing a clang-tidy
hicpp-named-parameter/readability-named-parameter failure.
- Update parse.md, accept.md, sax_parse.md, and the five from_*.md pages
to document the merged overload instead of separate (2)/(3) overloads,
also fixing an over-160-char line that broke the documentation
style_check CI job.
- Rework the BSON iterator+sentinel test to parse a BSON file already
present in the test suite instead of writing/deleting a temp file.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Wunneeded-internal-declaration for CustomSentinel in test
CustomSentinel lives in an anonymous namespace (internal linkage), and
the library's parse loop only ever evaluates the iterator-first
direction (it != last), so the reversed-order friend operator!= was
never referenced. Clang's -Weverything flags such unused internal
declarations as an error. Drop the unused overload; the used direction
is enough to satisfy can_compare_ne's either-order detection.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy hicpp-named-parameter and misc-const-correctness
- Drop the unused reversed-order operator!= overload from
utils::istreambuf_sentinel (only iterator != sentinel is ever
evaluated) and name the remaining friend's sentinel parameter, fixing
hicpp-named-parameter/readability-named-parameter.
- Mark the istreambuf_iterator first/last helper variable const in the
five binary-format sentinel tests, fixing misc-const-correctness.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy misc-const-correctness in heterogeneous sentinel test
json_str is only read via .data()/.size() and never reassigned, so
clang-tidy correctly flags it as const-able. Verified against the exact
CI job (silkeh/clang:dev, ci_clang_tidy target) by running clang-tidy
directly on this file plus the five binary-format sentinel tests
touched by prior commits; all are now clean.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add implementation to retrieve start and end positions of json during parse
* Add more unit tests and add start/stop parsing for arrays
* Add raw value for all types
* Add more tests and fix compiler warning
* Amalgamate
* Fix CLang GCC warnings
* Fix error in build
* Style using astyle 3.1
* Fix whitespace changes
* revert
* more whitespace reverts
* Address PR comments
* Fix failing issues
* More whitespace reverts
* Address remaining PR comments
* Address comments
* Switch to using custom base class instead of default basic_json
* Adding a basic using for a json using the new base class. Also address PR comments and fix CI failures
* Address decltype comments
* Diagnostic positions macro (#4)
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Fix missed include deletion
* Add docs and address other PR comments (#5)
* Add docs and address other PR comments
---------
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Address new PR comments and fix CI tests for documentation
* Update documentation based on feedback (#6)
---------
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>
* Address std::size_t and other comments
* Fix new CI issues
* Fix lcov
* Improve lcov case with update to handle_diagnostic_positions call for discarded values
* Fix indentation of LCOV_EXCL_STOP comments
* fix amalgamation astyle issue
---------
Co-authored-by: Sush Shringarputale <sushring@linux.microsoft.com>