Compare commits

...
28 Commits
Author SHA1 Message Date
dependabot[bot]andGitHub 21af527e75 ⬆️ Bump the codeql-action group with 4 updates (#5365) 2026-08-07 19:30:42 +02:00
Niels LohmannandGitHub 23518f54fe Add an Ecosystem page for third-party projects built on nlohmann::json (#5369) 2026-08-07 19:29:58 +02:00
DmitryandGitHub 1c136a66c4 Move the CBOR doc block to the function it describes (#5363)
The block documenting get_char and tag_handler sat above
get_cbor_negative_integer(), which takes neither, so Doxygen attached it
there and parse_cbor_internal() was left undocumented.

Comment placement only.

Signed-off-by: Dmitry <45711841+darkdi@users.noreply.github.com>
2026-08-06 08:30:15 +02:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
c1c19a7bcd ⬆️ Bump lukka/get-cmake from 4.4.0 to 4.4.1 (#5364)
Bumps [lukka/get-cmake](https://github.com/lukka/get-cmake) from 4.4.0 to 4.4.1.
- [Release notes](https://github.com/lukka/get-cmake/releases)
- [Changelog](https://github.com/lukka/get-cmake/blob/main/RELEASE_PROCESS.md)
- [Commits](https://github.com/lukka/get-cmake/compare/e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3...4a7d025fc60f00db0c7b44ebf783d19b52444830)

---
updated-dependencies:
- dependency-name: lukka/get-cmake
  dependency-version: 4.4.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-06 08:30:00 +02:00
Niels LohmannandGitHub bacdabd176 Fix start_pos() for strings containing escape sequences (#5361)
The diagnostic position of a string value was derived by subtracting the
parsed value's length from the end position. Escape sequences make the
source token longer than the value it parses to, so the reported start
position landed inside the string, one byte off per escape sequence:

    input: {"a":"\n\n\n\n\n\n"}
      start_pos() == 11, so the reported range covered  n\n\n\n"
      instead of the documented "\n\n\n\n\n\n"

This contradicts the documented behavior of start_pos(), which is the
position of the opening quote, and it also corrupted the "(bytes N-M)"
part of JSON_DIAGNOSTICS exception messages. Strings with multi-byte
UTF-8 but no escapes were unaffected, which is why this went unnoticed.

Record the offset of the token in the lexer when it starts scanning and
use that, instead of reconstructing it from the parsed value. Booleans,
null and numbers already reported correct positions and are unchanged.

The new lexer member and accessor are compiled only when
JSON_DIAGNOSTIC_POSITIONS is enabled, which is already part of the ABI
tag, so the default build is unaffected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-05 16:01:23 +02:00
Niels LohmannandGitHub d5647e6a3b Resolve the TODO(niels) in get_ubjson_string (#5355)
The comment asked whether the no-op marker 'N' may be ignored when a
string is read. It may not: at that point the next byte must be a string
length type specification, and 'N' is not one. No-ops at positions where
a value may start are already consumed by the callers through
get_ignore_noop(), so nothing is lost by not skipping them here.

Replace the TODO with a comment stating that, and add regression tests
pinning both directions: a no-op is accepted at top level (also
repeated), before and after an array element, and before an object key,
between key and value, and before the closing brace of an object of
unknown size; it is rejected where a length type specification is
expected, i.e. after the 'S' marker of a string value and as the key
length of an object of known size.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-05 14:47:40 +02:00
Niels LohmannandGitHub 9a091d2b82 Do not write BJData ndarrays whose size overflows std::size_t (#5362)
* Do not write BJData ndarrays whose size overflows std::size_t

write_bjdata_ndarray() multiplied the _ArraySize_ dimensions into a
std::size_t without checking for overflow. A product that wraps around
to a value that happens to match the size of _ArrayData_ passed the
length check, and the writer emitted an ndarray header announcing an
element count that cannot be represented:

    {"_ArrayType_":"uint8","_ArraySize_":[9223372036854775808,2],"_ArrayData_":[]}

was encoded as 5b 24 55 23 5b 4d 00 00 00 00 00 00 00 80 69 02 5d, an
ndarray of 2^64 elements followed by no data. Reading that back throws
out_of_range.408 ("excessive ndarray size caused overflow"), so to_bjdata
produced output that from_bjdata rejects. This is reachable by parsing
untrusted JSON and re-encoding it as BJData.

Mirror the overflow check the binary reader already performs, and also
reject a single dimension that does not fit into std::size_t, which the
previous cast silently truncated where std::size_t is narrower than 64
bits. Such objects now fall back to a plain object encoding, which is
what the surrounding type and length validation already does for
annotations it cannot represent, and they round-trip unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document when to_bjdata converts a JData annotation to an ND-array

The BJData page described the 1-D vector case as the only situation in
which an object carrying _ArrayType_/_ArraySize_/_ArrayData_ is not
written as a compact ND-array. The writer has always had several other
fallbacks -- an unknown _ArrayType_, a dimension that is not a
non-negative integer, an _ArrayData_ whose length does not match the
product of the dimensions, and elements that are not numbers of the
annotated kind -- all of which cause the value to be serialized as a
regular JSON object instead.

Spell out the conditions, including the size-overflow check added in the
preceding commit, so the documented behavior matches the implementation.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-05 13:44:15 +02:00
Niels LohmannandGitHub b890b4cba3 CI: build the MinGW Clang matrix without debug info (#5360)
Linking test-regression2_cpp20 intermittently fails with

  unit-regression2.cpp.obj:(.debug_info+0x16): relocation truncated to
  fit: IMAGE_REL_AMD64_SECREL against `.debug_line'

The failure moves between matrix entries from run to run, and the same
commit can pass and fail on consecutive runs, so it is the size of the
debug sections rather than any one Clang version.

The jobs only build and run the tests, so override CMAKE_CXX_FLAGS_DEBUG
to drop the default -g. Everything else about the Debug build is
unchanged: no optimization flag is added and NDEBUG stays undefined, so
JSON_ASSERT remains active.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-05 13:44:05 +02:00
Angadi56andGitHub dca9d49a33 reject out-of-range code points in UTF-32 wide-string input (#5348)
* reject out-of-range code points in UTF-32 wide-string input

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

* remove useless cast to char_traits<char>::int_type

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

---------

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
2026-08-05 13:43:36 +02:00
Petr BělohlávekandGitHub acd87e2336 CI: Add clang 21&22 to Ubuntu CLang build matrix (#5347)
* Add clang 21 to ubuntu build matrix (CI)

Signed-off-by: Petr Belohlavek <me@petrbel.cz>

* Add clang 22 to ubuntu build matrix (CI)

Signed-off-by: Petr Belohlavek <me@petrbel.cz>

* Register Clang 22.1.8 to quality_assurance.md

Signed-off-by: Petr Belohlavek <me@petrbel.cz>

---------

Signed-off-by: Petr Belohlavek <me@petrbel.cz>
2026-08-04 16:14:35 +02:00
Niels LohmannandGitHub ad94fb01cc docs: document size-mismatch behavior of fixed-size conversions (#5352)
Conversions whose element count is fixed by the destination C++ type --
`std::pair`, `std::tuple`, `std::array<T, N>`, C arrays, and
`std::map`/`std::unordered_map` with a non-string key -- read exactly the
elements they need via `at` and never compare the JSON array's size to
that number. Excess elements are silently discarded, while a shortfall
throws `out_of_range.401` rather than a `type_error`. Neither direction
was documented in `conversions.md`, `get.md`, or `from_json.md`.

The existing warning covered only `std::array` and stated that a too-short
JSON array leaves the remaining elements default-constructed with no
exception thrown; that is not what happens. Generalize it to all
fixed-size destinations and correct the shortfall direction.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-04 08:45:50 +02:00
Niels LohmannandGitHub c2e1cc50e0 docs: document the complexity of ordered_map operations (#5353)
* docs: document the complexity of ordered_map operations

ordered_map stores its elements in a std::vector in insertion order and
has no lookup index, so emplace, operator[], at, find, count, erase, and
insert are all linear scans. The documentation stated no complexity for
any operation, neither in ordered_map.md nor in ordered_json.md.

Add a per-operation complexity table and note the consequence: building
or parsing an ordered_json object of n keys is O(n^2). Measured with
-O2 -DNDEBUG for parsing a flat object of n keys, ordered_json is 5x
slower than json at n=2000 and 54x slower at n=16000, with the timings
quadrupling per doubling of n. Cross-reference the table from
ordered_json.md and from the object order page, which recommends
ordered_json without mentioning the cost.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: move the Complexity section after Member functions

scripts/check_structure.py enforces a fixed section order for pages under
docs/mkdocs/docs/api, in which Complexity comes after Member functions.
The section had been placed right after Iterator invalidation, which made
ci_test_build_documentation fail with structure/section_order.

No content change beyond the move; the table columns are realigned to the
narrower content.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-04 08:45:35 +02:00
Niels LohmannandGitHub 173f2a7407 Documentation review: exceptions from binary-format hardening, and tuple reference types (#5359)
* 📝 Document exceptions newly thrown by the binary-format hardening

A round of binary-format input validation (#5274, #5284, #5287, #5332)
added new failure modes without updating exceptions.md, and left two
descriptions factually narrower than the code:

- parse_error.110 said "CBOR or MessagePack"; BSON and UBJSON also
  throw it. Generalized, and added the BSON EOF example (#5332).
- parse_error.112: added the BSON document-size mismatch example
  (#5287).
- parse_error.113 said "while parsing a map key", but its own existing
  UBJSON char example already contradicted that. Broadened to cover
  invalid length specifications, and added the negative-string-length
  example (#5284).
- out_of_range.408 said "of an UBJSON array or object"; CBOR now throws
  it too (#5274). Generalized and added both CBOR examples.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 📝 Correct which types may be referenced in a tuple extraction

The note added in #5271 said a referenced type must be one the library
stores "or an arithmetic type it can convert to/from". The parenthetical
is wrong: is_compatible_reference_type requires an exact match against
the stored types, so std::tuple<int&> is rejected by static_assert even
though int converts fine as a value. Only the value case is permissive.

Spell out the eight admissible types, give the int& counter-example, and
separate the reference restriction from by-value conversion.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-04 08:45:18 +02:00
Niels LohmannandGitHub 1c63a120b6 docs: document how discarded values are removed by the parser callback (#5354)
Follow-up to #5342, which fixed the parser callback leaving a discarded
member behind when an array or a value under an object key was rejected.
The documentation of parser_callback_t only stated that discarded values
in structured types are skipped, without saying that this covers object
parents and that the key is removed along with the value, so there was no
way to tell the fixed behavior from the buggy one.

Spell out the discarding rules, add an example that exercises the cases
the fix repaired, and correct the return value description: a discarded
top-level value is replaced by null, not by "an empty discarded object".

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-04 08:45:02 +02:00
Niels LohmannandGitHub 85889e8843 docs: use HTTPS for the astyle and cppcheck links in README (#5351)
* docs: use HTTPS for the astyle and cppcheck links in README

Both links were still `http://`. `astyle.sourceforge.net` serves HTTPS
directly; `cppcheck.sourceforge.net` redirects to
`https://cppcheck.sourceforge.io`, which is also the URL already used in
`docs/mkdocs/docs/community/quality_assurance.md`, so the redirect is
skipped here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* CI: suppress -Wc2y-extensions for Clang

Clang 22.1 (now shipped by silkeh/clang:latest) diagnoses __COUNTER__ as a
C2y extension, and does so in C++ mode as well. Under -Weverything -Werror
this breaks every ci_test_clang_cxx* / ci_test_clang_libcxx_cxx* target,
independently of the code under test.

The library itself does not use __COUNTER__; all diagnostics originate in
vendored Doctest (DOCTEST_ANONYMOUS, used by TEST_CASE and SECTION).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-04 08:43:59 +02:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3c0a9a99fd ⬆️ Bump actions/stale from 10.4.0 to 11.0.0 (#5350)
Bumps [actions/stale](https://github.com/actions/stale) from 10.4.0 to 11.0.0.
- [Release notes](https://github.com/actions/stale/releases)
- [Changelog](https://github.com/actions/stale/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/stale/compare/1e223db275d687790206a7acac4d1a11bd6fe629...4391f3da665fdf50b6810c1a66712fb9ba21aa93)

---
updated-dependencies:
- dependency-name: actions/stale
  dependency-version: 11.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-03 20:50:15 +02:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
e82724d87f ⬆️ Bump coverallsapp/github-action from 2.3.7 to 2.3.8 (#5349)
Bumps [coverallsapp/github-action](https://github.com/coverallsapp/github-action) from 2.3.7 to 2.3.8.
- [Release notes](https://github.com/coverallsapp/github-action/releases)
- [Commits](https://github.com/coverallsapp/github-action/compare/5cbfd81b66ca5d10c19b062c04de0199c215fb6e...8d6379e14d29928660c4ba802d8e85393440b329)

---
updated-dependencies:
- dependency-name: coverallsapp/github-action
  dependency-version: 2.3.8
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-03 20:50:05 +02:00
78821cd9c2 docs: document standards compliance and parse() vs operator>> strictness (#5326)
* docs: document RFC 8259 / JSONTestSuite compliance and parse() vs operator>> strictness

The compliance story lived only in tests/src/unit-testsuites.cpp, so
drive-by comparisons kept claiming the library "does not fully pass
JSONTestSuite". Make it discoverable:

- README: add a "Standards compliance" note stating that both nst
  JSONTestSuite revisions run in CI, that all mandatory y_/n_ cases pass
  through the strict parse() entry point, and listing the deliberate
  implementation-defined i_ choices (unbounded nesting, silent BOM
  stripping, noncharacters forwarded, strict rejection of invalid UTF-8
  and lone surrogates, out_of_range.406 on numeric overflow).
- features/parsing: add a "Strictness and trailing data" section
  documenting that parse() is strict and rejects trailing data while
  operator>> follows relaxed iostream semantics (parses one value and
  leaves the stream positioned after it) -- the single place a naive
  test yields a "non-compliant" result.

Documentation only; no parser behavior change. Closes #5290.

Signed-off-by: manon <youdie006@users.noreply.github.com>

* docs: correct test-data vendoring and parse()/operator>> claims per review

- README: the JSONTestSuite data is downloaded from nlohmann/json_test_data at
  configure time, not vendored/committed; say so.
- README: only the updated suite runs y_ and n_ cases through strict parse();
  the original suite's y_ cases go through operator>>. Narrow the claim.
- parsing/index.md and operator_gtgt.md: note that operator>> consumes a number's
  terminating byte, so concatenated numbers must be whitespace-separated (1 2
  works, 1true does not); structural and literal values are unaffected.

Signed-off-by: manon <youdie006@users.noreply.github.com>

---------

Signed-off-by: manon <youdie006@users.noreply.github.com>
Co-authored-by: manon <youdie006@users.noreply.github.com>
2026-08-03 19:15:53 +02:00
Angadi56andGitHub 68f0722a19 remove discarded array from parent object in end_array (#5342)
* remove discarded array from parent object in end_array

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

* remove discarded scalar value from parent object in handle_value

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

---------

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
2026-08-03 19:13:31 +02:00
Angadi56andGitHub 5f121d8c50 avoid sign extension in char_traits<signed char>::to_int_type (#5336)
* avoid sign extension in char_traits<signed char>::to_int_type

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

* spell out-of-range signed char constants as negative values (MSVC C4309)

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

---------

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
2026-08-03 08:35:53 +02:00
Niels LohmannandGitHub 585929bff9 Fix Clang deprecation warning for json_pointer operator== with ordered_json (#5289)
is_comparable used a flat && chain to both exclude json_pointer/string
comparisons (added for #4621) and check whether Compare(A, B) is well-formed.
Naming std::is_constructible<decltype(...)> as a later operand of that chain
still causes the decltype to be substituted regardless of the first
operand's value, since the operands aren't lazily deferred like
std::conjunction would defer them. That instantiates the transparent
std::equal_to<>::operator() used by ordered_json, whose noexcept-specifier
evaluates the deprecated json_pointer/string operator==, which Clang (unlike
GCC in this case) warns about even though the result is discarded.

Split is_comparable so the Compare(A, B) checks live in a separate helper
that is only referenced from the specialization selected when
is_json_pointer_of is false, so the decltype is never written when A/B are
a json_pointer/string pair, regardless of compiler.

Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
2026-08-03 08:20:41 +02:00
Niels LohmannandGitHub 31ba5208c8 docs: qualify the operator>> stream positioning guarantee (#5343)
operator>>'s notes state that it leaves the stream positioned right
after the parsed value, so that concatenated JSON values can be read
back to back. That does not hold when the value is a number: a number
is only terminated by the character that follows it, and the lexer's
unget() is simulated (it rewinds only the lexer's own bookkeeping),
so that character stays consumed from the stream.

Document the actual behaviour: the guarantee holds for all value types
except numbers, which must be followed by whitespace. Also qualify the
cross-reference on the JSON Lines page, which repeated the unqualified
claim.

Documentation only; the behaviour itself is tracked in #5340.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-03 08:18:03 +02:00
tomatotomataandGitHub 2222d386c9 fix: check CBOR tagged subtype reads (#5339) 2026-07-31 22:06:55 +02:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
eaedec859a ⬆️ Bump the codeql-action group with 4 updates (#5341)
Bumps the codeql-action group with 4 updates: [github/codeql-action/init](https://github.com/github/codeql-action), [github/codeql-action/autobuild](https://github.com/github/codeql-action), [github/codeql-action/analyze](https://github.com/github/codeql-action) and [github/codeql-action/upload-sarif](https://github.com/github/codeql-action).


Updates `github/codeql-action/init` from 4.37.2 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e0647621c2984b5ed2f768cb892365bf2a616ad1...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81)

Updates `github/codeql-action/autobuild` from 4.37.2 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e0647621c2984b5ed2f768cb892365bf2a616ad1...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81)

Updates `github/codeql-action/analyze` from 4.37.2 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e0647621c2984b5ed2f768cb892365bf2a616ad1...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81)

Updates `github/codeql-action/upload-sarif` from 4.37.2 to 4.37.3
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/e0647621c2984b5ed2f768cb892365bf2a616ad1...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81)

---
updated-dependencies:
- dependency-name: github/codeql-action/init
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/autobuild
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.37.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-31 20:45:32 +02:00
Angadi56andGitHub d94cbd99dc reject CBOR array/map length equal to the indefinite-length marker (#5274)
* reject CBOR array/map length equal to the indefinite-length marker

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

* reject CBOR lengths that do not fit in std::size_t via value_in_range_of

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

---------

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
2026-07-31 18:04:30 +02:00
bc48951128 Fix UBJSON high-precision floating-point overflow handling (#5323)
This adds a std::isfinite check to the UBJSON floating-point parsing path, throwing out_of_range.406 on overflow. This makes the UBJSON parser's behavior consistent with the normal JSON parser. Fixes #5322.

Signed-off-by: AJ369ninja <abhishek.j@iitg.ac.in>
Co-authored-by: AJ369ninja <abhishek.j@iitg.ac.in>
2026-07-31 18:04:13 +02:00
Angadi56andGitHub fd72ecfc8c validate ndarray element types in write_bjdata_ndarray (#5301)
* validate ndarray element types in write_bjdata_ndarray

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

* read ndarray elements through get<> instead of a fixed union member

_ArrayType_ names the wire type, not how the value is stored: parsing
keeps a non-negative integer as number_unsigned while the C++ API keeps
an int literal as number_integer. Selecting the union member from the
type marker therefore reads the inactive alternative for one of the two,
so read through get<> instead, which dispatches on the active member.

Also reject a negative _ArraySize_ entry, which is not a usable
dimension, and cover the parse-built path in the tests.

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>

---------

Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
2026-07-31 08:52:27 +02:00
YingqiDuanandGitHub de8a099ba5 Document BSON interoperability and subtype-less binary round trips (#5330)
- warn about BSON marker 0x11 interoperability in both directions
- explain subtype-less binary normalization to subtype 0x00
- add a round-trip test for binary values without a subtype

Signed-off-by: YingqiDuan <141370165+YingqiDuan@users.noreply.github.com>
2026-07-31 08:33:07 +02:00
44 changed files with 1244 additions and 176 deletions
+3 -3
View File
@@ -38,14 +38,14 @@ jobs:
# Initializes the CodeQL tools for scanning. # Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL - name: Initialize CodeQL
uses: github/codeql-action/init@e0647621c2984b5ed2f768cb892365bf2a616ad1 # v4.37.2 uses: github/codeql-action/init@f205ea1c3313d32999d8d6a48b4f6530d4437b38 # v4.37.4
with: with:
languages: c-cpp languages: c-cpp
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java). # Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below) # If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild - name: Autobuild
uses: github/codeql-action/autobuild@e0647621c2984b5ed2f768cb892365bf2a616ad1 # v4.37.2 uses: github/codeql-action/autobuild@f205ea1c3313d32999d8d6a48b4f6530d4437b38 # v4.37.4
- name: Perform CodeQL Analysis - name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@e0647621c2984b5ed2f768cb892365bf2a616ad1 # v4.37.2 uses: github/codeql-action/analyze@f205ea1c3313d32999d8d6a48b4f6530d4437b38 # v4.37.4
+1 -1
View File
@@ -43,6 +43,6 @@ jobs:
output: 'flawfinder_results.sarif' output: 'flawfinder_results.sarif'
- name: Upload analysis results to GitHub Security tab - name: Upload analysis results to GitHub Security tab
uses: github/codeql-action/upload-sarif@e0647621c2984b5ed2f768cb892365bf2a616ad1 # v4.37.2 uses: github/codeql-action/upload-sarif@f205ea1c3313d32999d8d6a48b4f6530d4437b38 # v4.37.4
with: with:
sarif_file: ${{github.workspace}}/flawfinder_results.sarif sarif_file: ${{github.workspace}}/flawfinder_results.sarif
+1 -1
View File
@@ -76,6 +76,6 @@ jobs:
# Upload the results to GitHub's code scanning dashboard. # Upload the results to GitHub's code scanning dashboard.
- name: "Upload to code-scanning" - name: "Upload to code-scanning"
uses: github/codeql-action/upload-sarif@e0647621c2984b5ed2f768cb892365bf2a616ad1 # v4.37.2 uses: github/codeql-action/upload-sarif@f205ea1c3313d32999d8d6a48b4f6530d4437b38 # v4.37.4
with: with:
sarif_file: results.sarif sarif_file: results.sarif
+1 -1
View File
@@ -61,7 +61,7 @@ jobs:
# Upload SARIF file generated in previous step # Upload SARIF file generated in previous step
- name: Upload SARIF file - name: Upload SARIF file
uses: github/codeql-action/upload-sarif@e0647621c2984b5ed2f768cb892365bf2a616ad1 # v4.37.2 uses: github/codeql-action/upload-sarif@f205ea1c3313d32999d8d6a48b4f6530d4437b38 # v4.37.4
with: with:
sarif_file: semgrep.sarif sarif_file: semgrep.sarif
if: always() if: always()
+1 -1
View File
@@ -20,7 +20,7 @@ jobs:
with: with:
egress-policy: audit egress-policy: audit
- uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v10.4.0 - uses: actions/stale@4391f3da665fdf50b6810c1a66712fb9ba21aa93 # v11.0.0
with: with:
stale-issue-label: 'state: stale' stale-issue-label: 'state: stale'
stale-pr-label: 'state: stale' stale-pr-label: 'state: stale'
+18 -18
View File
@@ -25,7 +25,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -47,7 +47,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -70,7 +70,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -89,7 +89,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -108,7 +108,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -142,7 +142,7 @@ jobs:
name: code-coverage-report name: code-coverage-report
path: ${{ github.workspace }}/build/html path: ${{ github.workspace }}/build/html
- name: Publish report to Coveralls - name: Publish report to Coveralls
uses: coverallsapp/github-action@5cbfd81b66ca5d10c19b062c04de0199c215fb6e # v2.3.7 uses: coverallsapp/github-action@8d6379e14d29928660c4ba802d8e85393440b329 # v2.3.8
with: with:
github-token: ${{ secrets.GITHUB_TOKEN }} github-token: ${{ secrets.GITHUB_TOKEN }}
path-to-lcov: ${{ github.workspace }}/build/json.info.filtered.noexcept path-to-lcov: ${{ github.workspace }}/build/json.info.filtered.noexcept
@@ -184,7 +184,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: CXX=g++-${{ matrix.compiler }} cmake -S . -B build -DJSON_CI=On run: CXX=g++-${{ matrix.compiler }} cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -202,7 +202,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -212,14 +212,14 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
strategy: strategy:
matrix: matrix:
compiler: ['3.4', '3.5', '3.6', '3.7', '3.8', '3.9', '4', '5', '6', '7', '8', '9', '10', '11', '12', '13', '14', '15-bullseye', '16', '17', '18', '19', '20', 'latest'] compiler: ['3.4', '3.5', '3.6', '3.7', '3.8', '3.9', '4', '5', '6', '7', '8', '9', '10', '11', '12', '13', '14', '15-bullseye', '16', '17', '18', '19', '20', '21', '22', 'latest']
container: silkeh/clang:${{ matrix.compiler }} container: silkeh/clang:${{ matrix.compiler }}
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Set env FORCE_STDCPPFS_FLAG for clang 7 / 8 / 9 / 10 - name: Set env FORCE_STDCPPFS_FLAG for clang 7 / 8 / 9 / 10
run: echo "JSON_FORCED_GLOBAL_COMPILE_OPTIONS=-DJSON_HAS_FILESYSTEM=0;-DJSON_HAS_EXPERIMENTAL_FILESYSTEM=0" >> "$GITHUB_ENV" run: echo "JSON_FORCED_GLOBAL_COMPILE_OPTIONS=-DJSON_HAS_FILESYSTEM=0;-DJSON_HAS_EXPERIMENTAL_FILESYSTEM=0" >> "$GITHUB_ENV"
if: ${{ matrix.compiler == '7' || matrix.compiler == '8' || matrix.compiler == '9' || matrix.compiler == '10' }} if: ${{ matrix.compiler == '7' || matrix.compiler == '8' || matrix.compiler == '9' || matrix.compiler == '10' }}
@@ -239,7 +239,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -259,7 +259,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build with libc++ - name: Build with libc++
@@ -286,7 +286,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -306,7 +306,7 @@ jobs:
# import-std support. Its opt-in token is CMake-version-specific, so pin # import-std support. Its opt-in token is CMake-version-specific, so pin
# CMake to the version whose token is set in tests/module_cpp20/CMakeLists.txt. # CMake to the version whose token is set in tests/module_cpp20/CMakeLists.txt.
- name: Get pinned CMake and ninja - name: Get pinned CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
with: with:
cmakeVersion: 4.3.4 cmakeVersion: 4.3.4
# Clang: the std library module is provided by libc++ (the image's libstdc++ # Clang: the std library module is provided by libc++ (the image's libstdc++
@@ -332,7 +332,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -347,7 +347,7 @@ jobs:
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -359,7 +359,7 @@ jobs:
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DJSON_CI=On run: cmake -S . -B build -DJSON_CI=On
- name: Build - name: Build
@@ -379,7 +379,7 @@ jobs:
with: with:
persist-credentials: false persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Run CMake - name: Run CMake
run: cmake -S . -B build -DCMAKE_TOOLCHAIN_FILE=$EMSDK/upstream/emscripten/cmake/Modules/Platform/Emscripten.cmake -GNinja run: cmake -S . -B build -DCMAKE_TOOLCHAIN_FILE=$EMSDK/upstream/emscripten/cmake/Modules/Platform/Emscripten.cmake -GNinja
- name: Build - name: Build
+8 -2
View File
@@ -88,7 +88,7 @@ jobs:
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
- name: Set extra CXX_FLAGS for latest std_version - name: Set extra CXX_FLAGS for latest std_version
# /wd5285 silences C5285 emitted by the bundled third-party doctest.h, which # /wd5285 silences C5285 emitted by the bundled third-party doctest.h, which
# specializes std::tuple (newly diagnosed by the VS2026 v145 toolset) # specializes std::tuple (newly diagnosed by the VS2026 v145 toolset)
@@ -153,10 +153,16 @@ jobs:
with: with:
platform: x64 platform: x64
version: 12.2.0 # https://github.com/egor-tensin/setup-mingw/issues/14 version: 12.2.0 # https://github.com/egor-tensin/setup-mingw/issues/14
# CMAKE_CXX_FLAGS_DEBUG is overridden to drop the default -g: linking
# test-regression2_cpp20 intermittently fails with "relocation truncated
# to fit: IMAGE_REL_AMD64_SECREL against `.debug_line'" because the
# MinGW linker cannot relocate the debug sections this test produces.
# The tests are only built and run here, so the debug info is not used.
- name: Run CMake - name: Run CMake
run: cmake -S . -B build ^ run: cmake -S . -B build ^
-DCMAKE_CXX_COMPILER="C:/Program Files/LLVM/bin/clang++.exe" ^ -DCMAKE_CXX_COMPILER="C:/Program Files/LLVM/bin/clang++.exe" ^
-DCMAKE_CXX_FLAGS="--target=x86_64-w64-mingw32 -stdlib=libstdc++ -pthread" ^ -DCMAKE_CXX_FLAGS="--target=x86_64-w64-mingw32 -stdlib=libstdc++ -pthread" ^
-DCMAKE_CXX_FLAGS_DEBUG="-g0" ^
-DCMAKE_EXE_LINKER_FLAGS="-lwinpthread" ^ -DCMAKE_EXE_LINKER_FLAGS="-lwinpthread" ^
-G"MinGW Makefiles" ^ -G"MinGW Makefiles" ^
-DCMAKE_BUILD_TYPE=Debug ^ -DCMAKE_BUILD_TYPE=Debug ^
@@ -193,7 +199,7 @@ jobs:
# import-std support. Its opt-in token is CMake-version-specific, so pin # import-std support. Its opt-in token is CMake-version-specific, so pin
# CMake to the version whose token is set in tests/module_cpp20/CMakeLists.txt. # CMake to the version whose token is set in tests/module_cpp20/CMakeLists.txt.
- name: Get pinned CMake and ninja - name: Get pinned CMake and ninja
uses: lukka/get-cmake@e6906078ebd1ccb8ce51ab4626ac46a1b5a517e3 # v4.4.0 uses: lukka/get-cmake@4a7d025fc60f00db0c7b44ebf783d19b52444830 # v4.4.1
with: with:
cmakeVersion: 4.3.4 cmakeVersion: 4.3.4
- name: Run CMake (Debug) - name: Run CMake (Debug)
+17 -2
View File
@@ -42,6 +42,7 @@
- [Specializing enum conversion](#specializing-enum-conversion) - [Specializing enum conversion](#specializing-enum-conversion)
- [Binary formats (BSON, CBOR, MessagePack, UBJSON, and BJData)](#binary-formats-bson-cbor-messagepack-ubjson-and-bjdata) - [Binary formats (BSON, CBOR, MessagePack, UBJSON, and BJData)](#binary-formats-bson-cbor-messagepack-ubjson-and-bjdata)
- [Customers](#customers) - [Customers](#customers)
- [Ecosystem](#ecosystem)
- [Supported compilers](#supported-compilers) - [Supported compilers](#supported-compilers)
- [Integration](#integration) - [Integration](#integration)
- [CMake](#cmake) - [CMake](#cmake)
@@ -1186,6 +1187,11 @@ The library is used in multiple projects, applications, operating systems, etc.
[![logos of customers using the library](docs/mkdocs/docs/images/customers.png)](https://json.nlohmann.me/home/customers/) [![logos of customers using the library](docs/mkdocs/docs/images/customers.png)](https://json.nlohmann.me/home/customers/)
## Ecosystem
Beyond projects that use the library, there are third-party projects that build on top of it - schema validators,
language bindings, format converters, and the like. See the curated [Ecosystem](https://json.nlohmann.me/community/ecosystem/) page.
## Supported compilers ## Supported compilers
Though it's 2026 already, the support for C++11 is still a bit sparse. Currently, the following compilers are known to work: Though it's 2026 already, the support for C++11 is still a bit sparse. Currently, the following compilers are known to work:
@@ -1801,13 +1807,13 @@ The library itself consists of a single header file licensed under the MIT licen
- [**amalgamate.py - Amalgamate C source and header files**](https://github.com/edlund/amalgamate) to create a single header file - [**amalgamate.py - Amalgamate C source and header files**](https://github.com/edlund/amalgamate) to create a single header file
- [**American fuzzy lop**](https://lcamtuf.coredump.cx/afl/) for fuzz testing - [**American fuzzy lop**](https://lcamtuf.coredump.cx/afl/) for fuzz testing
- [**AppVeyor**](https://www.appveyor.com) for [continuous integration](https://ci.appveyor.com/project/nlohmann/json) on Windows - [**AppVeyor**](https://www.appveyor.com) for [continuous integration](https://ci.appveyor.com/project/nlohmann/json) on Windows
- [**Artistic Style**](http://astyle.sourceforge.net) for automatic source code indentation - [**Artistic Style**](https://astyle.sourceforge.net) for automatic source code indentation
- [**Clang**](https://clang.llvm.org) for compilation with code sanitizers - [**Clang**](https://clang.llvm.org) for compilation with code sanitizers
- [**CMake**](https://cmake.org) for build automation - [**CMake**](https://cmake.org) for build automation
- [**Codacy**](https://www.codacy.com) for further [code analysis](https://app.codacy.com/gh/nlohmann/json/dashboard) - [**Codacy**](https://www.codacy.com) for further [code analysis](https://app.codacy.com/gh/nlohmann/json/dashboard)
- [**Coveralls**](https://coveralls.io) to measure [code coverage](https://coveralls.io/github/nlohmann/json) - [**Coveralls**](https://coveralls.io) to measure [code coverage](https://coveralls.io/github/nlohmann/json)
- [**Coverity Scan**](https://scan.coverity.com) for [static analysis](https://scan.coverity.com/projects/nlohmann-json) - [**Coverity Scan**](https://scan.coverity.com) for [static analysis](https://scan.coverity.com/projects/nlohmann-json)
- [**cppcheck**](http://cppcheck.sourceforge.net) for static analysis - [**cppcheck**](https://cppcheck.sourceforge.io) for static analysis
- [**doctest**](https://github.com/onqtam/doctest) for the unit tests - [**doctest**](https://github.com/onqtam/doctest) for the unit tests
- [**GitHub Changelog Generator**](https://github.com/skywinder/github-changelog-generator) to generate the [ChangeLog](https://github.com/nlohmann/json/blob/develop/ChangeLog.md) - [**GitHub Changelog Generator**](https://github.com/skywinder/github-changelog-generator) to generate the [ChangeLog](https://github.com/nlohmann/json/blob/develop/ChangeLog.md)
- [**Google Benchmark**](https://github.com/google/benchmark) to implement the benchmarks - [**Google Benchmark**](https://github.com/google/benchmark) to implement the benchmarks
@@ -1822,6 +1828,15 @@ The library itself consists of a single header file licensed under the MIT licen
## Notes ## Notes
### Standards compliance
The library targets strict conformance with [RFC 8259](https://tools.ietf.org/html/rfc8259.html). Both the original [JSONTestSuite](https://github.com/nst/JSONTestSuite) and its updated revision are exercised in CI; their test data is downloaded from [`nlohmann/json_test_data`](https://github.com/nlohmann/json_test_data) at configure time rather than committed to this repository (see [`tests/src/unit-testsuites.cpp`](https://github.com/nlohmann/json/blob/develop/tests/src/unit-testsuites.cpp)):
- The updated revision runs all mandatory `y_` (must-accept) and `n_` (must-reject) cases through the strict [`parse()`](https://json.nlohmann.me/api/basic_json/parse/) entry point; the original suite runs its `n_` cases through `parse()` and its `y_` cases through [`operator>>`](https://json.nlohmann.me/api/operator_gtgt/).
- The `i_` (implementation-defined) cases are, by RFC 8259, free to be accepted *or* rejected, so "passing all `i_` cases" is not a meaningful conformance metric. The library makes deliberate, documented choices there: nesting depth is not artificially limited, a leading UTF-8 byte order mark is silently ignored, [Unicode noncharacters](https://www.unicode.org/faq/private_use.html#nonchar1) are forwarded unchanged, invalid UTF-8 and lone/unpaired UTF-16 surrogates are rejected (stricter than required), and a number that cannot be stored without becoming `NaN`/`INF` raises [`out_of_range.406`](https://json.nlohmann.me/home/exceptions/#jsonexceptionout_of_range406).
One behavioral nuance is worth calling out, because a superficial test often misreads it as non-compliance: [`parse()`](https://json.nlohmann.me/api/basic_json/parse/) is strict and rejects trailing data after a value, whereas [`operator>>`](https://json.nlohmann.me/api/operator_gtgt/) follows relaxed iostream semantics — it parses a single value and leaves the stream positioned right after it. Feeding "a valid document followed by trailing bytes" through `operator>>` reports success; the same input through `parse()` is rejected. This is a documented two-API design, not a conformance gap. See [**parsing**](https://json.nlohmann.me/features/parsing/) for details.
### Character encoding ### Character encoding
The library supports **Unicode input** as follows: The library supports **Unicode input** as follows:
+4
View File
@@ -5,6 +5,9 @@
# -Wno-extra-semi-stmt The library uses assert which triggers this warning. # -Wno-extra-semi-stmt The library uses assert which triggers this warning.
# -Wno-padded We do not care about padding warnings. # -Wno-padded We do not care about padding warnings.
# -Wno-covered-switch-default All switches list all cases and a default case. # -Wno-covered-switch-default All switches list all cases and a default case.
# -Wno-c2y-extensions Clang 22.1 diagnoses __COUNTER__ as a C2y extension, also in
# C++ mode. The library does not use __COUNTER__; the warnings
# all come from vendored Doctest (SECTION/TEST_CASE macros).
# -Wno-unsafe-buffer-usage Pervasive: the library's own low-level numeric/buffer code # -Wno-unsafe-buffer-usage Pervasive: the library's own low-level numeric/buffer code
# (to_chars, serializer, lexer, binary reader/writer, input # (to_chars, serializer, lexer, binary reader/writer, input
# adapters, json_pointer) plus vendored Doctest itself (~208 # adapters, json_pointer) plus vendored Doctest itself (~208
@@ -20,5 +23,6 @@ set(CLANG_CXXFLAGS
-Wno-extra-semi-stmt -Wno-extra-semi-stmt
-Wno-padded -Wno-padded
-Wno-covered-switch-default -Wno-covered-switch-default
-Wno-c2y-extensions
-Wno-unsafe-buffer-usage -Wno-unsafe-buffer-usage
) )
@@ -29,7 +29,14 @@ Discarding a value (i.e., returning `#!cpp false`) has different effects dependi
called: called:
- Discarded values in structured types are skipped. That is, the parser will behave as if the discarded value was never - Discarded values in structured types are skipped. That is, the parser will behave as if the discarded value was never
read. read. This holds for every value type and for both kinds of parent: a discarded element is removed from the
surrounding array, and a discarded member is removed from the surrounding object together with its key.
- Arrays and objects can be discarded either at their `parse_event_t::array_start`/`parse_event_t::object_start` event
or at their `parse_event_t::array_end`/`parse_event_t::object_end` event, and both remove the whole value. Discarding
it at the start event also means the callback is called neither for the content of the value nor for its matching end
event.
- Discarding a `parse_event_t::key` event discards the whole object member. The callback is still called for the
associated value, but its return value has no further effect.
- In case a value outside a structured type is skipped, it is replaced with `null`. This case happens if the top-level - In case a value outside a structured type is skipped, it is replaced with `null`. This case happens if the top-level
element is skipped. element is skipped.
@@ -49,7 +56,7 @@ called:
## Return value ## Return value
Whether the JSON value which called the function during parsing should be kept (`#!cpp true`) or not (`#!cpp false`). In Whether the JSON value which called the function during parsing should be kept (`#!cpp true`) or not (`#!cpp false`). In
the latter case, it is either skipped completely or replaced by an empty discarded object. the latter case, it is skipped completely, or replaced by `null` if it is the top-level value.
## Examples ## Examples
@@ -68,6 +75,21 @@ the latter case, it is either skipped completely or replaced by an empty discard
--8<-- "examples/parse__string__parser_callback_t.output" --8<-- "examples/parse__string__parser_callback_t.output"
``` ```
??? example
The example below shows where discarded values are removed. The array and the number are discarded in different
ways, but in each case the parse result contains neither the value nor its key.
```cpp
--8<-- "examples/parser_callback_t.cpp"
```
Output:
```json
--8<-- "examples/parser_callback_t.output"
```
## See also ## See also
- [parse](parse.md) deserialize from a compatible input - [parse](parse.md) deserialize from a compatible input
@@ -76,3 +98,5 @@ the latter case, it is either skipped completely or replaced by an empty discard
## Version history ## Version history
- Added in version 1.0.0. - Added in version 1.0.0.
- Fixed in version 3.13.0 to also remove discarded values from a parent object; before, discarding an array or a value
stored under an object key left a discarded member behind, which made the parse result serialize to invalid JSON.
+32 -5
View File
@@ -33,17 +33,44 @@ A UTF-8 byte order mark is silently ignored.
Invalid Unicode escapes and unpaired surrogates in the input are reported as Invalid Unicode escapes and unpaired surrogates in the input are reported as
[`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) with a detailed message. [`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) with a detailed message.
`operator>>` parses exactly one JSON value and leaves the stream positioned right after it, so it can be called `operator>>` parses exactly one JSON value, so it can be called repeatedly to read a sequence of concatenated JSON
repeatedly to read a sequence of concatenated JSON values from the same stream: values from the same stream:
```cpp ```cpp
json j1, j2; json j1, j2;
input >> j1; // parses the first value, stream now positioned right after it input >> j1; // parses the first value
input >> j2; // parses the next value input >> j2; // parses the next value
``` ```
Note this does **not** work for [JSON Lines](../features/parsing/json_lines.md) (newline-delimited JSON) input -- !!! warning "A number must be followed by whitespace"
see that page for why and for the recommended alternative.
A number is only terminated by the character that follows it. That character is read from the stream to detect the
end of the number, and it is **not** put back. When a value that is a number is immediately followed by the next
value, the first character of that next value is lost:
```cpp
std::istringstream input("1true");
json j1, j2;
input >> j1; // j1 == 1
input >> j2; // throws parse_error.101: the stream now starts at "rue"
```
Separating the values with whitespace avoids this, because the character that is eaten is then the separator:
```cpp
std::istringstream input("1 true");
json j1, j2;
input >> j1; // j1 == 1
input >> j2; // j2 == true
```
Only numbers are affected. Values ending in a self-delimiting character do not read past themselves, so
`truefalse`, `[1][2]`, `{"a":1}{"b":2}`, and `"a""b"` can be read back to back without a separator.
This is tracked in [#5340](https://github.com/nlohmann/json/issues/5340).
Note that reading concatenated values does **not** work for [JSON Lines](../features/parsing/json_lines.md)
(newline-delimited JSON) input -- see that page for why and for the recommended alternative.
!!! warning "Deprecation" !!! warning "Deprecation"
+6
View File
@@ -13,6 +13,12 @@ Therefore, adding object elements can yield a reallocation in which case all ite
[`end()`](basic_json/end.md) iterator) and all references to the elements are invalidated. Also, any iterator or [`end()`](basic_json/end.md) iterator) and all references to the elements are invalidated. Also, any iterator or
reference after the insertion point will point to the same index, which is now a different value. reference after the insertion point will point to the same index, which is now a different value.
## Complexity
[`ordered_map`](ordered_map.md) has no lookup index: every key-based object operation is a linear scan, so building or
parsing an object of `n` keys costs O(n²) rather than O(n log n). See
[`ordered_map` complexity](ordered_map.md#complexity) for the per-operation table and for measured numbers.
## Examples ## Examples
??? example ??? example
+42
View File
@@ -56,6 +56,48 @@ std::equal_to<> // since C++14
- **find** - **find**
- **insert** - **insert**
## Complexity
Because the elements are stored in a `std::vector` in insertion order, there is no index to look a key up by. Every
key-based operation performs a **linear scan** over the stored elements. With `n` denoting the number of elements in the
container:
| Operation | Complexity | Note |
|----------------------------------------|----------------|----------------------------------------------------------|
| **emplace** | O(n) | scans for an existing key, then appends (amortized O(1)) |
| **operator\[\]** | O(n) | delegates to **emplace** (non-const) or **at** (const) |
| **at** | O(n) | throws `#!cpp std::out_of_range` if the key is not found |
| **find** | O(n) | |
| **count** | O(n) | the result is always 0 or 1 |
| **erase(key)** | O(n) | scan, then move the remaining elements one position down |
| **erase(pos)**, **erase(first, last)** | O(n) | moves all elements after the erased range |
| **insert(value)** | O(n) | equivalent to **emplace** |
| **insert(first, last)** | O((n + m) * m) | for `m` inserted elements |
This differs from `#!cpp std::map`, where the same operations are O(log n).
!!! warning "Quadratic cost of building large objects"
Because every insertion scans all elements inserted so far, building an object of `n` distinct keys costs
**O(n²)** in total. This applies to filling an [`ordered_json`](ordered_json.md) object key by key as well as to
parsing one, since the parser inserts each key as it is read.
The cost is negligible for the object sizes typically found in configuration files or API payloads, but it grows
steeply for machine-generated objects with many thousands of keys. Measured with `-O2 -DNDEBUG` for parsing a flat
object of `n` keys, relative to `#!cpp nlohmann::json` (which uses `#!cpp std::map`):
| `n` | `json` | `ordered_json` | factor |
|--------|--------|----------------|--------|
| 2000 | 0.7 ms | 3.6 ms | 5× |
| 4000 | 0.8 ms | 14.0 ms | 19× |
| 8000 | 1.6 ms | 67.8 ms | 43× |
| 16 000 | 3.3 ms | 181.6 ms | 54× |
If key order matters for objects of that size, consider a container with a lookup index, such as
[`tsl::ordered_map`](https://github.com/Tessil/ordered-map)
([integration](https://github.com/nlohmann/json/issues/546#issuecomment-304447518)), as the object type -- see
[object order](../features/object_order.md).
## Examples ## Examples
??? example ??? example
+40
View File
@@ -0,0 +1,40 @@
# Ecosystem
The projects below build on top of `nlohmann::json` rather than merely using it - schema validators, language
bindings, format converters, and similar building blocks. The list is not exhaustive, and is curated rather than
automatically generated. If you maintain or know of a project that belongs here,
[please let me know](mailto:mail@nlohmann.me).
For products, applications, and organizations that use the library, see [Customers](../home/customers.md) instead.
## Schema validation
- [**json-schema-validator**](https://github.com/pboettch/json-schema-validator), a JSON Schema (draft 7) validator
with human-readable error messages
## Serialization and reflection
- [**nlohmann_json_reflect**](https://github.com/1261385937/nlohmann_json_reflect), a reflection extension for
(de)serializing nested containers-in-structs-in-containers
## Encodings
- [**base-encode-decode**](https://github.com/saxonnicholls/base-encode-decode), a header-only Base64/32/16/8/4/2
(and DNA/RNA) encoding library, with an adapter that serializes binary data through `nlohmann::json`
## Language bindings and interop
- [**pybind11_json**](https://github.com/pybind/pybind11_json), a bidirectional type caster between
`nlohmann::json` and Python objects for [pybind11](https://github.com/pybind/pybind11) bindings
- [**nanobind_json**](https://github.com/ianhbell/nanobind_json), the same idea for
[nanobind](https://github.com/wjakob/nanobind) bindings
- [**nlohmann_json_qt**](https://github.com/dpurgin/nlohmann_json_qt), deserialization helpers for Qt types
(`QString`, `QUrl`, `QDateTime`, `QVector`, ...) from `nlohmann::json`
- [**vulkan2json**](https://github.com/Fadis/vulkan2json), serialization and deserialization of Vulkan API structs
## Format converters
- [**tojson**](https://github.com/mircodz/tojson), a header-only converter between YAML/XML documents and
`nlohmann::json`
- [**json2xml**](https://github.com/testillano/json2xml), a header-only converter from `nlohmann::json` to XML for
simple configuration documents
+1
View File
@@ -1,5 +1,6 @@
# Community # Community
- [Ecosystem](ecosystem.md) - third-party projects built on top of this library
- [Code of Conduct](code_of_conduct.md) - the rules and norms of this project - [Code of Conduct](code_of_conduct.md) - the rules and norms of this project
- [Contribution Guidelines](contribution_guidelines.md) - guidelines how to contribute to this project - [Contribution Guidelines](contribution_guidelines.md) - guidelines how to contribute to this project
- [Governance](governance.md) - the governance model of this project - [Governance](governance.md) - the governance model of this project
@@ -66,6 +66,7 @@ Note: Some modern features (like C++20 ranges or filesystem support) may be disa
| Clang 20.1.1 | x86_64 | Ubuntu 22.04.1 LTS | GitHub | | Clang 20.1.1 | x86_64 | Ubuntu 22.04.1 LTS | GitHub |
| Clang 20.1.8 with GNU-like command-line | x86_64 | Windows Server 2022 (Build 20348) | GitHub | | Clang 20.1.8 with GNU-like command-line | x86_64 | Windows Server 2022 (Build 20348) | GitHub |
| Clang 21.1.8 | x86_64 | Ubuntu 22.04.1 LTS | GitHub | | Clang 21.1.8 | x86_64 | Ubuntu 22.04.1 LTS | GitHub |
| Clang 22.1.8 | x86_64 | Ubuntu 22.04.1 LTS | GitHub |
| CUDA 11.8.0 (nvcc) | x86_64 | Ubuntu 22.04 LTS | GitHub | | CUDA 11.8.0 (nvcc) | x86_64 | Ubuntu 22.04 LTS | GitHub |
| CUDA 12.1.1 (nvcc) | x86_64 | Ubuntu 22.04 LTS | GitHub | | CUDA 12.1.1 (nvcc) | x86_64 | Ubuntu 22.04 LTS | GitHub |
| CUDA 12.6.3 (nvcc) | x86_64 | Ubuntu 22.04 LTS | GitHub | | CUDA 12.6.3 (nvcc) | x86_64 | Ubuntu 22.04 LTS | GitHub |
@@ -0,0 +1,47 @@
#include <iostream>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
// a JSON text with an array and a number inside an object
auto text = R"({"IDs": [116, 943], "Width": 800})";
// discard the array when the parser reads its opening bracket
json j_array_start = json::parse(text, [](int /*depth*/, json::parse_event_t event, json & /*parsed*/)
{
return event != json::parse_event_t::array_start;
});
// discard the same array when the parser reads its closing bracket
json j_array_end = json::parse(text, [](int /*depth*/, json::parse_event_t event, json & /*parsed*/)
{
return event != json::parse_event_t::array_end;
});
// discard the number, but keep its key
json j_value = json::parse(text, [](int /*depth*/, json::parse_event_t event, json & parsed)
{
return !(event == json::parse_event_t::value && parsed == json(800));
});
// discard the key of the number
json j_key = json::parse(text, [](int /*depth*/, json::parse_event_t event, json & parsed)
{
return !(event == json::parse_event_t::key && parsed == json("Width"));
});
// discard the top-level object
json j_root = json::parse(text, [](int /*depth*/, json::parse_event_t event, json & /*parsed*/)
{
return event != json::parse_event_t::object_end;
});
// in every case, the discarded value is removed together with its key
std::cout << j_array_start << '\n'
<< j_array_end << '\n'
<< j_value << '\n'
<< j_key << '\n'
<< j_root << '\n';
}
@@ -0,0 +1,5 @@
{"Width":800}
{"Width":800}
{"IDs":[116,943]}
{"IDs":[116,943]}
null
@@ -116,9 +116,19 @@ The library uses the following mapping from JSON values types to BJData types ac
``` ```
Likewise, when a JSON object in the above form is serialized using Likewise, when a JSON object in the above form is serialized using
[`to_bjdata`](../../api/basic_json/to_bjdata.md), it is automatically converted into a compact BJData ND-array. The [`to_bjdata`](../../api/basic_json/to_bjdata.md), it is automatically converted into a compact BJData ND-array. When
only exception is, that when the 1-dimensional vector stored in `"_ArraySize_"` contains a single integer or two the 1-dimensional vector stored in `"_ArraySize_"` contains a single integer or two integers with one being 1, a
integers with one being 1, a regular 1-D optimized array is generated. regular 1-D optimized array is generated instead.
An object is only converted if the annotation actually describes a packed array; otherwise it is serialized as a
regular JSON object. This requires all of the following:
- `"_ArrayType_"` is one of `uint8`, `int8`, `uint16`, `int16`, `uint32`, `int32`, `uint64`, `int64`, `single`,
`double`, `char`, or `byte`,
- every entry of `"_ArraySize_"` is a non-negative integer, and their product is representable as a `std::size_t`,
- `"_ArrayData_"` holds exactly that many elements, and
- every element of `"_ArrayData_"` is a number of the kind named by `"_ArrayType_"` (a floating-point number for
`single` and `double`, an integer otherwise).
The current version of this library does not yet support automatic detection of and conversion from a nested JSON The current version of this library does not yet support automatic detection of and conversion from a nested JSON
array input to a BJData ND-array. array input to a BJData ND-array.
@@ -35,6 +35,19 @@ The library uses the following mapping from JSON values types to BSON types:
The mapping is **incomplete**, since only JSON-objects (and things contained therein) can be serialized to BSON. The mapping is **incomplete**, since only JSON-objects (and things contained therein) can be serialized to BSON.
Also, keys may not contain U+0000, since they are serialized a zero-terminated c-strings. Also, keys may not contain U+0000, since they are serialized a zero-terminated c-strings.
!!! warning "BSON type 0x11 interoperability"
The BSON specification defines type `0x11` as a Timestamp. This library uses marker `0x11` when serializing
`number_unsigned` values in the range `9223372036854775808..18446744073709551615`. Other BSON implementations may
therefore interpret these values as Timestamps instead of unsigned integers.
!!! info "Binary values without a subtype"
BSON requires every binary value to have a subtype. If a binary value has no subtype, this library serializes it
with the generic subtype `0x00`. After deserialization, `has_subtype()` returns `true` and `subtype()` returns `0`.
As a result, serializing and deserializing a JSON object containing such a value produces a different JSON object,
even though the binary data is unchanged.
??? example ??? example
```cpp ```cpp
@@ -82,8 +95,8 @@ The library maps BSON record types to JSON value types as follows:
!!! note "Handling of BSON type 0x11" !!! note "Handling of BSON type 0x11"
BSON type 0x11 is used to represent uint64 numbers. This library treats these values purely as uint64 numbers This library deserializes BSON type `0x11` (Timestamp) as a `number_unsigned` value. The 64-bit value is preserved,
and does not parse them into date-related formats. but the Timestamp type information is not.
??? example ??? example
+30 -7
View File
@@ -66,8 +66,14 @@ auto t = j.get<std::tuple<double, std::string, int>>(); // {1.0, "hello", 42}
std::get<1>(refs) = "world"; // modifies j[1] in place std::get<1>(refs) = "world"; // modifies j[1] in place
``` ```
A referenced type must be one the library actually stores (or an arithmetic type it can convert to/from); A referenced element must name the type the library actually *stores* — one of [`boolean_t`](../api/basic_json/boolean_t.md),
otherwise this is a compile error. [`number_integer_t`](../api/basic_json/number_integer_t.md), [`number_unsigned_t`](../api/basic_json/number_unsigned_t.md),
[`number_float_t`](../api/basic_json/number_float_t.md), [`string_t`](../api/basic_json/string_t.md),
[`binary_t`](../api/basic_json/binary_t.md), [`array_t`](../api/basic_json/array_t.md), or
[`object_t`](../api/basic_json/object_t.md). There is nothing else to refer to, so a reference to any other type is a
compile error even when a conversion would exist: `#!cpp std::tuple<int&>` is rejected, because the library stores a
`#!cpp number_integer_t` (`#!cpp std::int64_t` by default) and not an `#!cpp int`. This restriction applies only to
reference elements — a plain `#!cpp std::tuple<int>` converts by value as usual.
## Implicit conversions ## Implicit conversions
@@ -116,17 +122,34 @@ which forces the explicit `get` form and can catch unintended conversions at com
with a custom `adl_serializer<std::optional<T>>` specialization. Prefer `get<std::optional<T>>()`/`get_to()` with a custom `adl_serializer<std::optional<T>>` specialization. Prefer `get<std::optional<T>>()`/`get_to()`
over `static_cast` for optional types. over `static_cast` for optional types.
!!! warning "Converting to a fixed-size `std::array` does not check length" !!! warning "Converting to a fixed-size destination does not check the array size"
Converting a JSON array to `#!cpp std::array<T, N>` does not check that the JSON array's size matches `N`: Some destination types have a size that is fixed by their C++ type rather than by the JSON value:
if the JSON array is longer, the extra elements are silently dropped; if it is shorter, the remaining `#!cpp std::pair<A, B>`, `#!cpp std::tuple<Ts...>`, `#!cpp std::array<T, N>`, C arrays `#!cpp T[N]`, and
`std::array` elements are left default-constructed. No exception is thrown in either case. `#!cpp std::map`/`#!cpp std::unordered_map` with a non-string key type (which is read from an array of
two-element arrays). All of them read exactly as many elements as they need via
[`at`](../api/basic_json/at.md) and **never compare the JSON array's size to that number**. The two
mismatch directions therefore behave differently:
- The JSON array has **too many** elements: the surplus is **silently discarded**, and no exception is
thrown.
- The JSON array has **too few** elements: `at` throws
[`out_of_range.401`](../home/exceptions.md#jsonexceptionout_of_range401) for the first missing index --
an out-of-range error, not a [`type_error`](../home/exceptions.md#type-errors), even though the cause
is a shape mismatch.
```cpp ```cpp
json j = {1, 2, 3, 4, 5}; json j = {1, 2, 3, 4, 5};
auto a = j.get<std::array<int, 3>>(); // {1, 2, 3} -- elements 4 and 5 silently dropped
auto a = j.get<std::array<int, 3>>(); // {1, 2, 3} -- elements 4 and 5 silently dropped
auto p = j.get<std::pair<int, int>>(); // (1, 2) -- elements 3, 4, and 5 silently dropped
json k = {1};
auto q = k.get<std::pair<int, int>>(); // ❌ throws out_of_range.401
``` ```
If a size mismatch is an error in your application, check the size yourself before converting.
## Omitting a field when serializing `std::optional` ## Omitting a field when serializing `std::optional`
By default, `to_json` for `std::optional<T>` writes either the value or `#!json null` -- there is no built-in way By default, `to_json` for `std::optional<T>` writes either the value or `#!json null` -- there is no built-in way
@@ -53,6 +53,12 @@ If you do want to preserve the **insertion order**, you can use the type [`nlohm
Alternatively, you can use a more sophisticated ordered map like [`tsl::ordered_map`](https://github.com/Tessil/ordered-map) ([integration](https://github.com/nlohmann/json/issues/546#issuecomment-304447518)) or [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) ([integration](https://github.com/nlohmann/json/issues/485#issuecomment-333652309)). Alternatively, you can use a more sophisticated ordered map like [`tsl::ordered_map`](https://github.com/Tessil/ordered-map) ([integration](https://github.com/nlohmann/json/issues/546#issuecomment-304447518)) or [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) ([integration](https://github.com/nlohmann/json/issues/485#issuecomment-333652309)).
The [`ordered_map`](../api/ordered_map.md) behind `nlohmann::ordered_json` is deliberately minimal and has no lookup
index, so every key access is a linear scan and building an object of `n` keys costs O(n²). This is unnoticeable at
typical object sizes but becomes significant for objects with many thousands of keys; see
[`ordered_map` complexity](../api/ordered_map.md#complexity). The alternatives above keep a lookup index and do not
have this cost.
### Notes on parsing ### Notes on parsing
Note that you also need to call the right [`parse`](../api/basic_json/parse.md) function when reading from a file. Note that you also need to call the right [`parse`](../api/basic_json/parse.md) function when reading from a file.
@@ -28,6 +28,22 @@ Inputs consisting of multiple values separated by newlines are handled by the [J
By default, the library rejects comments and trailing commas. Both can be enabled with parameters of the `parse` By default, the library rejects comments and trailing commas. Both can be enabled with parameters of the `parse`
function — see [comments](../comments.md) and [trailing commas](../trailing_commas.md). function — see [comments](../comments.md) and [trailing commas](../trailing_commas.md).
## Strictness and trailing data
[`parse`](../../api/basic_json/parse.md) reads a single JSON value and requires the whole input to be consumed: any
non-whitespace data after the value is reported as a parse error. Use it when you want to guarantee that an input is
exactly one complete JSON document.
[`operator>>`](../../api/operator_gtgt.md) follows relaxed `#!cpp std::istream` semantics instead: it parses one JSON
value and leaves the stream positioned right after it, without requiring the rest of the stream to be consumed. This is
what makes it possible to read several concatenated values from the same stream, but it also means that "a valid
document followed by trailing bytes" is accepted rather than rejected. If you are validating conformance, or need to
reject any input that is not exactly one JSON document, prefer `parse`.
When using `operator>>` to read several concatenated values this way, a value that is a number must be followed by
whitespace, because `operator>>` consumes the character that terminates a number — see the
[`operator>>` notes](../../api/operator_gtgt.md#notes) for details and examples.
## SAX vs. DOM parsing ## SAX vs. DOM parsing
The library offers two parsing models: The library offers two parsing models:
@@ -49,4 +49,5 @@ JSON Lines input with more than one value is treated as invalid JSON by the [`pa
with a JSON Lines input does not work, because the parser will try to parse one value after the last one. with a JSON Lines input does not work, because the parser will try to parse one value after the last one.
This is different from parsing a stream of *concatenated* (non-newline-delimited) JSON values, for which This is different from parsing a stream of *concatenated* (non-newline-delimited) JSON values, for which
`operator>>` does work -- see its [notes](../../api/operator_gtgt.md#notes) for details. `operator>>` does work, provided that a value that is a number is followed by whitespace -- see its
[notes](../../api/operator_gtgt.md#notes) for details.
+24 -5
View File
@@ -291,9 +291,10 @@ A JSON Pointer array index must be a number.
### json.exception.parse_error.110 ### json.exception.parse_error.110
When parsing CBOR or MessagePack, the byte vector ends before the complete value has been read. When parsing a [binary format](../features/binary_formats/index.md), the byte vector ends before the complete value has
been read.
!!! failure "Example message" !!! failure "Example messages"
``` ```
[json.exception.parse_error.110] parse error at byte 5: syntax error while parsing CBOR string: unexpected end of input [json.exception.parse_error.110] parse error at byte 5: syntax error while parsing CBOR string: unexpected end of input
@@ -301,6 +302,9 @@ When parsing CBOR or MessagePack, the byte vector ends before the complete value
``` ```
[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing UBJSON value: expected end of input; last byte: 0x5A [json.exception.parse_error.110] parse error at byte 2: syntax error while parsing UBJSON value: expected end of input; last byte: 0x5A
``` ```
```
[json.exception.parse_error.110] parse error at byte 8: syntax error while parsing BSON number: unexpected end of input
```
### json.exception.parse_error.112 ### json.exception.parse_error.112
@@ -329,10 +333,14 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
``` ```
[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow [json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow
``` ```
```
[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)
```
### json.exception.parse_error.113 ### json.exception.parse_error.113
While parsing a map key, a value that is not a string has been read. A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), or the string's length specification is invalid.
!!! failure "Example messages" !!! failure "Example messages"
@@ -345,6 +353,9 @@ While parsing a map key, a value that is not a string has been read.
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82 [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82
``` ```
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative
```
### json.exception.parse_error.114 ### json.exception.parse_error.114
@@ -853,13 +864,21 @@ and this exception no longer occurs.
### json.exception.out_of_range.408 ### json.exception.out_of_range.408
The size (following `#`) of an UBJSON array or object exceeds the maximal capacity. The size of an array or object in a [binary format](../features/binary_formats/index.md) exceeds the maximal capacity:
the size following `#` for [UBJSON](../features/binary_formats/ubjson.md)/[BJData](../features/binary_formats/bjdata.md),
or the encoded length for [CBOR](../features/binary_formats/cbor.md).
!!! failure "Example message" !!! failure "Example messages"
``` ```
excessive array size: 8658170730974374167 excessive array size: 8658170730974374167
``` ```
```
[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive array size
```
```
[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size
```
### json.exception.out_of_range.409 ### json.exception.out_of_range.409
+1
View File
@@ -308,6 +308,7 @@ nav:
- 'NLOHMANN_JSON_VERSION_MAJOR, NLOHMANN_JSON_VERSION_MINOR, NLOHMANN_JSON_VERSION_PATCH': api/macros/nlohmann_json_version_major.md - 'NLOHMANN_JSON_VERSION_MAJOR, NLOHMANN_JSON_VERSION_MINOR, NLOHMANN_JSON_VERSION_PATCH': api/macros/nlohmann_json_version_major.md
- Community: - Community:
- community/index.md - community/index.md
- community/ecosystem.md
- "Code of Conduct": community/code_of_conduct.md - "Code of Conduct": community/code_of_conduct.md
- community/contribution_guidelines.md - community/contribution_guidelines.md
- community/quality_assurance.md - community/quality_assurance.md
+89 -23
View File
@@ -465,15 +465,6 @@ class binary_reader
// CBOR // // CBOR //
////////// //////////
/*!
@param[in] get_char whether a new character should be retrieved from the
input (true) or whether the last read character should
be considered instead (false)
@param[in] tag_handler how CBOR tags should be treated
@return whether a valid CBOR value was passed to the SAX parser
*/
template<typename NumberType> template<typename NumberType>
bool get_cbor_negative_integer() bool get_cbor_negative_integer()
{ {
@@ -492,6 +483,14 @@ class binary_reader
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number)); return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
} }
/*!
@param[in] get_char whether a new character should be retrieved from the
input (true) or whether the last read character should
be considered instead (false)
@param[in] tag_handler how CBOR tags should be treated
@return whether a valid CBOR value was passed to the SAX parser
*/
bool parse_cbor_internal(const bool get_char, bool parse_cbor_internal(const bool get_char,
const cbor_tag_handler_t tag_handler) const cbor_tag_handler_t tag_handler)
{ {
@@ -704,13 +703,15 @@ class binary_reader
case 0x9A: // array (four-byte uint32_t for n follow) case 0x9A: // array (four-byte uint32_t for n follow)
{ {
std::uint32_t len{}; std::uint32_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_array(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "array") && get_cbor_array(size, tag_handler);
} }
case 0x9B: // array (eight-byte uint64_t for n follow) case 0x9B: // array (eight-byte uint64_t for n follow)
{ {
std::uint64_t len{}; std::uint64_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_array(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "array") && get_cbor_array(size, tag_handler);
} }
case 0x9F: // array (indefinite length) case 0x9F: // array (indefinite length)
@@ -758,13 +759,15 @@ class binary_reader
case 0xBA: // map (four-byte uint32_t for n follow) case 0xBA: // map (four-byte uint32_t for n follow)
{ {
std::uint32_t len{}; std::uint32_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_object(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "map") && get_cbor_object(size, tag_handler);
} }
case 0xBB: // map (eight-byte uint64_t for n follow) case 0xBB: // map (eight-byte uint64_t for n follow)
{ {
std::uint64_t len{}; std::uint64_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_object(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "map") && get_cbor_object(size, tag_handler);
} }
case 0xBF: // map (indefinite length) case 0xBF: // map (indefinite length)
@@ -807,25 +810,37 @@ class binary_reader
case 0xD8: case 0xD8:
{ {
std::uint8_t subtype_to_ignore{}; std::uint8_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
case 0xD9: case 0xD9:
{ {
std::uint16_t subtype_to_ignore{}; std::uint16_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
case 0xDA: case 0xDA:
{ {
std::uint32_t subtype_to_ignore{}; std::uint32_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
case 0xDB: case 0xDB:
{ {
std::uint64_t subtype_to_ignore{}; std::uint64_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
default: default:
@@ -843,28 +858,40 @@ class binary_reader
case 0xD8: case 0xD8:
{ {
std::uint8_t subtype{}; std::uint8_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
case 0xD9: case 0xD9:
{ {
std::uint16_t subtype{}; std::uint16_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
case 0xDA: case 0xDA:
{ {
std::uint32_t subtype{}; std::uint32_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
case 0xDB: case 0xDB:
{ {
std::uint64_t subtype{}; std::uint64_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
@@ -1155,6 +1182,31 @@ class binary_reader
} }
} }
/*!
@brief narrow a definite CBOR array/map length to std::size_t
A definite length is rejected if it does not fit in std::size_t or if it
equals detail::unknown_size(), which is reserved to mark an indefinite-
length container and would otherwise make the length read as indefinite.
Both cases exceed any container's max_size(), so no representable input
is affected.
@param[in] len the declared length
@param[out] result the length narrowed to std::size_t
@param[in] context "array" or "map", for the error message
@return whether the length is usable
*/
bool get_cbor_container_size(const std::uint64_t len, std::size_t& result, const char* context)
{
if (JSON_HEDLEY_UNLIKELY(!value_in_range_of<std::size_t>(len) || len == detail::unknown_size()))
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408,
exception_message(input_format_t::cbor, concat("excessive ", context, " size"), "size"), nullptr));
}
result = conditional_static_cast<std::size_t>(len);
return true;
}
/*! /*!
@param[in] len the length of the array or detail::unknown_size() for an @param[in] len the length of the array or detail::unknown_size() for an
array of indefinite size array of indefinite size
@@ -1935,7 +1987,11 @@ class binary_reader
{ {
if (get_char) if (get_char)
{ {
get(); // TODO(niels): may we ignore N here? // no get_ignore_noop() here: the byte read next must be a string
// length type specification, and a no-op ('N') is not valid in
// that position. No-ops at positions where a value may appear are
// already consumed by the callers via get_ignore_noop().
get();
} }
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format, "value"))) if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format, "value")))
@@ -2825,7 +2881,17 @@ class binary_reader
case token_type::value_unsigned: case token_type::value_unsigned:
return sax->number_unsigned(number_lexer.get_number_unsigned()); return sax->number_unsigned(number_lexer.get_number_unsigned());
case token_type::value_float: case token_type::value_float:
return sax->number_float(number_lexer.get_number_float(), std::move(number_string)); {
const auto parsed_float = number_lexer.get_number_float();
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(parsed_float)))
{
return sax->parse_error(
chars_read,
number_string,
out_of_range::create(406, concat("number overflow parsing '", number_string, '\''), nullptr));
}
return sax->number_float(parsed_float, std::move(number_string));
}
case token_type::uninitialized: case token_type::uninitialized:
case token_type::literal_true: case token_type::literal_true:
case token_type::literal_false: case token_type::literal_false:
@@ -345,8 +345,12 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
} }
else else
{ {
// unknown character // A code point above U+10FFFF has no UTF-8 encoding. Passing the
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc); // unit through would narrow it to int, where 0xFFFFFFFF becomes
// char_traits<char>::eof() and would end the input silently, so
// emit a byte that is never valid UTF-8 and let the decoder
// reject it.
utf8_bytes[0] = 0xFF;
utf8_bytes_filled = 1; utf8_bytes_filled = 1;
} }
} }
+48 -15
View File
@@ -370,8 +370,10 @@ class json_sax_dom_parser
case value_t::string: case value_t::string:
{ {
// include the length of the quotes, which is 2 // escape sequences make the token longer than the value it
v.start_position = v.end_position - v.m_data.m_value.string->size() - 2; // parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = m_lexer_ref->get_token_start_position();
break; break;
} }
@@ -626,14 +628,7 @@ class json_sax_dom_callback_parser
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_structured()) if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_structured())
{ {
// remove discarded value // remove discarded value
for (auto it = ref_stack.back()->begin(); it != ref_stack.back()->end(); ++it) remove_discarded_value(*ref_stack.back());
{
if (it->is_discarded())
{
ref_stack.back()->erase(it);
break;
}
}
} }
return true; return true;
@@ -674,8 +669,9 @@ class json_sax_dom_callback_parser
bool end_array() bool end_array()
{ {
bool keep = true; bool keep = true;
const bool stored = ref_stack.back() != nullptr;
if (ref_stack.back()) if (stored)
{ {
keep = callback(static_cast<int>(ref_stack.size()) - 1, parse_event_t::array_end, *ref_stack.back()); keep = callback(static_cast<int>(ref_stack.size()) - 1, parse_event_t::array_end, *ref_stack.back());
if (keep) if (keep)
@@ -709,9 +705,19 @@ class json_sax_dom_callback_parser
keep_stack.pop_back(); keep_stack.pop_back();
// remove discarded value // remove discarded value
if (!keep && !ref_stack.empty() && ref_stack.back()->is_array()) if (!ref_stack.empty() && ref_stack.back())
{ {
ref_stack.back()->m_data.m_value.array->pop_back(); if (!keep && ref_stack.back()->is_array())
{
ref_stack.back()->m_data.m_value.array->pop_back();
}
else if ((!keep || !stored) && ref_stack.back()->is_object())
{
// the array is either still stored under its key or was never
// stored, leaving the placeholder key() wrote; both show up as
// a discarded member of the parent object
remove_discarded_value(*ref_stack.back());
}
} }
return true; return true;
@@ -765,8 +771,10 @@ class json_sax_dom_callback_parser
case value_t::string: case value_t::string:
{ {
// include the length of the quotes, which is 2 // escape sequences make the token longer than the value it
v.start_position = v.end_position - v.m_data.m_value.string->size() - 2; // parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = m_lexer_ref->get_token_start_position();
break; break;
} }
@@ -801,6 +809,19 @@ class json_sax_dom_callback_parser
} }
#endif #endif
/// remove the discarded value the callback rejected from its parent
static void remove_discarded_value(BasicJsonType& parent)
{
for (auto it = parent.begin(); it != parent.end(); ++it)
{
if (it->is_discarded())
{
parent.erase(it);
break;
}
}
}
/*! /*!
@param[in] v value to add to the JSON value we build during parsing @param[in] v value to add to the JSON value we build during parsing
@param[in] skip_callback whether we should skip calling the callback @param[in] skip_callback whether we should skip calling the callback
@@ -841,6 +862,18 @@ class json_sax_dom_callback_parser
// do not handle this value if we just learnt it shall be discarded // do not handle this value if we just learnt it shall be discarded
if (!keep) if (!keep)
{ {
// if the value was to become an object member, key() already
// stored a placeholder for it that has to be removed again
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_object())
{
JSON_ASSERT(!key_keep_stack.empty());
const bool placeholder_stored = key_keep_stack.back();
key_keep_stack.pop_back();
if (placeholder_stored)
{
remove_discarded_value(*ref_stack.back());
}
}
return {false, nullptr}; return {false, nullptr};
} }
+20
View File
@@ -1357,6 +1357,11 @@ scan_number_done:
token_buffer.clear(); token_buffer.clear();
decimal_point_position = std::string::npos; decimal_point_position = std::string::npos;
#if JSON_DIAGNOSTIC_POSITIONS
// the first character of the token has already been read, hence the -1
token_start_position = position.chars_read_total - 1;
#endif
note_token_start(std::integral_constant<bool, lazy_token_string> {}); note_token_start(std::integral_constant<bool, lazy_token_string> {});
} }
@@ -1519,6 +1524,15 @@ scan_number_done:
return position; return position;
} }
#if JSON_DIAGNOSTIC_POSITIONS
/// return the offset of the first character of the last read token; unlike
/// the token's parsed value, this accounts for escape sequences
constexpr std::size_t get_token_start_position() const noexcept
{
return token_start_position;
}
#endif
/// seekable adapter: rebuild the last read token from the input on demand /// seekable adapter: rebuild the last read token from the input on demand
const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const
{ {
@@ -1719,6 +1733,12 @@ scan_number_done:
/// the last read token on error for seekable adapters (see collect_token_chars) /// the last read token on error for seekable adapters (see collect_token_chars)
std::size_t token_string_start = 0; std::size_t token_string_start = 0;
#if JSON_DIAGNOSTIC_POSITIONS
/// start offset of the current token within the input, used to report
/// diagnostic positions (see reset())
std::size_t token_start_position = 0;
#endif
/// buffer for variable-length tokens (numbers, strings) /// buffer for variable-length tokens (numbers, strings)
string_t token_buffer {}; string_t token_buffer {};
+26 -10
View File
@@ -231,7 +231,9 @@ struct char_traits<signed char> : std::char_traits<char>
// Redefine to_int_type function // Redefine to_int_type function
static int_type to_int_type(char_type c) noexcept static int_type to_int_type(char_type c) noexcept
{ {
return static_cast<int_type>(c); // cast via unsigned char: sign-extending a negative char_type would make
// byte 0xFF indistinguishable from eof()
return static_cast<int_type>(static_cast<unsigned char>(c));
} }
static char_type to_char_type(int_type i) noexcept static char_type to_char_type(int_type i) noexcept
@@ -699,21 +701,35 @@ struct is_json_pointer_of<A, ::nlohmann::json_pointer<A>> : std::true_type {};
template <typename A> template <typename A>
struct is_json_pointer_of<A, ::nlohmann::json_pointer<A>&> : std::true_type {}; struct is_json_pointer_of<A, ::nlohmann::json_pointer<A>&> : std::true_type {};
// checks if A and B are comparable using Compare functor // checks if A and B are comparable using Compare functor, assuming that
// neither A nor B is a json_pointer type (that case is handled by
// is_comparable below, which never instantiates this helper otherwise)
template<typename Compare, typename A, typename B, typename = void> template<typename Compare, typename A, typename B, typename = void>
struct is_comparable : std::false_type {}; struct is_comparable_no_json_pointer : std::false_type {};
// We exclude json_pointer here, because the checks using Compare(A, B) will
// use json_pointer::operator string_t() which triggers a deprecation warning
// for GCC. See https://github.com/nlohmann/json/issues/4621. The call to
// is_json_pointer_of can be removed once the deprecated function has been
// removed.
template<typename Compare, typename A, typename B> template<typename Compare, typename A, typename B>
struct is_comparable < Compare, A, B, enable_if_t < !is_json_pointer_of<A, B>::value struct is_comparable_no_json_pointer < Compare, A, B, enable_if_t <
&& std::is_constructible <decltype(std::declval<Compare>()(std::declval<A>(), std::declval<B>()))>::value std::is_constructible <decltype(std::declval<Compare>()(std::declval<A>(), std::declval<B>()))>::value
&& std::is_constructible <decltype(std::declval<Compare>()(std::declval<B>(), std::declval<A>()))>::value && std::is_constructible <decltype(std::declval<Compare>()(std::declval<B>(), std::declval<A>()))>::value
>> : std::true_type {}; >> : std::true_type {};
// checks if A and B are comparable using Compare functor
// We dispatch on is_json_pointer_of as a plain bool (rather than folding it
// into a single enable_if_t condition together with the checks below) so
// that the Compare(A, B) checks are only ever written - and thus only ever
// instantiated - when A/B are not a json_pointer/string pair. Those checks
// use json_pointer::operator string_t() (GCC, see #4621) resp. the
// deprecated json_pointer/string operator== (Clang, see #5288), and merely
// naming them as later operands of a plain && chain is not sufficient to
// avoid their instantiation on all compilers, even when the first operand
// is false. The dispatch on is_json_pointer_of can be removed once the
// deprecated json_pointer comparison operators have been removed.
template<typename Compare, typename A, typename B, bool = is_json_pointer_of<A, B>::value>
struct is_comparable : std::false_type {};
template<typename Compare, typename A, typename B>
struct is_comparable<Compare, A, B, false> : is_comparable_no_json_pointer<Compare, A, B> {};
template<typename T> template<typename T>
using detect_is_transparent = typename T::is_transparent; using detect_is_transparent = typename T::is_transparent;
@@ -1662,7 +1662,31 @@ class binary_writer
std::size_t len = (value.at(key).empty() ? 0 : 1); std::size_t len = (value.at(key).empty() ? 0 : 1);
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
len *= static_cast<std::size_t>(el.m_data.m_value.number_unsigned); // a dimension is read as an unsigned value below, so anything that
// is not a non-negative integer is rejected: a non-integer entry
// would pun unrelated bytes as the dimension, and a negative one
// would wrap into a nonsensical length
if (!el.is_number_integer() || (!el.is_number_unsigned() && el.template get<std::int64_t>() < 0))
{
return true;
}
// a dimension that does not fit into std::size_t, or a product that
// overflows it, would wrap around and could match the size of
// _ArrayData_ by accident; the resulting header announces an
// element count that no reader can honor (the binary reader rejects
// it with out_of_range.408), so encode as a plain object instead
const auto dim = el.template get<std::uint64_t>();
if (!value_in_range_of<std::size_t>(dim))
{
return true;
}
const auto dim_size = static_cast<std::size_t>(dim);
if (dim_size != 0 && len > (std::numeric_limits<std::size_t>::max)() / dim_size)
{
return true;
}
len *= dim_size;
} }
key = "_ArrayData_"; key = "_ArrayData_";
@@ -1671,6 +1695,24 @@ class binary_writer
return true; return true;
} }
// every element is written below as the number kind dtype names, so it
// has to actually be a number of that category: an element of any other
// type would reinterpret unrelated bytes, e.g. a string's heap pointer,
// as that number. Such an object falls back to a plain object encoding.
// dtype names the wire type, not the storage type: whether an integer
// is held as number_integer or number_unsigned depends on how the value
// was built (parsing stores non-negative integers as unsigned, the C++
// API stores int literals as signed), so both are accepted here and the
// writes below go through get<>, which reads the member that is active.
const bool ndarray_is_float = (dtype == 'd' || dtype == 'D');
for (const auto& el : value.at(key))
{
if (ndarray_is_float ? !el.is_number_float() : !el.is_number_integer())
{
return true;
}
}
oa->write_character('['); oa->write_character('[');
oa->write_character('$'); oa->write_character('$');
oa->write_character(dtype); oa->write_character(dtype);
@@ -1684,70 +1726,70 @@ class binary_writer
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint8_t>(el.m_data.m_value.number_unsigned), true); write_number(static_cast<std::uint8_t>(el.template get<std::uint64_t>()), true);
} }
} }
else if (dtype == 'i') else if (dtype == 'i')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int8_t>(el.m_data.m_value.number_integer), true); write_number(static_cast<std::int8_t>(el.template get<std::int64_t>()), true);
} }
} }
else if (dtype == 'u') else if (dtype == 'u')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint16_t>(el.m_data.m_value.number_unsigned), true); write_number(static_cast<std::uint16_t>(el.template get<std::uint64_t>()), true);
} }
} }
else if (dtype == 'I') else if (dtype == 'I')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int16_t>(el.m_data.m_value.number_integer), true); write_number(static_cast<std::int16_t>(el.template get<std::int64_t>()), true);
} }
} }
else if (dtype == 'm') else if (dtype == 'm')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint32_t>(el.m_data.m_value.number_unsigned), true); write_number(static_cast<std::uint32_t>(el.template get<std::uint64_t>()), true);
} }
} }
else if (dtype == 'l') else if (dtype == 'l')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int32_t>(el.m_data.m_value.number_integer), true); write_number(static_cast<std::int32_t>(el.template get<std::int64_t>()), true);
} }
} }
else if (dtype == 'M') else if (dtype == 'M')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint64_t>(el.m_data.m_value.number_unsigned), true); write_number(el.template get<std::uint64_t>(), true);
} }
} }
else if (dtype == 'L') else if (dtype == 'L')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int64_t>(el.m_data.m_value.number_integer), true); write_number(el.template get<std::int64_t>(), true);
} }
} }
else if (dtype == 'd') else if (dtype == 'd')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<float>(el.m_data.m_value.number_float), true); write_number(static_cast<float>(el.template get<double>()), true);
} }
} }
else if (dtype == 'D') else if (dtype == 'D')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<double>(el.m_data.m_value.number_float), true); write_number(el.template get<double>(), true);
} }
} }
return false; return false;
+242 -61
View File
@@ -3994,7 +3994,9 @@ struct char_traits<signed char> : std::char_traits<char>
// Redefine to_int_type function // Redefine to_int_type function
static int_type to_int_type(char_type c) noexcept static int_type to_int_type(char_type c) noexcept
{ {
return static_cast<int_type>(c); // cast via unsigned char: sign-extending a negative char_type would make
// byte 0xFF indistinguishable from eof()
return static_cast<int_type>(static_cast<unsigned char>(c));
} }
static char_type to_char_type(int_type i) noexcept static char_type to_char_type(int_type i) noexcept
@@ -4462,21 +4464,35 @@ struct is_json_pointer_of<A, ::nlohmann::json_pointer<A>> : std::true_type {};
template <typename A> template <typename A>
struct is_json_pointer_of<A, ::nlohmann::json_pointer<A>&> : std::true_type {}; struct is_json_pointer_of<A, ::nlohmann::json_pointer<A>&> : std::true_type {};
// checks if A and B are comparable using Compare functor // checks if A and B are comparable using Compare functor, assuming that
// neither A nor B is a json_pointer type (that case is handled by
// is_comparable below, which never instantiates this helper otherwise)
template<typename Compare, typename A, typename B, typename = void> template<typename Compare, typename A, typename B, typename = void>
struct is_comparable : std::false_type {}; struct is_comparable_no_json_pointer : std::false_type {};
// We exclude json_pointer here, because the checks using Compare(A, B) will
// use json_pointer::operator string_t() which triggers a deprecation warning
// for GCC. See https://github.com/nlohmann/json/issues/4621. The call to
// is_json_pointer_of can be removed once the deprecated function has been
// removed.
template<typename Compare, typename A, typename B> template<typename Compare, typename A, typename B>
struct is_comparable < Compare, A, B, enable_if_t < !is_json_pointer_of<A, B>::value struct is_comparable_no_json_pointer < Compare, A, B, enable_if_t <
&& std::is_constructible <decltype(std::declval<Compare>()(std::declval<A>(), std::declval<B>()))>::value std::is_constructible <decltype(std::declval<Compare>()(std::declval<A>(), std::declval<B>()))>::value
&& std::is_constructible <decltype(std::declval<Compare>()(std::declval<B>(), std::declval<A>()))>::value && std::is_constructible <decltype(std::declval<Compare>()(std::declval<B>(), std::declval<A>()))>::value
>> : std::true_type {}; >> : std::true_type {};
// checks if A and B are comparable using Compare functor
// We dispatch on is_json_pointer_of as a plain bool (rather than folding it
// into a single enable_if_t condition together with the checks below) so
// that the Compare(A, B) checks are only ever written - and thus only ever
// instantiated - when A/B are not a json_pointer/string pair. Those checks
// use json_pointer::operator string_t() (GCC, see #4621) resp. the
// deprecated json_pointer/string operator== (Clang, see #5288), and merely
// naming them as later operands of a plain && chain is not sufficient to
// avoid their instantiation on all compilers, even when the first operand
// is false. The dispatch on is_json_pointer_of can be removed once the
// deprecated json_pointer comparison operators have been removed.
template<typename Compare, typename A, typename B, bool = is_json_pointer_of<A, B>::value>
struct is_comparable : std::false_type {};
template<typename Compare, typename A, typename B>
struct is_comparable<Compare, A, B, false> : is_comparable_no_json_pointer<Compare, A, B> {};
template<typename T> template<typename T>
using detect_is_transparent = typename T::is_transparent; using detect_is_transparent = typename T::is_transparent;
@@ -7332,8 +7348,12 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
} }
else else
{ {
// unknown character // A code point above U+10FFFF has no UTF-8 encoding. Passing the
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc); // unit through would narrow it to int, where 0xFFFFFFFF becomes
// char_traits<char>::eof() and would end the input silently, so
// emit a byte that is never valid UTF-8 and let the decoder
// reject it.
utf8_bytes[0] = 0xFF;
utf8_bytes_filled = 1; utf8_bytes_filled = 1;
} }
} }
@@ -9055,6 +9075,11 @@ scan_number_done:
token_buffer.clear(); token_buffer.clear();
decimal_point_position = std::string::npos; decimal_point_position = std::string::npos;
#if JSON_DIAGNOSTIC_POSITIONS
// the first character of the token has already been read, hence the -1
token_start_position = position.chars_read_total - 1;
#endif
note_token_start(std::integral_constant<bool, lazy_token_string> {}); note_token_start(std::integral_constant<bool, lazy_token_string> {});
} }
@@ -9217,6 +9242,15 @@ scan_number_done:
return position; return position;
} }
#if JSON_DIAGNOSTIC_POSITIONS
/// return the offset of the first character of the last read token; unlike
/// the token's parsed value, this accounts for escape sequences
constexpr std::size_t get_token_start_position() const noexcept
{
return token_start_position;
}
#endif
/// seekable adapter: rebuild the last read token from the input on demand /// seekable adapter: rebuild the last read token from the input on demand
const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const
{ {
@@ -9417,6 +9451,12 @@ scan_number_done:
/// the last read token on error for seekable adapters (see collect_token_chars) /// the last read token on error for seekable adapters (see collect_token_chars)
std::size_t token_string_start = 0; std::size_t token_string_start = 0;
#if JSON_DIAGNOSTIC_POSITIONS
/// start offset of the current token within the input, used to report
/// diagnostic positions (see reset())
std::size_t token_start_position = 0;
#endif
/// buffer for variable-length tokens (numbers, strings) /// buffer for variable-length tokens (numbers, strings)
string_t token_buffer {}; string_t token_buffer {};
@@ -9793,8 +9833,10 @@ class json_sax_dom_parser
case value_t::string: case value_t::string:
{ {
// include the length of the quotes, which is 2 // escape sequences make the token longer than the value it
v.start_position = v.end_position - v.m_data.m_value.string->size() - 2; // parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = m_lexer_ref->get_token_start_position();
break; break;
} }
@@ -10049,14 +10091,7 @@ class json_sax_dom_callback_parser
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_structured()) if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_structured())
{ {
// remove discarded value // remove discarded value
for (auto it = ref_stack.back()->begin(); it != ref_stack.back()->end(); ++it) remove_discarded_value(*ref_stack.back());
{
if (it->is_discarded())
{
ref_stack.back()->erase(it);
break;
}
}
} }
return true; return true;
@@ -10097,8 +10132,9 @@ class json_sax_dom_callback_parser
bool end_array() bool end_array()
{ {
bool keep = true; bool keep = true;
const bool stored = ref_stack.back() != nullptr;
if (ref_stack.back()) if (stored)
{ {
keep = callback(static_cast<int>(ref_stack.size()) - 1, parse_event_t::array_end, *ref_stack.back()); keep = callback(static_cast<int>(ref_stack.size()) - 1, parse_event_t::array_end, *ref_stack.back());
if (keep) if (keep)
@@ -10132,9 +10168,19 @@ class json_sax_dom_callback_parser
keep_stack.pop_back(); keep_stack.pop_back();
// remove discarded value // remove discarded value
if (!keep && !ref_stack.empty() && ref_stack.back()->is_array()) if (!ref_stack.empty() && ref_stack.back())
{ {
ref_stack.back()->m_data.m_value.array->pop_back(); if (!keep && ref_stack.back()->is_array())
{
ref_stack.back()->m_data.m_value.array->pop_back();
}
else if ((!keep || !stored) && ref_stack.back()->is_object())
{
// the array is either still stored under its key or was never
// stored, leaving the placeholder key() wrote; both show up as
// a discarded member of the parent object
remove_discarded_value(*ref_stack.back());
}
} }
return true; return true;
@@ -10188,8 +10234,10 @@ class json_sax_dom_callback_parser
case value_t::string: case value_t::string:
{ {
// include the length of the quotes, which is 2 // escape sequences make the token longer than the value it
v.start_position = v.end_position - v.m_data.m_value.string->size() - 2; // parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = m_lexer_ref->get_token_start_position();
break; break;
} }
@@ -10224,6 +10272,19 @@ class json_sax_dom_callback_parser
} }
#endif #endif
/// remove the discarded value the callback rejected from its parent
static void remove_discarded_value(BasicJsonType& parent)
{
for (auto it = parent.begin(); it != parent.end(); ++it)
{
if (it->is_discarded())
{
parent.erase(it);
break;
}
}
}
/*! /*!
@param[in] v value to add to the JSON value we build during parsing @param[in] v value to add to the JSON value we build during parsing
@param[in] skip_callback whether we should skip calling the callback @param[in] skip_callback whether we should skip calling the callback
@@ -10264,6 +10325,18 @@ class json_sax_dom_callback_parser
// do not handle this value if we just learnt it shall be discarded // do not handle this value if we just learnt it shall be discarded
if (!keep) if (!keep)
{ {
// if the value was to become an object member, key() already
// stored a placeholder for it that has to be removed again
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_object())
{
JSON_ASSERT(!key_keep_stack.empty());
const bool placeholder_stored = key_keep_stack.back();
key_keep_stack.pop_back();
if (placeholder_stored)
{
remove_discarded_value(*ref_stack.back());
}
}
return {false, nullptr}; return {false, nullptr};
} }
@@ -11014,15 +11087,6 @@ class binary_reader
// CBOR // // CBOR //
////////// //////////
/*!
@param[in] get_char whether a new character should be retrieved from the
input (true) or whether the last read character should
be considered instead (false)
@param[in] tag_handler how CBOR tags should be treated
@return whether a valid CBOR value was passed to the SAX parser
*/
template<typename NumberType> template<typename NumberType>
bool get_cbor_negative_integer() bool get_cbor_negative_integer()
{ {
@@ -11041,6 +11105,14 @@ class binary_reader
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number)); return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
} }
/*!
@param[in] get_char whether a new character should be retrieved from the
input (true) or whether the last read character should
be considered instead (false)
@param[in] tag_handler how CBOR tags should be treated
@return whether a valid CBOR value was passed to the SAX parser
*/
bool parse_cbor_internal(const bool get_char, bool parse_cbor_internal(const bool get_char,
const cbor_tag_handler_t tag_handler) const cbor_tag_handler_t tag_handler)
{ {
@@ -11253,13 +11325,15 @@ class binary_reader
case 0x9A: // array (four-byte uint32_t for n follow) case 0x9A: // array (four-byte uint32_t for n follow)
{ {
std::uint32_t len{}; std::uint32_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_array(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "array") && get_cbor_array(size, tag_handler);
} }
case 0x9B: // array (eight-byte uint64_t for n follow) case 0x9B: // array (eight-byte uint64_t for n follow)
{ {
std::uint64_t len{}; std::uint64_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_array(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "array") && get_cbor_array(size, tag_handler);
} }
case 0x9F: // array (indefinite length) case 0x9F: // array (indefinite length)
@@ -11307,13 +11381,15 @@ class binary_reader
case 0xBA: // map (four-byte uint32_t for n follow) case 0xBA: // map (four-byte uint32_t for n follow)
{ {
std::uint32_t len{}; std::uint32_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_object(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "map") && get_cbor_object(size, tag_handler);
} }
case 0xBB: // map (eight-byte uint64_t for n follow) case 0xBB: // map (eight-byte uint64_t for n follow)
{ {
std::uint64_t len{}; std::uint64_t len{};
return get_number(input_format_t::cbor, len) && get_cbor_object(conditional_static_cast<std::size_t>(len), tag_handler); std::size_t size{};
return get_number(input_format_t::cbor, len) && get_cbor_container_size(len, size, "map") && get_cbor_object(size, tag_handler);
} }
case 0xBF: // map (indefinite length) case 0xBF: // map (indefinite length)
@@ -11356,25 +11432,37 @@ class binary_reader
case 0xD8: case 0xD8:
{ {
std::uint8_t subtype_to_ignore{}; std::uint8_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
case 0xD9: case 0xD9:
{ {
std::uint16_t subtype_to_ignore{}; std::uint16_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
case 0xDA: case 0xDA:
{ {
std::uint32_t subtype_to_ignore{}; std::uint32_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
case 0xDB: case 0xDB:
{ {
std::uint64_t subtype_to_ignore{}; std::uint64_t subtype_to_ignore{};
get_number(input_format_t::cbor, subtype_to_ignore); if (!get_number(input_format_t::cbor, subtype_to_ignore))
{
return false;
}
break; break;
} }
default: default:
@@ -11392,28 +11480,40 @@ class binary_reader
case 0xD8: case 0xD8:
{ {
std::uint8_t subtype{}; std::uint8_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
case 0xD9: case 0xD9:
{ {
std::uint16_t subtype{}; std::uint16_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
case 0xDA: case 0xDA:
{ {
std::uint32_t subtype{}; std::uint32_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
case 0xDB: case 0xDB:
{ {
std::uint64_t subtype{}; std::uint64_t subtype{};
get_number(input_format_t::cbor, subtype); if (!get_number(input_format_t::cbor, subtype))
{
return false;
}
b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype)); b.set_subtype(detail::conditional_static_cast<typename binary_t::subtype_type>(subtype));
break; break;
} }
@@ -11704,6 +11804,31 @@ class binary_reader
} }
} }
/*!
@brief narrow a definite CBOR array/map length to std::size_t
A definite length is rejected if it does not fit in std::size_t or if it
equals detail::unknown_size(), which is reserved to mark an indefinite-
length container and would otherwise make the length read as indefinite.
Both cases exceed any container's max_size(), so no representable input
is affected.
@param[in] len the declared length
@param[out] result the length narrowed to std::size_t
@param[in] context "array" or "map", for the error message
@return whether the length is usable
*/
bool get_cbor_container_size(const std::uint64_t len, std::size_t& result, const char* context)
{
if (JSON_HEDLEY_UNLIKELY(!value_in_range_of<std::size_t>(len) || len == detail::unknown_size()))
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408,
exception_message(input_format_t::cbor, concat("excessive ", context, " size"), "size"), nullptr));
}
result = conditional_static_cast<std::size_t>(len);
return true;
}
/*! /*!
@param[in] len the length of the array or detail::unknown_size() for an @param[in] len the length of the array or detail::unknown_size() for an
array of indefinite size array of indefinite size
@@ -12484,7 +12609,11 @@ class binary_reader
{ {
if (get_char) if (get_char)
{ {
get(); // TODO(niels): may we ignore N here? // no get_ignore_noop() here: the byte read next must be a string
// length type specification, and a no-op ('N') is not valid in
// that position. No-ops at positions where a value may appear are
// already consumed by the callers via get_ignore_noop().
get();
} }
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format, "value"))) if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format, "value")))
@@ -13374,7 +13503,17 @@ class binary_reader
case token_type::value_unsigned: case token_type::value_unsigned:
return sax->number_unsigned(number_lexer.get_number_unsigned()); return sax->number_unsigned(number_lexer.get_number_unsigned());
case token_type::value_float: case token_type::value_float:
return sax->number_float(number_lexer.get_number_float(), std::move(number_string)); {
const auto parsed_float = number_lexer.get_number_float();
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(parsed_float)))
{
return sax->parse_error(
chars_read,
number_string,
out_of_range::create(406, concat("number overflow parsing '", number_string, '\''), nullptr));
}
return sax->number_float(parsed_float, std::move(number_string));
}
case token_type::uninitialized: case token_type::uninitialized:
case token_type::literal_true: case token_type::literal_true:
case token_type::literal_false: case token_type::literal_false:
@@ -18457,7 +18596,31 @@ class binary_writer
std::size_t len = (value.at(key).empty() ? 0 : 1); std::size_t len = (value.at(key).empty() ? 0 : 1);
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
len *= static_cast<std::size_t>(el.m_data.m_value.number_unsigned); // a dimension is read as an unsigned value below, so anything that
// is not a non-negative integer is rejected: a non-integer entry
// would pun unrelated bytes as the dimension, and a negative one
// would wrap into a nonsensical length
if (!el.is_number_integer() || (!el.is_number_unsigned() && el.template get<std::int64_t>() < 0))
{
return true;
}
// a dimension that does not fit into std::size_t, or a product that
// overflows it, would wrap around and could match the size of
// _ArrayData_ by accident; the resulting header announces an
// element count that no reader can honor (the binary reader rejects
// it with out_of_range.408), so encode as a plain object instead
const auto dim = el.template get<std::uint64_t>();
if (!value_in_range_of<std::size_t>(dim))
{
return true;
}
const auto dim_size = static_cast<std::size_t>(dim);
if (dim_size != 0 && len > (std::numeric_limits<std::size_t>::max)() / dim_size)
{
return true;
}
len *= dim_size;
} }
key = "_ArrayData_"; key = "_ArrayData_";
@@ -18466,6 +18629,24 @@ class binary_writer
return true; return true;
} }
// every element is written below as the number kind dtype names, so it
// has to actually be a number of that category: an element of any other
// type would reinterpret unrelated bytes, e.g. a string's heap pointer,
// as that number. Such an object falls back to a plain object encoding.
// dtype names the wire type, not the storage type: whether an integer
// is held as number_integer or number_unsigned depends on how the value
// was built (parsing stores non-negative integers as unsigned, the C++
// API stores int literals as signed), so both are accepted here and the
// writes below go through get<>, which reads the member that is active.
const bool ndarray_is_float = (dtype == 'd' || dtype == 'D');
for (const auto& el : value.at(key))
{
if (ndarray_is_float ? !el.is_number_float() : !el.is_number_integer())
{
return true;
}
}
oa->write_character('['); oa->write_character('[');
oa->write_character('$'); oa->write_character('$');
oa->write_character(dtype); oa->write_character(dtype);
@@ -18479,70 +18660,70 @@ class binary_writer
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint8_t>(el.m_data.m_value.number_unsigned), true); write_number(static_cast<std::uint8_t>(el.template get<std::uint64_t>()), true);
} }
} }
else if (dtype == 'i') else if (dtype == 'i')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int8_t>(el.m_data.m_value.number_integer), true); write_number(static_cast<std::int8_t>(el.template get<std::int64_t>()), true);
} }
} }
else if (dtype == 'u') else if (dtype == 'u')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint16_t>(el.m_data.m_value.number_unsigned), true); write_number(static_cast<std::uint16_t>(el.template get<std::uint64_t>()), true);
} }
} }
else if (dtype == 'I') else if (dtype == 'I')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int16_t>(el.m_data.m_value.number_integer), true); write_number(static_cast<std::int16_t>(el.template get<std::int64_t>()), true);
} }
} }
else if (dtype == 'm') else if (dtype == 'm')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint32_t>(el.m_data.m_value.number_unsigned), true); write_number(static_cast<std::uint32_t>(el.template get<std::uint64_t>()), true);
} }
} }
else if (dtype == 'l') else if (dtype == 'l')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int32_t>(el.m_data.m_value.number_integer), true); write_number(static_cast<std::int32_t>(el.template get<std::int64_t>()), true);
} }
} }
else if (dtype == 'M') else if (dtype == 'M')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::uint64_t>(el.m_data.m_value.number_unsigned), true); write_number(el.template get<std::uint64_t>(), true);
} }
} }
else if (dtype == 'L') else if (dtype == 'L')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<std::int64_t>(el.m_data.m_value.number_integer), true); write_number(el.template get<std::int64_t>(), true);
} }
} }
else if (dtype == 'd') else if (dtype == 'd')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<float>(el.m_data.m_value.number_float), true); write_number(static_cast<float>(el.template get<double>()), true);
} }
} }
else if (dtype == 'D') else if (dtype == 'D')
{ {
for (const auto& el : value.at(key)) for (const auto& el : value.at(key))
{ {
write_number(static_cast<double>(el.m_data.m_value.number_float), true); write_number(el.template get<double>(), true);
} }
} }
return false; return false;
+34
View File
@@ -132,3 +132,37 @@ TEST_CASE("BJData")
} }
} }
} }
TEST_CASE("CBOR")
{
SECTION("parse errors")
{
SECTION("array/map size larger than std::size_t")
{
// declared lengths do not fit in a 32-bit std::size_t and must not be truncated
std::vector<uint8_t> const varr = {0x9B, 0x00, 0x00, 0x00, 0x01, 0x00, 0x00, 0x00, 0x05};
std::vector<uint8_t> const vmap = {0xBB, 0x00, 0x00, 0x00, 0x01, 0x00, 0x00, 0x00, 0x05};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(varr), "[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive array size", json::out_of_range&);
CHECK(json::from_cbor(varr, true, false).is_discarded());
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vmap), "[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size", json::out_of_range&);
CHECK(json::from_cbor(vmap, true, false).is_discarded());
}
SECTION("array/map size equal to the indefinite-length sentinel")
{
// on 32-bit platforms a four-byte length of 0xFFFFFFFF aliases unknown_size()
std::vector<uint8_t> const varr = {0x9A, 0xFF, 0xFF, 0xFF, 0xFF};
std::vector<uint8_t> const vmap = {0xBA, 0xFF, 0xFF, 0xFF, 0xFF};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(varr), "[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive array size", json::out_of_range&);
CHECK(json::from_cbor(varr, true, false).is_discarded());
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vmap), "[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size", json::out_of_range&);
CHECK(json::from_cbor(vmap, true, false).is_discarded());
}
}
}
+86
View File
@@ -1347,6 +1347,8 @@ TEST_CASE("BJData")
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec2), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing BJData high-precision number: invalid number text: 1A", json::parse_error); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec2), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing BJData high-precision number: invalid number text: 1A", json::parse_error);
std::vector<uint8_t> const vec3 = {'H', 'i', 2, '1', '.'}; std::vector<uint8_t> const vec3 = {'H', 'i', 2, '1', '.'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec3), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing BJData high-precision number: invalid number text: 1.", json::parse_error); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec3), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing BJData high-precision number: invalid number text: 1.", json::parse_error);
std::vector<uint8_t> const vec_overflow = {'H', 'i', 5, '1', 'e', '4', '0', '0'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec_overflow), "[json.exception.out_of_range.406] number overflow parsing '1e400'", json::out_of_range);
std::vector<uint8_t> const vec4 = {'H', 2, '1', '0'}; std::vector<uint8_t> const vec4 = {'H', 2, '1', '0'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec4), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing BJData size: expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x02", json::parse_error); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vec4), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing BJData size: expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x02", json::parse_error);
} }
@@ -2587,6 +2589,69 @@ TEST_CASE("BJData")
CHECK(json::to_bjdata(json::from_bjdata(v_B), true, true) == v_B); CHECK(json::to_bjdata(json::from_bjdata(v_B), true, true) == v_B);
} }
SECTION("ndarray with data not matching _ArrayType_ is written as an object")
{
// A JData-annotated object is only serialized as an ndarray when
// its _ArrayData_ elements are actually stored as the number kind
// named by _ArrayType_. Otherwise the writer would read the wrong
// union member (e.g. a std::string's heap pointer as a uint64) and
// emit it, so such an object falls back to a plain object encoding
// that still round-trips.
// string data declared as a uint64 array
json const j_str = json({{"_ArrayType_", "uint64"}, {"_ArraySize_", {1}}, {"_ArrayData_", {"pointer"}}});
const auto out_str = json::to_bjdata(j_str);
CHECK(out_str.at(0) == '{');
CHECK(json::from_bjdata(out_str) == j_str);
// integer data declared as a double array
json const j_float = json({{"_ArrayType_", "double"}, {"_ArraySize_", {2}}, {"_ArrayData_", {1, 2}}});
const auto out_float = json::to_bjdata(j_float);
CHECK(out_float.at(0) == '{');
CHECK(json::from_bjdata(out_float) == j_float);
// a non-integer shape entry is likewise not treated as an ndarray
json const j_size = json({{"_ArrayType_", "uint8"}, {"_ArraySize_", {"x"}}, {"_ArrayData_", {1}}});
const auto out_size = json::to_bjdata(j_size);
CHECK(out_size.at(0) == '{');
CHECK(json::from_bjdata(out_size) == j_size);
// a negative shape entry is not a usable dimension either
json const j_neg = json::parse(R"({"_ArrayType_":"uint8","_ArraySize_":[-1],"_ArrayData_":[1]})");
const auto out_neg = json::to_bjdata(j_neg);
CHECK(out_neg.at(0) == '{');
CHECK(json::from_bjdata(out_neg) == j_neg);
}
SECTION("ndarray parsed from text is written as a typed array")
{
// json::parse stores a non-negative integer as number_unsigned while
// the C++ API stores an int literal as number_integer, so _ArrayType_
// names the wire type rather than the storage. Both storages have to
// produce the same typed array for every type.
for (const char* type :
{"uint8", "int8", "uint16", "int16", "uint32", "int32", "uint64", "int64", "char", "byte"
})
{
CAPTURE(type);
const std::string text = std::string(R"({"_ArrayType_":")") + type +
R"(","_ArraySize_":[2,3],"_ArrayData_":[1,2,3,4,5,6]})";
const auto from_text = json::to_bjdata(json::parse(text));
CHECK(from_text.at(0) == '[');
CHECK(from_text == json::to_bjdata(json({{"_ArrayType_", type}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2, 3, 4, 5, 6}}})));
}
// negative values under a signed type behave the same way
const auto from_neg = json::to_bjdata(json::parse(R"({"_ArrayType_":"int32","_ArraySize_":[2],"_ArrayData_":[-5,7]})"));
CHECK(from_neg.at(0) == '[');
CHECK(from_neg == json::to_bjdata(json({{"_ArrayType_", "int32"}, {"_ArraySize_", {2}}, {"_ArrayData_", {-5, 7}}})));
// and so do the floating point types
const auto from_float = json::to_bjdata(json::parse(R"({"_ArrayType_":"double","_ArraySize_":[2],"_ArrayData_":[1.5,2.5]})"));
CHECK(from_float.at(0) == '[');
CHECK(from_float == json::to_bjdata(json({{"_ArrayType_", "double"}, {"_ArraySize_", {2}}, {"_ArrayData_", {1.5, 2.5}}})));
}
SECTION("optimized ndarray (type and vector-size as 1D array)") SECTION("optimized ndarray (type and vector-size as 1D array)")
{ {
// create vector with two elements of the same type // create vector with two elements of the same type
@@ -2665,6 +2730,27 @@ TEST_CASE("BJData")
CHECK(json::from_bjdata(json::to_bjdata(j_type), true, true) == j_type); CHECK(json::from_bjdata(json::to_bjdata(j_type), true, true) == j_type);
CHECK(json::from_bjdata(json::to_bjdata(j_size), true, true) == j_size); CHECK(json::from_bjdata(json::to_bjdata(j_size), true, true) == j_size);
} }
SECTION("ndarray whose dimensions overflow stays as object")
{
// the product of the dimensions wraps around std::size_t to 0
// and so matches the size of the empty _ArrayData_; writing this
// as an ndarray would announce an element count no reader can
// honor, so it has to stay a plain object
json j_overflow = json({{"_ArrayData_", json::array()}, {"_ArraySize_", {9223372036854775808ull, 2}}, {"_ArrayType_", "uint8"}});
CHECK(json::from_bjdata(json::to_bjdata(j_overflow), true, true) == j_overflow);
// a single dimension that does not fit into std::size_t is
// rejected for the same reason (only observable where
// std::size_t is narrower than 64 bit)
json j_huge = json({{"_ArrayData_", json::array()}, {"_ArraySize_", {18446744073709551615ull}}, {"_ArrayType_", "uint8"}});
CHECK(json::from_bjdata(json::to_bjdata(j_huge), true, true) == j_huge);
// a well-formed ndarray is still encoded as one
json j_ok = json({{"_ArrayData_", {1, 2, 3, 4, 5, 6}}, {"_ArraySize_", {2, 3}}, {"_ArrayType_", "uint8"}});
CHECK(json::to_bjdata(j_ok) == std::vector<uint8_t>({'[', '$', 'U', '#', '[', 'i', 2, 'i', 3, ']', 1, 2, 3, 4, 5, 6}));
CHECK(json::from_bjdata(json::to_bjdata(j_ok), true, true) == j_ok);
}
} }
} }
+35
View File
@@ -529,6 +529,41 @@ TEST_CASE("BSON")
CHECK(json::from_bson(result, true, false) == j); CHECK(json::from_bson(result, true, false) == j);
} }
SECTION("non-empty object with binary member without subtype")
{
const size_t N = 10;
const auto s = std::vector<std::uint8_t>(N, 'x');
json const j =
{
{ "entry", json::binary(s) }
};
CHECK(!j.at("entry").get_binary().has_subtype());
std::vector<std::uint8_t> const expected =
{
0x1B, 0x00, 0x00, 0x00, // size (little endian)
0x05, // entry: binary
'e', 'n', 't', 'r', 'y', '\x00',
0x0A, 0x00, 0x00, 0x00, // size of binary (little endian)
0x00, // Generic binary subtype
0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip adds the generic binary subtype
const auto roundtrip = json::from_bson(result);
CHECK(roundtrip != j);
CHECK(roundtrip.at("entry").get_binary().has_subtype());
CHECK(roundtrip.at("entry").get_binary().subtype() == 0);
CHECK(json::from_bson(result, true, false) == roundtrip);
}
SECTION("non-empty object with binary member with subtype") SECTION("non-empty object with binary member with subtype")
{ {
// an MD5 hash // an MD5 hash
+36
View File
@@ -1999,6 +1999,42 @@ TEST_CASE("CBOR regressions")
} }
#endif #endif
TEST_CASE("CBOR definite length equal to the indefinite-length sentinel")
{
// A definite-length array or map whose declared element count equals the
// reserved unknown_size() sentinel (SIZE_MAX) must be rejected. Otherwise
// it is read as an indefinite-length container and the following bytes are
// silently accepted instead of the (impossible) count being reported.
json _;
SECTION("array")
{
// 0x9B: array with eight-byte length; length = 0xFFFFFFFFFFFFFFFF
const std::vector<uint8_t> input = {0x9B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0x01, 0x02, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), "[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive array size", json::out_of_range&);
}
SECTION("map")
{
// 0xBB: map with eight-byte length; length = 0xFFFFFFFFFFFFFFFF
const std::vector<uint8_t> input = {0xBB, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0x61, 0x61, 0x01, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), "[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size", json::out_of_range&);
}
SECTION("indefinite-length containers are unaffected")
{
CHECK(json::from_cbor(std::vector<uint8_t>({0x9F, 0x01, 0x02, 0xFF})) == json({1, 2}));
CHECK(json::from_cbor(std::vector<uint8_t>({0xBF, 0x61, 0x61, 0x01, 0xFF})) == json({{"a", 1}}));
}
SECTION("ordinary four-byte length containers are unaffected")
{
// 0x9A/0xBA carry a four-byte length; a normal count still parses
CHECK(json::from_cbor(std::vector<uint8_t>({0x9A, 0x00, 0x00, 0x00, 0x02, 0x01, 0x02})) == json({1, 2}));
CHECK(json::from_cbor(std::vector<uint8_t>({0xBA, 0x00, 0x00, 0x00, 0x01, 0x61, 0x61, 0x01})) == json({{"a", 1}}));
}
}
TEST_CASE("CBOR roundtrips" * doctest::skip()) TEST_CASE("CBOR roundtrips" * doctest::skip())
{ {
SECTION("input from flynn") SECTION("input from flynn")
+49
View File
@@ -1440,6 +1440,13 @@ TEST_CASE("parser class")
] ]
)"; )";
const auto* structured_object = R"(
{
"foo": [1, 2],
"bar": 3
}
)";
SECTION("filter nothing") SECTION("filter nothing")
{ {
const json j_object = json::parse(s_object, [](int /*unused*/, json::parse_event_t /*unused*/, const json& /*unused*/) noexcept const json j_object = json::parse(s_object, [](int /*unused*/, json::parse_event_t /*unused*/, const json& /*unused*/) noexcept
@@ -1515,6 +1522,48 @@ TEST_CASE("parser class")
CHECK (j_filtered2 == json({1})); CHECK (j_filtered2 == json({1}));
} }
SECTION("filter array in object")
{
// the array is discarded once it is already stored under its key
const json j_filtered1 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json& /*parsed*/) noexcept
{
return e != json::parse_event_t::array_end;
});
CHECK (j_filtered1 == json({{"bar", 3}}));
// the array is discarded before it is stored, leaving the
// placeholder the key event wrote
const json j_filtered2 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json& /*parsed*/) noexcept
{
return e != json::parse_event_t::array_start;
});
CHECK (j_filtered2 == json({{"bar", 3}}));
}
SECTION("filter value in object")
{
// the value is discarded after its key was kept, leaving the
// placeholder the key event wrote
const json j_filtered1 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json & parsed) noexcept
{
return !(e == json::parse_event_t::value && parsed == json(3));
});
CHECK (j_filtered1 == json({{"foo", {1, 2}}}));
// the same value is discarded together with its key, so no
// placeholder was stored for it
const json j_filtered2 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json & parsed) noexcept
{
return !((e == json::parse_event_t::key && parsed == json("bar")) ||
(e == json::parse_event_t::value && parsed == json(3)));
});
CHECK (j_filtered2 == json({{"foo", {1, 2}}}));
}
SECTION("filter specific events") SECTION("filter specific events")
{ {
SECTION("first closing event") SECTION("first closing event")
+24
View File
@@ -427,6 +427,30 @@ TEST_CASE("deserialization")
CHECK(l.events == std::vector<std::string>({"boolean(true)"})); CHECK(l.events == std::vector<std::string>({"boolean(true)"}));
} }
SECTION("from std::vector<signed char>")
{
std::vector<signed char> const v = {'t', 'r', 'u', 'e'};
CHECK(json::parse(v) == json(true));
CHECK(json::accept(v));
SaxEventLogger l;
CHECK(json::sax_parse(v, &l));
CHECK(l.events.size() == 1);
CHECK(l.events == std::vector<std::string>({"boolean(true)"}));
// bytes outside ASCII are negative here and must not be sign-extended;
// 0xC3 and 0xA9 do not fit in signed char (MSVC C4309), so spell them as negative values
std::vector<signed char> const umlaut = {'"', static_cast<signed char>(0xC3 - 0x100), static_cast<signed char>(0xA9 - 0x100), '"'};
CHECK(json::parse(umlaut) == json("\xC3\xA9"));
CHECK(json::accept(umlaut));
// 0xFF (spelled as -1 to stay in range) must not be reported as end of input
std::vector<signed char> const trailing = {'t', 'r', 'u', 'e', static_cast<signed char>(0xFF - 0x100)};
json _;
CHECK_THROWS_WITH_AS(_ = json::parse(trailing), "[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: 'true\xFF'; expected end of input", json::parse_error&);
CHECK(!json::accept(trailing));
}
SECTION("from std::array") SECTION("from std::array")
{ {
std::array<uint8_t, 5> const v { {'t', 'r', 'u', 'e'} }; std::array<uint8_t, 5> const v { {'t', 'r', 'u', 'e'} };
+30
View File
@@ -38,6 +38,36 @@ TEST_CASE("Better diagnostics with positions")
"[json.exception.type_error.302] type must be number, but is string", json::type_error); "[json.exception.type_error.302] type must be number, but is string", json::type_error);
} }
SECTION("positions of strings containing escape sequences")
{
// escape sequences make the token longer than the string it parses to,
// so the positions must not be derived from the parsed value's length
const auto check = [](const std::string & text, const std::string & token)
{
CAPTURE(text)
CAPTURE(token)
const json j = json::parse(text);
const json& v = j.at("a");
CHECK(text.substr(v.start_pos(), v.end_pos() - v.start_pos()) == token);
};
check(R"({"a":"plain"})", R"("plain")");
check(R"({"a":"tab\there"})", R"("tab\there")");
check(R"({"a":"\n\n\n\n\n\n"})", R"("\n\n\n\n\n\n")");
check(R"({"a":"\""})", R"("\"")");
check(R"({"a":"\\"})", R"("\\")");
check(R"({"a":"é"})", R"("é")");
check(R"({"a":"🌞"})", R"("🌞")");
check("{\"a\":\"\xc3\xa9\"}", "\"\xc3\xa9\""); // multi-byte UTF-8, no escapes
// a string at the root, where an escape would otherwise push the
// reported start position past the opening quote
const std::string root = R"("a\tb")";
const json j = json::parse(root);
CHECK(j.start_pos() == 0);
CHECK(j.end_pos() == root.size());
}
SECTION("JSON patch add to primitive parent (#4292)") SECTION("JSON patch add to primitive parent (#4292)")
{ {
// the JSON Patch "add" target /foo/bar/baz has a string parent // the JSON Patch "add" target /foo/bar/baz has a string parent
+25
View File
@@ -1530,4 +1530,29 @@ TEST_CASE("issue #4320 - custom base class must not leak nlohmann::detail into A
CHECK(j == json({{"x", 1.0}, {"y", 2.0}, {"z", 3.0}})); CHECK(j == json({{"x", 1.0}, {"y", 2.0}, {"z", 3.0}}));
} }
TEST_CASE("issue #5338 - truncated CBOR tagged binary subtype is rejected")
{
const std::vector<std::vector<std::uint8_t>> truncated_tags =
{
{0xD8},
{0xD9, 0x00},
{0xDA, 0x00, 0x00, 0x00},
{0xDB, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}
};
for (const auto& data : truncated_tags)
{
CAPTURE(data);
for (const auto tag_handler :
{
json::cbor_tag_handler_t::ignore, json::cbor_tag_handler_t::store
})
{
CAPTURE(tag_handler);
const auto result = json::from_cbor(data, true, false, tag_handler);
CHECK(result.is_discarded());
}
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP DOCTEST_CLANG_SUPPRESS_WARNING_POP
+30
View File
@@ -53,4 +53,34 @@ TEST_CASE("type traits")
// NOLINTEND(hicpp-avoid-c-arrays,modernize-avoid-c-arrays,cppcoreguidelines-avoid-c-arrays) // NOLINTEND(hicpp-avoid-c-arrays,modernize-avoid-c-arrays,cppcoreguidelines-avoid-c-arrays)
} }
} }
SECTION("char_traits")
{
SECTION("to_int_type does not sign-extend")
{
using unsigned_traits = nlohmann::detail::char_traits<unsigned char>;
using signed_traits = nlohmann::detail::char_traits<signed char>;
CHECK(unsigned_traits::to_int_type(static_cast<unsigned char>(0x7F)) == 0x7F);
CHECK(unsigned_traits::to_int_type(static_cast<unsigned char>(0x80)) == 0x80);
CHECK(unsigned_traits::to_int_type(static_cast<unsigned char>(0xFF)) == 0xFF);
CHECK(signed_traits::to_int_type(static_cast<signed char>(0x7F)) == 0x7F);
// 0x80 and 0xFF do not fit in signed char (MSVC C4309), so spell them as negative values
CHECK(signed_traits::to_int_type(static_cast<signed char>(0x80 - 0x100)) == 0x80);
CHECK(signed_traits::to_int_type(static_cast<signed char>(0xFF - 0x100)) == 0xFF);
}
SECTION("no byte value collides with eof")
{
using unsigned_traits = nlohmann::detail::char_traits<unsigned char>;
using signed_traits = nlohmann::detail::char_traits<signed char>;
for (int i = 0; i < 256; ++i)
{
CHECK(unsigned_traits::to_int_type(static_cast<unsigned char>(i)) != unsigned_traits::eof());
CHECK(signed_traits::to_int_type(static_cast<signed char>(i)) != signed_traits::eof());
}
}
}
} }
+40
View File
@@ -819,6 +819,8 @@ TEST_CASE("UBJSON")
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec2), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1A", json::parse_error); CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec2), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1A", json::parse_error);
std::vector<uint8_t> const vec3 = {'H', 'i', 2, '1', '.'}; std::vector<uint8_t> const vec3 = {'H', 'i', 2, '1', '.'};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec3), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1.", json::parse_error); CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec3), "[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1.", json::parse_error);
std::vector<uint8_t> const vec_overflow = {'H', 'i', 5, '1', 'e', '4', '0', '0'};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec_overflow), "[json.exception.out_of_range.406] number overflow parsing '1e400'", json::out_of_range&);
std::vector<uint8_t> const vec4 = {'H', 2, '1', '0'}; std::vector<uint8_t> const vec4 = {'H', 2, '1', '0'};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec4), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON size: expected length type specification (U, i, I, l, L) after '#'; last byte: 0x02", json::parse_error); CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vec4), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON size: expected length type specification (U, i, I, l, L) after '#'; last byte: 0x02", json::parse_error);
} }
@@ -1711,6 +1713,44 @@ TEST_CASE("UBJSON")
CHECK(json::to_ubjson(json::from_ubjson(s_L)) == s_i); CHECK(json::to_ubjson(json::from_ubjson(s_L)) == s_i);
} }
SECTION("no-op markers")
{
// A no-op ('N') is valid wherever a value may start; it is consumed
// by get_ignore_noop() before the value is read. It is not valid
// where a string length type specification is expected.
SECTION("accepted where a value may start")
{
// at top level, also repeated
CHECK(json::from_ubjson(std::vector<uint8_t>({'N', 'i', 1})) == json(1));
CHECK(json::from_ubjson(std::vector<uint8_t>({'N', 'N', 'N', 'i', 1})) == json(1));
// inside an array of unknown size, before and after an element
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', 'N', 'i', 1, ']'})) == json({1}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', 'i', 1, 'N', ']'})) == json({1}));
// inside an object of unknown size: before a key, between key
// and value, and before the closing '}'
CHECK(json::from_ubjson(std::vector<uint8_t>({'{', 'N', 'U', 1, 'a', 'i', 1, '}'})) == json({{"a", 1}}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'{', 'U', 1, 'a', 'N', 'i', 1, '}'})) == json({{"a", 1}}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'{', 'U', 1, 'a', 'i', 1, 'N', '}'})) == json({{"a", 1}}));
}
SECTION("rejected where a length type specification is expected")
{
json _;
// after the 'S' marker of a string value
std::vector<uint8_t> const v_S = {'S', 'N', 'U', 1, 'a'};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v_S), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON string: expected length type specification (U, i, I, l, L); last byte: 0x4E", json::parse_error&);
// as the key length of an object with a known size, where
// no-ops are not permitted in the first place
std::vector<uint8_t> const v_key = {'{', '#', 'i', 1, 'N', 'U', 1, 'a', 'i', 1};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v_key), "[json.exception.parse_error.113] parse error at byte 5: syntax error while parsing UBJSON string: expected length type specification (U, i, I, l, L); last byte: 0x4E", json::parse_error&);
}
}
SECTION("number") SECTION("number")
{ {
SECTION("float") SECTION("float")
+10
View File
@@ -125,6 +125,16 @@ TEST_CASE("wide strings")
std::u32string const w = U"\"\x110000"; std::u32string const w = U"\"\x110000";
json _; json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&); CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// a code unit above U+10FFFF must not be narrowed onto the EOF
// sentinel: 0xFFFFFFFF would otherwise end the document silently and
// let everything following it pass the strict end-of-input check
std::u32string const trailing{U'[', U'1', U']', static_cast<char32_t>(0xFFFFFFFF), U'x'};
CHECK_THROWS_WITH_AS(_ = json::parse(trailing), "[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: '1]\xFF'; expected end of input", json::parse_error&);
CHECK(!json::accept(trailing));
// the same unit inside a string is reported as an ill-formed byte
CHECK_THROWS_WITH_AS(_ = json::parse(std::u32string{U'"', static_cast<char32_t>(0xFFFFFFFF), U'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xFF'", json::parse_error&);
} }
} }
} }