Compare commits

..
Author SHA1 Message Date
Niels Lohmann 5e93415d91 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/lexer.hpp
#	include/nlohmann/detail/input/parser.hpp
#	single_include/nlohmann/json.hpp
#	tests/src/fuzzer-parse_json.cpp
2026-10-01 07:36:30 +02:00
Niels Lohmann 7e8e8e219b Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/binary_reader.hpp
#	include/nlohmann/json.hpp
#	single_include/nlohmann/json.hpp
2026-09-30 22:53:03 +02:00
Niels Lohmann 5f1727cef2 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	tests/src/unit-alt-string.cpp
2026-09-30 20:11:28 +02:00
Niels Lohmann c16dd7e4f5 Fix CI: do not hide base class methods in the #3989 test SAX parsers
clang-tidy's bugprone-derived-method-shadowing-base-method rejected the
test SAX parsers that derive from json_sax_dom_parser (or the test's
SaxEventLogger) and redefine their non-virtual event functions. The
recovering DOM parsers of unit-class_parser.cpp and unit-regression2.cpp
now hold a json_sax_dom_parser and forward to it, SaxEventLogger gets a
flag to recover from errors instead of a derived class, and the one
function unit-alt-string.cpp redefines is marked as intended.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 04:50:11 +02:00
Niels Lohmann 792853d725 Fix CI: clang-tidy and JSON_NOEXCEPTION in the #3989 changes
ci_clang_tidy flagged nested conditional operators in the BON8 skip
code and in remove_incomplete_utf8_sequence()
(readability-avoid-nested-conditional-operator), and two branches with
the same body in recover_string() (bugprone-branch-clone). Use if
chains instead of the nested conditionals and merge the two branches
into one condition; the short-circuit order is unchanged.

ci_test_noexceptions aborted in the #3989 regression test: it compares
the first recovered error with the message of the exception that
from_cbor() and friends throw, and under JSON_NOEXCEPTION that call
aborts instead of throwing. Guard binary_error_message() and its uses
with #if !defined(JSON_NOEXCEPTION), as other tests in the suite do.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 16:42:57 +02:00
Niels Lohmann 4bdf1b7e74 Fix CI: no global std::string in the #3989 regression test
The U+FFFD constant was a global std::string, which the clang CI build
rejects with -Wglobal-constructors. It is now a function.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 09:27:46 +02:00
Niels Lohmann b1e9d98e41 Merge remote-tracking branch 'origin/claude/fix-issue-3989-db7e45' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:31:07 +02:00
Niels Lohmann f855d257df Fix CI: keep raw strings with backslashes out of test macros for MSVC
MSVC stringizes the arguments of doctest's CHECK() so that a raw string
literal becomes an ordinary one, and then reported the "\q" in the new
alt_string recovery test as warning C4129, an error with /WX. The input
is now a variable. The same pattern with "\u0000" in the parser's
recovery test is replaced by a JSON value built from a std::string.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:30:39 +02:00
Niels Lohmann 3e683e9c04 Merge branch 'develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:22:25 +02:00
Niels Lohmann d1d84ed9af Fix Flawfinder: format the BSON element type without snprintf
Moving the report of an unsupported BSON element type into
skip_unsupported_bson_element() moved its snprintf() call, which
Flawfinder then reported as a new CWE-134 finding. The two hexadecimal
digits are now computed directly; the message is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 20:36:28 +02:00
Niels Lohmann de8529f99b Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/binary_reader.hpp
#	single_include/nlohmann/json.hpp
2026-09-28 17:53:51 +02:00
Niels Lohmann 677794f076 Fix CI: recover numbers with the string operations every string_t has
recover_number() used back(), pop_back(), front(), and
find_first_of(const char*), which the minimal alt_string of
unit-alt-string.cpp does not provide. Since the binary readers recover
UBJSON/BJData high-precision numbers with it, from_ubjson() instantiated
it, too, and the test no longer compiled. It now uses only size(),
operator[], and resize(), and unit-alt-string.cpp recovers from errors
in JSON text, UBJSON, and CBOR.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 12:17:11 +02:00
Niels Lohmann 437a95cfdb Repair complete items in binary formats when parse_error() returns true (#3989)
When the SAX parser asks to recover, the binary readers now repair an
item whose end is known and read on after it, as RFC 8949, Section 5.3
describes for CBOR:

- CBOR: tags are ignored, and simple values other than false, true, and
  null become null (RFC 8949, Section 6.1); a negative integer below the
  range of number_integer_t becomes the nearest floating-point number.
- Strings that are not valid UTF-8 get U+FFFD for each ill-formed
  sequence, as in JSON text; so does a UBJSON/BJData char above 0x7F.
- UBJSON/BJData high-precision numbers keep their longest valid
  beginning (via the lexer's recover_token()), or become infinity.
- Members whose key is not a string are skipped (CBOR, MessagePack,
  BON8), like members without a key in JSON text.
- BSON elements of types the library does not read (ObjectId, datetime,
  decimal128, ...) become null; a string without its terminator and a
  document whose size does not match are kept.

Where the end of an item is unknown, reading stops as before, except
that BSON skips to the end of the document, whose size it knows.

The value read before such an error is now completed by the reader from
its container stack, as the JSON parser does, instead of by a proxy SAX
parser, which is removed. Like the parser, binary_reader gets an
AllowRecovery template parameter, so that from_*() compile without the
new code.

Tests: a table of repairs, numbers out of range, errors that stop, and
a sweep over changed and removed bytes of eight encodings that checks
balanced events and that the first error is the one from_*() reports.
All fuzzers now run a recovering checker; the binary ones also check
that it reports an error exactly when from_*() fails.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 21:40:59 +02:00
Niels Lohmann c8735246d0 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 21:02:38 +02:00
Niels Lohmann 9adb510a0d Recover from parse errors when parse_error() returns true (#3989)
The return value of json_sax::parse_error() was documented both as "must
return false" and as "whether the parsing should continue", and the code
just passed it on. For JSON text, parsing stopped anyway, but sax_parse()
could report success for invalid input. The binary readers read on after
the error, looping forever on a CBOR indefinite-length array without its
end.

Now false stops parsing, and true recovers from the error:

- JSON text is repaired with the smallest local edit (insert a missing
  ',' or ':', remove a stray token, keep the readable part of a broken
  string or number, null for a value that cannot be read, close the
  innermost container at a wrong closing bracket and all of them at the
  end of the input), and parsing continues. The SAX events stay balanced,
  every key is followed by exactly one value, and each token is reported
  at most once.
- The binary formats cannot resynchronize, so they stop, but complete
  the value read so far.

sax_parse() returns false after any error. parse(), accept(), and the
from_*() functions never recover and compile to the same code as before.

Supersedes #4522.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 20:16:42 +02:00
96 changed files with 7288 additions and 2628 deletions
+5 -10
View File
@@ -1,15 +1,9 @@
# bugprone-use-after-move (hicpp-invalid-access-moved is its alias) still flags
# the basic_json move constructor, which forwards the whole object to its base
# class (#5724), and two forwards in the error-message construction of
# at(KeyType&&) (json.hpp, both overloads: find(std::forward<KeyType>(key))
# followed by string_t(std::forward<KeyType>(key)) in the throw), which #5689
# rewrites. Re-enable both checks once those changes have landed.
# portability-avoid-pragma-once: kept disabled on purpose. #pragma once is accepted
# by every supported compiler, and tools/amalgamate/amalgamate.py strips it from
# single_include, so there is nothing left to fix here.
# TODO: The first three checks are only removed to get the CI going. They have to be addressed at some point.
# TODO: portability-avoid-pragma-once: should be fixed eventually
Checks: '*,
-portability-template-virtual-member-function,
-bugprone-use-after-move,
-hicpp-invalid-access-moved,
@@ -43,6 +37,7 @@ Checks: '*,
-google-readability-function-size,
-google-runtime-float,
-google-runtime-int,
-google-runtime-references,
-hicpp-avoid-goto,
-hicpp-explicit-conversions,
-hicpp-function-size,
@@ -76,7 +71,6 @@ Checks: '*,
-readability-magic-numbers,
-readability-redundant-access-specifiers,
-readability-redundant-parentheses,
-readability-redundant-typename,
-readability-simplify-boolean-expr,
-readability-uppercase-literal-suffix,
-readability-use-concise-preprocessor-directives'
@@ -87,4 +81,5 @@ CheckOptions:
WarningsAsErrors: '*'
#HeaderFilterRegex: '.*nlohmann.*'
HeaderFilterRegex: '.*hpp$'
+1 -1
View File
@@ -85,7 +85,7 @@ jobs:
ci_static_analysis_clang:
runs-on: ubuntu-latest
container: silkeh/clang:22
container: silkeh/clang:dev
strategy:
matrix:
target: [ci_test_clang, ci_clang_tidy, ci_test_clang_sanitizer, ci_clang_analyze, ci_single_binaries]
+2 -2
View File
@@ -495,7 +495,7 @@ bool key(string_t& val);
bool parse_error(std::size_t position, const std::string& last_token, const detail::exception& ex);
```
The return value of each function determines whether parsing should proceed.
The return value of each function determines whether parsing should proceed. For `parse_error`, returning `true` [recovers from the error](https://json.nlohmann.me/features/parsing/error_recovery/): the parser repairs the input and continues.
To implement your own SAX handler, proceed as follows:
@@ -503,7 +503,7 @@ To implement your own SAX handler, proceed as follows:
2. Create an object of your SAX interface class, e.g. `my_sax`.
3. Call `bool json::sax_parse(input, &my_sax)`; where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface.
Note the `sax_parse` function only returns a `bool` indicating the result of the last executed SAX event. It does not return a `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error -- it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file [`json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp).
Note the `sax_parse` function only returns a `bool` indicating whether the input was parsed without errors and no SAX event returned `false`. It does not return a `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error -- it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file [`json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp).
### STL-like access
+4 -4
View File
@@ -8,24 +8,24 @@ set(N 10)
include(FindPython3)
find_package(Python3 COMPONENTS Interpreter)
find_program(CLANG_TOOL NAMES clang++-HEAD clang++ clang++-22 clang++-21 clang++-20 clang++-19 clang++-18 clang++-17 clang++-16 clang++-15 clang++-14 clang++-13 clang++-12 clang++-11 clang++)
find_program(CLANG_TOOL NAMES clang++-HEAD clang++ clang++-20 clang++-19 clang++-18 clang++-17 clang++-16 clang++-15 clang++-14 clang++-13 clang++-12 clang++-11 clang++)
execute_process(COMMAND ${CLANG_TOOL} --version OUTPUT_VARIABLE CLANG_TOOL_VERSION ERROR_VARIABLE CLANG_TOOL_VERSION)
string(REGEX MATCH "[0-9]+(\\.[0-9]+)+" CLANG_TOOL_VERSION "${CLANG_TOOL_VERSION}")
message(STATUS "🔖 Clang ${CLANG_TOOL_VERSION} (${CLANG_TOOL})")
find_program(CLANG_TIDY_TOOL NAMES clang-tidy-22 clang-tidy-21 clang-tidy-20 clang-tidy-19 clang-tidy-18 clang-tidy-17 clang-tidy-16 clang-tidy-15 clang-tidy-14 clang-tidy-13 clang-tidy-12 clang-tidy-11 clang-tidy)
find_program(CLANG_TIDY_TOOL NAMES clang-tidy-20 clang-tidy-19 clang-tidy-18 clang-tidy-17 clang-tidy-16 clang-tidy-15 clang-tidy-14 clang-tidy-13 clang-tidy-12 clang-tidy-11 clang-tidy)
execute_process(COMMAND ${CLANG_TIDY_TOOL} --version OUTPUT_VARIABLE CLANG_TIDY_TOOL_VERSION ERROR_VARIABLE CLANG_TIDY_TOOL_VERSION)
string(REGEX MATCH "[0-9]+(\\.[0-9]+)+" CLANG_TIDY_TOOL_VERSION "${CLANG_TIDY_TOOL_VERSION}")
message(STATUS "🔖 Clang-Tidy ${CLANG_TIDY_TOOL_VERSION} (${CLANG_TIDY_TOOL})")
message(STATUS "🔖 CMake ${CMAKE_VERSION} (${CMAKE_COMMAND})")
find_program(GCC_TOOL NAMES g++-latest g++-HEAD g++ g++-16 g++-15 g++-14 g++-13 g++-12 g++-11 g++-10)
find_program(GCC_TOOL NAMES g++-latest g++-HEAD g++ g++-15 g++-14 g++-13 g++-12 g++-11 g++-10)
execute_process(COMMAND ${GCC_TOOL} --version OUTPUT_VARIABLE GCC_TOOL_VERSION ERROR_VARIABLE GCC_TOOL_VERSION)
string(REGEX MATCH "[0-9]+(\\.[0-9]+)+" GCC_TOOL_VERSION "${GCC_TOOL_VERSION}")
message(STATUS "🔖 GCC ${GCC_TOOL_VERSION} (${GCC_TOOL})")
find_program(GCOV_TOOL NAMES gcov-HEAD gcov gcov-16 gcov-15 gcov-14 gcov-13 gcov-12 gcov-11 gcov-10)
find_program(GCOV_TOOL NAMES gcov-HEAD gcov gcov-15 gcov-14 gcov-13 gcov-12 gcov-11 gcov-10)
execute_process(COMMAND ${GCOV_TOOL} --version OUTPUT_VARIABLE GCOV_TOOL_VERSION ERROR_VARIABLE GCOV_TOOL_VERSION)
string(REGEX MATCH "[0-9]+(\\.[0-9]+)+" GCOV_TOOL_VERSION "${GCOV_TOOL_VERSION}")
message(STATUS "🔖 GCOV ${GCOV_TOOL_VERSION} (${GCOV_TOOL})")
+2
View File
@@ -2,6 +2,7 @@
# -Wno-c++98-compat The library targets C++11.
# -Wno-c++98-compat-pedantic The library targets C++11.
# -Wno-deprecated-declarations The library contains annotations for deprecated functions.
# -Wno-extra-semi-stmt The library uses assert which triggers this warning.
# -Wno-padded We do not care about padding warnings.
# -Wno-covered-switch-default All switches list all cases and a default case.
# -Wno-c2y-extensions Clang 22.1 diagnoses __COUNTER__ as a C2y extension, also in
@@ -19,6 +20,7 @@ set(CLANG_CXXFLAGS
-Wno-c++98-compat
-Wno-c++98-compat-pedantic
-Wno-deprecated-declarations
-Wno-extra-semi-stmt
-Wno-padded
-Wno-covered-switch-default
-Wno-c2y-extensions
+13 -24
View File
@@ -1,4 +1,4 @@
# Warning flags determined for GCC 16.2.0 with https://github.com/nlohmann/gcc_flags:
# Warning flags determined for GCC 15.1.0 with https://github.com/nlohmann/gcc_flags:
# Ignored GCC warnings:
# -Wno-abi-tag We do not care about ABI tags.
# -Wno-aggregate-return The library uses aggregate returns.
@@ -16,8 +16,6 @@ set(GCC_CXXFLAGS
--extra-warnings
-W
-WNSObject-attribute
-Wabbreviated-auto-in-template-arg
-Wabi
-Wno-abi-tag
-Waddress
-Waddress-of-packed-member
@@ -66,7 +64,6 @@ set(GCC_CXXFLAGS
-Wanalyzer-tainted-divisor
-Wanalyzer-tainted-offset
-Wanalyzer-tainted-size
-Wanalyzer-throw-of-unexpected-type
-Wanalyzer-too-complex
-Wanalyzer-undefined-behavior-ptrdiff
-Wanalyzer-undefined-behavior-strtok
@@ -83,13 +80,10 @@ set(GCC_CXXFLAGS
-Warith-conversion
-Warray-bounds=2
-Warray-compare
-Warray-parameter
-Warray-parameter=2
-Wattribute-alias=2
-Wattribute-warning
-Wattributes
-Wauto-profile
-Wbidi-chars=any
-Wbool-compare
-Wbool-operation
-Wbuiltin-declaration-mismatch
@@ -105,7 +99,6 @@ set(GCC_CXXFLAGS
-Wc++20-compat
-Wc++20-extensions
-Wc++23-extensions
-Wc++26-compat
-Wc++26-extensions
-Wc++2a-compat
-Wcalloc-transposed-args
@@ -149,7 +142,6 @@ set(GCC_CXXFLAGS
-Wdeprecated-enum-enum-conversion
-Wdeprecated-enum-float-conversion
-Wdeprecated-literal-operator
-Wdeprecated-openmp
-Wdeprecated-variadic-comma-omission
-Wdisabled-optimization
-Wdiv-by-zero
@@ -164,18 +156,21 @@ set(GCC_CXXFLAGS
-Wenum-conversion
-Wexceptions
-Wexpansion-to-defined
-Wexperimental-fmv-target
-Wexpose-global-module-tu-local
-Wexternal-tu-local
-Wextra
-Wextra-semi
-Wflex-array-member-not-at-end
-Wfloat-conversion
-Wfloat-equal
-Wformat-diag
-Wformat-overflow=2
-Wformat-signedness
-Wformat-truncation=2
-Wformat -Wformat-contains-nul
-Wformat -Wformat-diag
-Wformat -Wformat-extra-args
-Wformat -Wformat-nonliteral
-Wformat -Wformat-overflow=2
-Wformat -Wformat-security
-Wformat -Wformat-signedness
-Wformat -Wformat-truncation=2
-Wformat -Wformat-y2k
-Wformat -Wformat-zero-length
-Wformat=2
-Wframe-address
-Wfree-nonheap-object
@@ -202,8 +197,6 @@ set(GCC_CXXFLAGS
-Winvalid-offsetof
-Winvalid-pch
-Winvalid-utf8
-Wkeyword-macro
-Wleading-whitespace=spaces
-Wliteral-suffix
-Wlogical-not-parentheses
-Wlogical-op
@@ -234,7 +227,6 @@ set(GCC_CXXFLAGS
-Wnarrowing
-Wnoexcept
-Wnoexcept-type
-Wnon-c-typedef-for-linkage
-Wnon-template-friend
-Wnon-virtual-dtor
-Wnonnull
@@ -277,8 +269,6 @@ set(GCC_CXXFLAGS
-Wscalar-storage-order
-Wself-move
-Wsequence-point
-Wsfinae-incomplete
-Wsfinae-incomplete=2
-Wshadow=compatible-local
-Wshadow=global
-Wshadow=local
@@ -299,7 +289,6 @@ set(GCC_CXXFLAGS
-Wstrict-aliasing=3
-Wstrict-null-sentinel
-Wstrict-overflow
-Wstrict-overflow=5
-Wstring-compare
-Wstringop-overflow
-Wstringop-overflow=4
@@ -344,8 +333,8 @@ set(GCC_CXXFLAGS
-Wunreachable-code
-Wunsafe-loop-optimizations
-Wunused
-Wunused-but-set-parameter=3
-Wunused-but-set-variable=3
-Wunused-but-set-parameter
-Wunused-but-set-variable
-Wunused-const-variable=2
-Wunused-function
-Wunused-label
+3 -6
View File
@@ -25,12 +25,10 @@ and `ensure_ascii` parameters.
result consists of ASCII characters only.
`error_handler` (in)
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md):
: how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
with U+FFFD), `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped), and `keep` (write the ill-formed bytes to the output as is,
without escaping them, even if `ensure_ascii` is `#!cpp true`; the result is then not valid UTF-8, but equals the
input bytes exactly, and well-formed characters around the ill-formed bytes are still escaped as usual)).
with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped)).
## Return value
@@ -96,4 +94,3 @@ Binary values are serialized as an object containing two keys:
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
- Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0.
- Error handler `keep` added in version 3.13.0.
@@ -4,30 +4,15 @@
enum class error_handler_t {
strict,
replace,
ignore,
keep
ignore
};
```
This enumeration is used to choose how to treat ill-formed UTF-8 in a string value or object key:
- [`dump`](dump.md) uses it while serializing a `basic_json` value to text.
- [`to_cbor`](to_cbor.md), [`to_ubjson`](to_ubjson.md), [`to_bjdata`](to_bjdata.md), and [`to_bson`](to_bson.md) use it
while serializing a `basic_json` value to that binary format; none of CBOR, UBJSON, BJData, or BSON requires a
decoder to reject ill-formed UTF-8, so by default (`strict`) the library checks on write instead. `to_msgpack` and
`to_bon8` do not take this parameter: MessagePack's specification explicitly allows a string to contain ill-formed
UTF-8, so `to_msgpack` always passes it through, while BON8 always validates, since UTF-8 lead bytes are structural
to that format.
- [`from_cbor`](from_cbor.md), [`from_msgpack`](from_msgpack.md), [`from_ubjson`](from_ubjson.md),
[`from_bjdata`](from_bjdata.md), and [`from_bson`](from_bson.md) use it while parsing that binary format, to decide
whether to check a string value or object key for well-formed UTF-8 at all; by default (`keep`) they do not, as no
binary reader did before this parameter was added. `from_bon8` does not take this parameter, for the same reason
`to_bon8` does not.
Four values are differentiated:
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Three values are differentiated:
strict
: throw a `type_error`/`parse_error` exception in case of invalid UTF-8
: throw a `type_error` exception in case of invalid UTF-8
replace
: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER)
@@ -35,12 +20,6 @@ replace
ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences unchanged; only meaningful for the binary formats mentioned above, since [`dump`]
(dump.md) itself must produce text, and `keep` there writes the ill-formed bytes to the output as is, so the
result is then not valid UTF-8 (but still equals the input bytes exactly, including around any well-formed
characters, which are still escaped as usual)
## Examples
??? example
@@ -61,5 +40,3 @@ keep
## Version history
- Added in version 3.4.0.
- Added `keep`, and made this enumeration apply to the binary readers and writers in addition to `dump`, in version
3.13.0.
+3 -12
View File
@@ -5,14 +5,12 @@
template<typename InputType>
static basic_json from_bjdata(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_bjdata(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
```
Deserializes a given input to a JSON value using the BJData (Binary JData) serialization format.
@@ -60,12 +58,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
BJData does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -81,7 +73,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if a parse error occurs
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string could not be parsed
successfully, or if a string value or object key is not valid UTF-8 and `error_handler` is `strict`
successfully
- Throws [out_of_range.408](../../home/exceptions.md#jsonexceptionout_of_range408) if the size of an optimized container
or n-dimensional array cannot be represented by `std::size_t`
@@ -119,4 +111,3 @@ Linear in the size of the input.
- Added in version 3.11.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+2 -13
View File
@@ -5,14 +5,12 @@
template<typename InputType>
static basic_json from_bson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_bson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
```
Deserializes a given input to a JSON value using the BSON (Binary JSON) serialization format.
@@ -60,12 +58,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
BSON does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -83,8 +75,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
invalid string or byte array length)
- Throws [`parse_error.114`](../../home/exceptions.md#jsonexceptionparse_error114) if an unsupported BSON record type is
encountered
- Throws [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) if a string value or object key is
not valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -121,7 +111,6 @@ Linear in the size of the input.
- Added in version 3.4.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+4 -14
View File
@@ -6,16 +6,14 @@ template<typename InputType>
static basic_json from_cbor(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep);
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_cbor(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep);
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error);
```
Deserializes a given input to a JSON value using the CBOR (Concise Binary Object Representation) serialization format.
@@ -67,12 +65,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: how to treat CBOR tags (optional, `error` by default); see [`cbor_tag_handler_t`](cbor_tag_handler_t.md) for more
information
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
CBOR does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -88,9 +80,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of
other types are not supported, as JSON object keys are always strings), or if a string value or object key is not
valid UTF-8 and `error_handler` is `strict`
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity
@@ -130,7 +121,6 @@ Linear in the size of the input.
- Added `tag_handler` parameter in version 3.9.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
@@ -5,14 +5,12 @@
template<typename InputType>
static basic_json from_msgpack(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_msgpack(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
```
Deserializes a given input to a JSON value using the MessagePack serialization format.
@@ -60,12 +58,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
MessagePack's specification explicitly allows ill-formed UTF-8, so checking is opt-in: the default, `keep`, does
not check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -81,9 +73,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of
other types are not supported, as JSON object keys are always strings), or if a string value or object key is not
valid UTF-8 and `error_handler` is `strict`
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity
@@ -122,7 +113,6 @@ Linear in the size of the input.
- Added `allow_exceptions` parameter in version 3.2.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+3 -12
View File
@@ -5,14 +5,12 @@
template<typename InputType>
static basic_json from_ubjson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_ubjson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
const bool allow_exceptions = true);
```
Deserializes a given input to a JSON value using the UBJSON (Universal Binary JSON) serialization format.
@@ -60,12 +58,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
UBJSON does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -81,7 +73,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if a parse error occurs
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string could not be parsed
successfully, or if a string value or object key is not valid UTF-8 and `error_handler` is `strict`
successfully
- Throws [out_of_range.408](../../home/exceptions.md#jsonexceptionout_of_range408) if the size of an optimized container
or n-dimensional array cannot be represented by `std::size_t`
@@ -120,7 +112,6 @@ Linear in the size of the input.
- Added `allow_exceptions` parameter in version 3.2.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+4 -1
View File
@@ -90,7 +90,9 @@ The SAX event lister must follow the interface of [`json_sax`](../json_sax/index
## Return value
return value of the last processed SAX event
`#!cpp true` if the input was parsed without errors and no SAX event returned `#!cpp false`; `#!cpp false` otherwise.
In particular, the result is `#!cpp false` for input with errors, even if the SAX parser recovered from all of them
(see [error recovery](../../features/parsing/error_recovery.md)).
## Exception safety
@@ -138,6 +140,7 @@ A UTF-8 byte order mark is silently ignored.
- Ignoring comments via `ignore_comments` added in version 3.9.0.
- Added `ignore_trailing_commas` in version 3.13.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Recovering from parse errors (see [`parse_error`](../json_sax/parse_error.md)) added in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_PRECISE_STREAM_POSITION` added in version 3.13.0 to optionally leave a `#!cpp std::istream` positioned right
after the parsed value when `strict` is `#!cpp false`.
+4 -17
View File
@@ -5,18 +5,15 @@
static std::vector<std::uint8_t> to_bjdata(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
const bjdata_version_t version = bjdata_version_t::draft2);
// (2)
static void to_bjdata(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
const bjdata_version_t version = bjdata_version_t::draft2);
static void to_bjdata(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
const bjdata_version_t version = bjdata_version_t::draft2);
```
Serializes a given JSON value `j` to a byte vector using the BJData (Binary JData) serialization format. BJData aims to
@@ -46,12 +43,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: which version of BJData to use (see note on "Binary values" on [BJData](../../features/binary_formats/bjdata.md));
optional, `#!cpp bjdata_version_t::draft2` by default.
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_bjdata` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. BJData serialization as byte vector
@@ -65,8 +56,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -100,6 +89,4 @@ Linear in the size of the JSON value `j`.
## Version history
- Added in version 3.11.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
+3 -17
View File
@@ -2,14 +2,11 @@
```cpp
// (1)
static std::vector<std::uint8_t> to_bson(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict);
static std::vector<std::uint8_t> to_bson(const basic_json& j);
// (2)
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_bson(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o);
static void to_bson(const basic_json& j, detail::output_adapter<char> o);
```
BSON (Binary JSON) is a binary format in which zero or more ordered key/value pairs are stored as a single entity (a
@@ -28,12 +25,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`o` (in)
: output adapter to write serialization to
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_bson` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. BSON serialization as a byte vector
@@ -55,8 +46,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`out_of_range.415`](../../home/exceptions.md#jsonexceptionout_of_range415) if the subtype of a binary value
exceeds 255, the maximum of the BSON binary subtype; example:
`"subtype 70000 is too large for the BSON binary subtype (max 255)"`
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -93,6 +82,3 @@ pass before anything is written.
- Added in version 3.4.0.
- Linear in the size of `j`, and no longer limited by the call stack for deeply nested values, since version 3.13.0.
- `out_of_range.415` is now detected before anything is written, like the other exceptions above, since version 3.13.0.
- Throwing `type_error.316` for a string value or object key that is not valid UTF-8, detected before anything is
written, added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+3 -19
View File
@@ -2,14 +2,11 @@
```cpp
// (1)
static std::vector<std::uint8_t> to_cbor(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict);
static std::vector<std::uint8_t> to_cbor(const basic_json& j);
// (2)
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_cbor(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o);
static void to_cbor(const basic_json& j, detail::output_adapter<char> o);
```
Serializes a given JSON value `j` to a byte vector using the CBOR (Concise Binary Object Representation) serialization
@@ -29,12 +26,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`o` (in)
: output adapter to write serialization to
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_cbor` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. CBOR serialization as a byte vector
@@ -44,11 +35,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
Strong guarantee: if an exception is thrown, there are no changes in the JSON value.
## Exceptions
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
Linear in the size of the JSON value `j`.
@@ -82,5 +68,3 @@ Linear in the size of the JSON value `j`.
- Added in version 2.0.9.
- Compact representation of floating-point numbers added in version 3.8.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+3 -16
View File
@@ -4,16 +4,13 @@
// (1)
static std::vector<std::uint8_t> to_ubjson(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
const bool use_type = false);
// (2)
static void to_ubjson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
const bool use_size = false, const bool use_type = false);
static void to_ubjson(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
const bool use_size = false, const bool use_type = false);
```
Serializes a given JSON value `j` to a byte vector using the UBJSON (Universal Binary JSON) serialization format. UBJSON
@@ -39,12 +36,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: whether to add type annotations to container types (must be combined with `#!cpp use_size = true`); optional,
`#!cpp false` by default.
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_ubjson` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. UBJSON serialization as a byte vector
@@ -58,8 +49,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -93,5 +82,3 @@ Linear in the size of the JSON value `j`.
## Version history
- Added in version 3.1.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+2 -1
View File
@@ -7,7 +7,8 @@ struct json_sax;
This class describes the SAX interface used by [sax_parse](../basic_json/sax_parse.md). Each function is called in
different situations while the input is parsed. The boolean return value informs the parser whether to continue
processing the input.
processing the input; for [`parse_error`](parse_error.md), it decides whether to
[recover from the error](../../features/parsing/error_recovery.md).
## Template parameters
+24 -1
View File
@@ -21,7 +21,14 @@ A parse error occurred.
## Return value
Whether parsing should proceed (**must return `#!cpp false`**).
Whether to recover from the error:
- `#!cpp false` stops parsing.
- `#!cpp true` recovers from the error: the error is repaired and parsing continues. If that is not possible, which
happens in the binary formats when the end of the item with the error is unknown, the value read so far is completed
and parsing stops. See [error recovery](../../features/parsing/error_recovery.md) for how errors are repaired.
Either way, [`sax_parse`](../basic_json/sax_parse.md) returns `#!cpp false`.
## Examples
@@ -39,6 +46,22 @@ Whether parsing should proceed (**must return `#!cpp false`**).
--8<-- "examples/sax_parse.output"
```
??? example
The example below shows how a SAX parser recovers from errors.
```cpp
--8<-- "examples/sax_parse__error_recovery.cpp"
```
Output:
```
--8<-- "examples/sax_parse__error_recovery.output"
```
## Version history
- Added in version 3.2.0.
- Returning `#!cpp true` recovers from the error since version 3.13.0; before, parsing stopped, but the result of
[`sax_parse`](../basic_json/sax_parse.md) could be wrong.
@@ -20,6 +20,5 @@ int main()
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::replace)
<< "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< "\nstring with the invalid byte kept as is (" << j_invalid.dump(-1, ' ', false, json::error_handler_t::keep).size()
<< " bytes, not valid UTF-8 itself)\n";
<< '\n';
}
@@ -1,4 +1,3 @@
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü"
string with the invalid byte kept as is (7 bytes, not valid UTF-8 itself)
@@ -0,0 +1,43 @@
#include <iostream>
#include <iomanip>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
// a SAX parser that creates a JSON value like json::parse does, but that
// recovers from parse errors instead of stopping at the first one
class recovering_parser : public nlohmann::detail::json_sax_dom_parser<json>
{
public:
explicit recovering_parser(json& result)
: nlohmann::detail::json_sax_dom_parser<json>(result, false)
{}
bool parse_error(std::size_t position,
const std::string& /*last_token*/,
const json::exception& ex)
{
std::cout << "byte " << position << ": " << ex.what() << '\n';
// repair the input and continue
return true;
}
};
int main()
{
// JSON text with several mistakes that ends too early
const std::string text = R"({
"name": "Hello World",
"tags": ["a" "b",],
"valid": tru,
"size": 1.,
"nested": {"x": 1)";
json result;
recovering_parser sax(result);
const bool valid = json::sax_parse(text, &sax);
std::cout << "\nvalid JSON: " << std::boolalpha << valid << '\n'
<< std::setw(4) << result << std::endl;
}
@@ -0,0 +1,19 @@
byte 49: [json.exception.parse_error.101] parse error at line 3, column 20: syntax error while parsing array - unexpected string literal; expected ']'
byte 51: [json.exception.parse_error.101] parse error at line 3, column 22: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal
byte 70: [json.exception.parse_error.101] parse error at line 4, column 17: syntax error while parsing value - invalid literal; last read: '"valid": tru,'
byte 86: [json.exception.parse_error.101] parse error at line 5, column 15: syntax error while parsing value - invalid number; expected digit after '.'; last read: '1.,'
byte 109: [json.exception.parse_error.101] parse error at line 6, column 22: syntax error while parsing object - unexpected end of input; expected '}'
valid JSON: false
{
"name": "Hello World",
"nested": {
"x": 1
},
"size": 1,
"tags": [
"a",
"b"
],
"valid": null
}
@@ -63,12 +63,6 @@ The library uses the following mapping from JSON values types to BJData types ac
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
BJData strings must use UTF-8 encoding. `to_bjdata()` validates the bytes of every string value and object key
and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so a
value with such a string cannot be serialized in the first place.
!!! info "Unused BJData markers"
The following markers are not used in the conversion:
@@ -214,20 +208,6 @@ The library maps BJData types to JSON value types as follows:
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_bjdata.md) parameter left at `keep` (the default), `from_bjdata()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_bjdata()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()`'s
own `error_handler` parameter defaults to `strict` (see above), so a value read this way cannot be written back
to BJData unless a non-strict handler is passed there too.
!!! info "Round trips"
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
@@ -109,22 +109,14 @@ The library maps BSON record types to JSON value types as follows:
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
`from_bson()`.
!!! warning "Ill-formed UTF-8 in string values"
!!! warning "UTF-8 validation of string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
decoder, so checking is opt-in: with the [`error_handler`](../../api/basic_json/from_bson.md) parameter left at
`keep` (the default), `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_bson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores
the ill-formed bytes. `to_bson()`'s own `error_handler` parameter defaults to `strict` and throws the same
exception for a string value or element (key) name that is not valid UTF-8, so an object with such a key or
value cannot be produced in the first place unless a non-strict handler is passed there, even though
`from_bson()` would accept it from another source with the default `keep` handler. Element (key) names are
never validated on read, since they are read byte-by-byte as a C string. `binary` values (type `0x05`) are
unaffected, since they are not required to hold text.
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
byte-by-byte as a C string, or are not required to hold text, respectively.
??? example
@@ -189,21 +189,15 @@ The library maps CBOR types to JSON value types as follows:
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "Ill-formed UTF-8 in text strings"
!!! warning "UTF-8 validation of text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
(major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this, so checking is
opt-in: with the [`error_handler`](../../api/basic_json/from_cbor.md) parameter left at `keep` (the default),
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_cbor()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores
the ill-formed bytes. `to_cbor()`'s own [`error_handler`](../../api/basic_json/to_cbor.md) parameter defaults
to `strict` and throws the same exception for a string value or object key that is not valid UTF-8, so such a
value cannot be written back to CBOR unless a non-strict handler is passed there too. Byte strings (major
type 2) are unaffected, since they are not required to hold text.
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
text.
!!! warning "Tagged items"
@@ -153,21 +153,14 @@ The library maps MessagePack types to JSON value types as follows:
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "Ill-formed UTF-8 in string values"
!!! warning "UTF-8 validation of string values"
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
This library follows that by default: with its
[`error_handler`](../../api/basic_json/from_msgpack.md) parameter left at `keep` (the default),
`from_msgpack()` reads `str` bytes (object keys included) as-is, without validating them, so such a value
round-trips through `from_msgpack(to_msgpack(j))` byte for byte. Passing `error_handler_t::strict` makes
`from_msgpack()` check anyway and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. `to_msgpack()` itself has no `error_handler`
parameter and always writes `str` bytes as-is, since the specification permits it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way with the
default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes.
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
This library validates the bytes of every such string (object keys included) at decode time and rejects
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
to hold text.
??? example
@@ -47,12 +47,6 @@ The library uses the following mapping from JSON values types to UBJSON types ac
- strings with more than 9223372036854775807 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
UBJSON's required string encoding is UTF-8. `to_ubjson()` validates the bytes of every string value and object
key and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so
a value with such a string cannot be serialized in the first place.
!!! info "Unused UBJSON markers"
The following markers are not used in the conversion:
@@ -126,20 +120,6 @@ The library maps UBJSON types to JSON value types as follows:
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_ubjson.md) parameter left at `keep` (the default), `from_ubjson()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_ubjson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()`'s
own `error_handler` parameter defaults to `strict` (see above), so a value read this way cannot be written back
to UBJSON unless a non-strict handler is passed there too.
??? example
```cpp
@@ -0,0 +1,121 @@
# Error Recovery
By default, parsing stops at the first error. With the [SAX interface](sax_interface.md), you can instead ask the
parser to *recover*: to repair the error and continue, so that you get as much as possible out of malformed input, for
instance a file that was cut off, JSON edited by hand, or the output of a language model.
## Recovering from errors
The SAX parser's [`parse_error`](../../api/json_sax/parse_error.md) function is called for every error. Its return value
decides what happens next:
- `#!cpp false` stops parsing. This is what the SAX parsers of the library do, so [`parse`](../../api/basic_json/parse.md)
and [`accept`](../../api/basic_json/accept.md) never recover.
- `#!cpp true` repairs the error and continues parsing.
When recovering, the SAX parser still receives well-formed events: every `start_object` or `start_array` is followed by
the matching `end_object` or `end_array`, and every `key` is followed by exactly one value. A SAX parser that creates a
JSON value, such as the one in the example below, therefore gets a complete value. Parsing always ends, and
[`sax_parse`](../../api/basic_json/sax_parse.md) returns `#!cpp false` for input that is not valid JSON, even if every
error was repaired. Each token is reported at most once, and the SAX parser can stop at any error by returning
`#!cpp false`.
!!! example
The example below derives a SAX parser from the library's parser for `json` values (`json_sax_dom_parser`),
and recovers from all errors.
```cpp
--8<-- "examples/sax_parse__error_recovery.cpp"
```
Output:
```
--8<-- "examples/sax_parse__error_recovery.output"
```
## How errors are repaired
Each error is repaired with the smallest local edit: a missing separator is inserted, a stray token is removed, what can
be read of a broken string or number is kept, and a value that cannot be read at all becomes `#!json null`.
| Mistake | Repair | Example | Result |
|---------------------------|--------------------------------------------------------------------------------|------------------------------------------|----------------------------|
| missing `,` or `:` | inserted | `#!json [1 2]`, `#!json {"a" 1}` | `[1,2]`, `{"a":1}` |
| missing value | `#!json null` for an object key or between commas in an array | `#!json {"a":}`, `#!json [1,,2]` | `{"a":null}`, `[1,null,2]` |
| trailing comma | removed | `#!json [1,2,]` | `[1,2]` |
| broken string | invalid escapes and bytes are replaced (see below); a line break ends the string | `#!json ["a\qb"]` | `["aqb"]` |
| broken number | the longest valid beginning is kept | `#!json [1., 2e+]` | `[1,2]` |
| unreadable value | `#!json null` | `#!json [1, NaN, tru]` | `[1,null,null]` |
| number too large | passed as infinity, together with its text | `#!json [1e999]` | infinity (see below) |
| stray `:` | removed | `#!json ["a":1]` | `["a",1]` |
| member without a key | skipped up to the next `,` or `}` | `#!json {1:2, "b":3}` | `{"b":3}` |
| wrong closing bracket | closes the innermost array or object | `#!json {"a":[1,2}, "b":3}` | `{"a":[1,2],"b":3}` |
| input ends too early | all open arrays and objects are closed | `#!json {"a":[1,2` | `{"a":[1,2]}` |
| text before the value | skipped | `#!json )]}'{"a":1}` | `{"a":1}` |
In a string, an unknown escape like `\q` stands for the escaped character (`q`), as in JavaScript. An invalid `\u`
escape, a lone surrogate, and ill-formed UTF-8 are each replaced by U+FFFD (REPLACEMENT CHARACTER), and control
characters are kept. A string without its closing quote ends at the next line break or at the end of the input.
The input after the top-level value is not repaired: as without recovery, it is reported as an error, and parsing stops.
## Binary formats
The binary formats ([BJData](../binary_formats/bjdata.md), [BON8](../binary_formats/bon8.md),
[BSON](../binary_formats/bson.md), [CBOR](../binary_formats/cbor.md), [MessagePack](../binary_formats/messagepack.md),
and [UBJSON](../binary_formats/ubjson.md)) have no delimiters to find the next value by. So what can be repaired depends
on whether the end of the item with the error is known, a distinction that
[RFC 8949, Section 5.3](https://www.rfc-editor.org/rfc/rfc8949.html#section-5.3) makes for CBOR, too.
If the item is complete, but cannot be passed on as it is, it is replaced, and parsing continues after it:
| Mistake | Formats | Repair |
|---------------------------------------------------------------------|-----------------------------------------|-------------------------------------------------------------------------|
| tag | CBOR | ignored |
| simple value other than `false`, `true`, and `null`, like undefined | CBOR | `#!json null` |
| negative integer below the range of `number_integer_t` | CBOR | the nearest floating-point number |
| string that is not valid UTF-8 | BJData, BSON, CBOR, MessagePack, UBJSON | each ill-formed sequence becomes U+FFFD |
| character (`C`) that is not ASCII | BJData, UBJSON | U+FFFD |
| invalid high-precision number (`H`) | BJData, UBJSON | the longest valid beginning is kept, as for JSON text, or `#!json null` |
| high-precision number too large | BJData, UBJSON | passed as infinity, together with its text |
| object key that is not a string | BON8, CBOR, MessagePack | the member is skipped |
| element of a type the library does not read, like ObjectId or date | BSON | `#!json null` |
| string without its terminator | BSON | kept |
| document whose size does not match its content | BSON | kept |
CBOR tags and simple values are repaired as [RFC 8949, Section 6.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-6.1)
suggests for converting CBOR to JSON. Note that [`sax_parse`](../../api/basic_json/sax_parse.md) has no parameter for
CBOR tags, so every tag is an error there; when recovering, tags are ignored like with
[`cbor_tag_handler_t::ignore`](../../api/basic_json/cbor_tag_handler_t.md).
After any other error, the end of the item is unknown: the input ended, a byte is not a valid type marker, or a size
cannot be right. Parsing then stops, and the value read so far is completed: a key that waits for its value gets
`#!json null`, and all open arrays and objects are closed. This keeps everything before the error of an input that was
cut off. The exception is BSON, which stores the size of every document: an element whose end is unknown gets
`#!json null`, the rest of its document is skipped, and parsing continues after the document.
## Limitations
- A repair is a guess. For example, `#!json {"a" "b": 1}` could be meant as `#!json {"a": "b"}` or as
`#!json {"a": null, "b": 1}`; it is repaired to the former. Treat recovered values as a best effort, and check the
reported errors.
- A closing bracket always closes the innermost array or object. If a bracket is missing rather than wrong, the
repair differs from the intention: `#!json {"a": {"b": [1, 2}, "c": 3}` is repaired to
`#!json {"a": {"b": [1, 2], "c": 3}}`, although `#!json {"a": {"b": [1, 2]}, "c": 3}` may have been meant.
- Keys without quotes, and strings in single quotes, are not supported; such members are skipped.
- In the binary formats, a member that is skipped because its key is not a string is lost, and so are the elements of a
BSON document after one whose end is unknown.
- A number that is too large for `number_float_t` is passed as positive or negative infinity. The SAX parser's
`number_float` also gets the number's text, but a JSON value cannot store it, and
[`dump`](../../api/basic_json/dump.md) serializes infinity as `#!json null`.
- When parsing is not strict (see [`sax_parse`](../../api/basic_json/sax_parse.md)), a repair may read parts of the
input after the value, for instance of the next value in a stream of concatenated values.
## See also
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [`parse_error`](../../api/json_sax/parse_error.md) - the SAX event for parse errors
- [`sax_parse`](../../api/basic_json/sax_parse.md) - generate SAX events
- [parsing and exceptions](parse_exceptions.md) - control error handling
+2 -1
View File
@@ -65,7 +65,7 @@ You can influence a DOM parse without switching to the SAX interface by passing
When the input is not valid JSON, the `parse` function throws an exception by default. If exceptions are undesired or
unavailable, the parser can instead return a discarded value, or [`accept`](../../api/basic_json/accept.md) can be used
to only check whether an input is valid JSON. See [parsing and exceptions](parse_exceptions.md) for the available
options.
options. To get as much as possible out of malformed input, a SAX parser can [recover from errors](error_recovery.md).
## See also
@@ -76,3 +76,4 @@ options.
- [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [parsing and exceptions](parse_exceptions.md) - control error handling
- [error recovery](error_recovery.md) - get as much as possible out of malformed input
@@ -64,7 +64,8 @@ bool parse_error(std::size_t position,
const json::exception& ex);
```
The return value indicates whether the parsing should continue, so the function should usually return `#!cpp false`.
The return value decides whether to stop parsing (`#!cpp false`) or to repair the error and continue
(`#!cpp true`); see [error recovery](error_recovery.md) for the latter.
??? example
@@ -60,7 +60,8 @@ bool key(string_t& val);
bool parse_error(std::size_t position, const std::string& last_token, const json::exception& ex);
```
The return value of each function determines whether parsing should proceed.
The return value of each function determines whether parsing should proceed. For `parse_error`, returning
`#!cpp true` [recovers from the error](error_recovery.md).
To implement your own SAX handler, proceed as follows:
@@ -68,7 +69,7 @@ To implement your own SAX handler, proceed as follows:
2. Create an object of your SAX interface class, e.g. `my_sax`.
3. Call `#!cpp bool json::sax_parse(input, &my_sax);` where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface.
Note the `sax_parse` function only returns a `#!cpp bool` indicating the result of the last executed SAX event. It does not return `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error - it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file `json_sax.hpp`.
Note the `sax_parse` function only returns a `#!cpp bool` indicating whether the input was parsed without errors and no SAX event returned `#!cpp false`. It does not return `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error - it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file `json_sax.hpp`.
## See also
+1 -4
View File
@@ -341,10 +341,7 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8 and the `error_handler` parameter of the corresponding `from_*` function is
set to `strict`. By default (`error_handler_t::keep`), the bytes of a string are not checked for valid UTF-8 on read;
see the ill-formed UTF-8 notes on the individual [binary format](../features/binary_formats/index.md) pages for how
such a string is handled depending on `error_handler`.
the string's bytes are not valid UTF-8.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
+1
View File
@@ -87,6 +87,7 @@ nav:
- features/object_order.md
- Parsing:
- features/parsing/index.md
- features/parsing/error_recovery.md
- features/parsing/json_lines.md
- features/parsing/parse_exceptions.md
- features/parsing/parser_callbacks.md
@@ -354,22 +354,22 @@ void())
}
template < typename BasicJsonType, typename T, std::size_t... Idx >
std::array<T, sizeof...(Idx)> from_json_inplace_array_impl(const BasicJsonType& j,
std::array<T, sizeof...(Idx)> from_json_inplace_array_impl(BasicJsonType&& j,
identity_tag<std::array<T, sizeof...(Idx)>> /*unused*/, index_sequence<Idx...> /*unused*/)
{
return { { j.at(Idx).template get<T>()... } };
return { { std::forward<BasicJsonType>(j).at(Idx).template get<T>()... } };
}
template < typename BasicJsonType, typename T, std::size_t N >
auto from_json(const BasicJsonType& j, identity_tag<std::array<T, N>> tag)
-> decltype(from_json_inplace_array_impl(j, tag, make_index_sequence<N> {}))
auto from_json(BasicJsonType&& j, identity_tag<std::array<T, N>> tag)
-> decltype(from_json_inplace_array_impl(std::forward<BasicJsonType>(j), tag, make_index_sequence<N> {}))
{
if (JSON_HEDLEY_UNLIKELY(!j.is_array()))
{
JSON_THROW(type_error::create(302, concat("type must be array, but is ", j.type_name()), &j));
}
return from_json_inplace_array_impl(j, tag, make_index_sequence<N> {});
return from_json_inplace_array_impl(std::forward<BasicJsonType>(j), tag, make_index_sequence<N> {});
}
template<typename BasicJsonType>
@@ -504,54 +504,54 @@ template<std::size_t PTagValue, typename BasicJsonType, typename... Types>
using tuple_type = std::tuple < decltype(from_json_tuple_get_impl(std::declval<BasicJsonType>(), detail::identity_tag<Types> {}, detail::priority_tag<PTagValue> {}))... >;
template<std::size_t PTagValue, typename... Args, typename BasicJsonType, std::size_t... Idx>
tuple_type<PTagValue, const BasicJsonType&, Args...> from_json_tuple_impl_base(const BasicJsonType& j, index_sequence<Idx...> /*unused*/)
tuple_type<PTagValue, BasicJsonType, Args...> from_json_tuple_impl_base(BasicJsonType&& j, index_sequence<Idx...> /*unused*/)
{
return tuple_type<PTagValue, const BasicJsonType&, Args...>(from_json_tuple_get_impl(j.at(Idx), detail::identity_tag<Args> {}, detail::priority_tag<PTagValue> {})...);
return tuple_type<PTagValue, BasicJsonType, Args...>(from_json_tuple_get_impl(std::forward<BasicJsonType>(j).at(Idx), detail::identity_tag<Args> {}, detail::priority_tag<PTagValue> {})...);
}
template<std::size_t PTagValue, typename BasicJsonType>
std::tuple<> from_json_tuple_impl_base(const BasicJsonType& /*unused*/, index_sequence<> /*unused*/)
std::tuple<> from_json_tuple_impl_base(BasicJsonType& /*unused*/, index_sequence<> /*unused*/)
{
return {};
}
template < typename BasicJsonType, class A1, class A2 >
std::pair<A1, A2> from_json_tuple_impl(const BasicJsonType& j, identity_tag<std::pair<A1, A2>> /*unused*/, priority_tag<0> /*unused*/)
std::pair<A1, A2> from_json_tuple_impl(BasicJsonType&& j, identity_tag<std::pair<A1, A2>> /*unused*/, priority_tag<0> /*unused*/)
{
return {j.at(0).template get<A1>(),
j.at(1).template get<A2>()};
return {std::forward<BasicJsonType>(j).at(0).template get<A1>(),
std::forward<BasicJsonType>(j).at(1).template get<A2>()};
}
template<typename BasicJsonType, typename A1, typename A2>
inline void from_json_tuple_impl(const BasicJsonType& j, std::pair<A1, A2>& p, priority_tag<1> /*unused*/)
inline void from_json_tuple_impl(BasicJsonType&& j, std::pair<A1, A2>& p, priority_tag<1> /*unused*/)
{
p = from_json_tuple_impl(j, identity_tag<std::pair<A1, A2>> {}, priority_tag<0> {});
p = from_json_tuple_impl(std::forward<BasicJsonType>(j), identity_tag<std::pair<A1, A2>> {}, priority_tag<0> {});
}
template<typename BasicJsonType, typename... Args>
std::tuple<Args...> from_json_tuple_impl(const BasicJsonType& j, identity_tag<std::tuple<Args...>> /*unused*/, priority_tag<2> /*unused*/)
std::tuple<Args...> from_json_tuple_impl(BasicJsonType&& j, identity_tag<std::tuple<Args...>> /*unused*/, priority_tag<2> /*unused*/)
{
static_assert(cxpr_and<cxpr_or<cxpr_not<std::is_reference<Args>>, is_compatible_reference_type<const BasicJsonType&, Args>>...>::value,
static_assert(cxpr_and<cxpr_or<cxpr_not<std::is_reference<Args>>, is_compatible_reference_type<BasicJsonType, Args>>...>::value,
"Can not return a tuple containing references to types not contained in a Json, try Json::get_to()");
return from_json_tuple_impl_base<1, Args...>(j, index_sequence_for<Args...> {});
return from_json_tuple_impl_base<1, Args...>(std::forward<BasicJsonType>(j), index_sequence_for<Args...> {});
}
template<typename BasicJsonType, typename... Args>
inline void from_json_tuple_impl(const BasicJsonType& j, std::tuple<Args...>& t, priority_tag<3> /*unused*/)
inline void from_json_tuple_impl(BasicJsonType&& j, std::tuple<Args...>& t, priority_tag<3> /*unused*/)
{
t = from_json_tuple_impl_base<2, Args...>(j, index_sequence_for<Args...> {});
t = from_json_tuple_impl_base<2, Args...>(std::forward<BasicJsonType>(j), index_sequence_for<Args...> {});
}
template<typename BasicJsonType, typename TupleRelated>
auto from_json(const BasicJsonType& j, TupleRelated&& t)
-> decltype(from_json_tuple_impl(j, std::forward<TupleRelated>(t), priority_tag<3> {}))
auto from_json(BasicJsonType&& j, TupleRelated&& t)
-> decltype(from_json_tuple_impl(std::forward<BasicJsonType>(j), std::forward<TupleRelated>(t), priority_tag<3> {}))
{
if (JSON_HEDLEY_UNLIKELY(!j.is_array()))
{
JSON_THROW(type_error::create(302, concat("type must be array, but is ", j.type_name()), &j));
}
return from_json_tuple_impl(j, std::forward<TupleRelated>(t), priority_tag<3> {});
return from_json_tuple_impl(std::forward<BasicJsonType>(j), std::forward<TupleRelated>(t), priority_tag<3> {});
}
template < typename BasicJsonType, typename Key, typename Value, typename Compare, typename Allocator,
@@ -636,7 +636,7 @@ struct from_json_fn
/// namespace to hold default `from_json` function
/// to see why this is required:
/// http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n4381.html
namespace // NOLINT(cert-dcl59-cpp,fuchsia-header-anon-namespaces,google-build-namespaces,misc-anonymous-namespace-in-header)
namespace // NOLINT(cert-dcl59-cpp,fuchsia-header-anon-namespaces,google-build-namespaces)
{
#endif
JSON_INLINE_VARIABLE constexpr const auto& from_json = // NOLINT(misc-definitions-in-headers)
@@ -547,7 +547,7 @@ struct to_json_fn
/// namespace to hold default `to_json` function
/// to see why this is required:
/// http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n4381.html
namespace // NOLINT(cert-dcl59-cpp,fuchsia-header-anon-namespaces,google-build-namespaces,misc-anonymous-namespace-in-header)
namespace // NOLINT(cert-dcl59-cpp,fuchsia-header-anon-namespaces,google-build-namespaces)
{
#endif
JSON_INLINE_VARIABLE constexpr const auto& to_json = // NOLINT(misc-definitions-in-headers)
File diff suppressed because it is too large Load Diff
@@ -744,9 +744,6 @@ struct container_input_adapter_factory< ContainerType,
static adapter_type create(ContainerType&& container)
{
// container is forwarded twice on purpose: the resulting begin/end
// iterator types must match adapter_type, computed the same way
// NOLINTNEXTLINE(bugprone-use-after-move)
return input_adapter(begin(std::forward<ContainerType>(container)), end(std::forward<ContainerType>(container)));
}
};
+9 -4
View File
@@ -132,7 +132,9 @@ struct json_sax
@param[in] position the position in the input where the error occurs
@param[in] last_token the last read token
@param[in] ex an exception object describing the error
@return whether parsing should proceed (must return false)
@return whether to recover from the error: false stops parsing; true
repairs the error and continues, or, if that is not possible,
stops after completing the value read so far
*/
virtual bool parse_error(std::size_t position,
const std::string& last_token,
@@ -269,9 +271,12 @@ a pointer to the respective array or object for each recursion depth.
After successful parsing, the value that is passed by reference to the
constructor contains the parsed value.
@tparam BasicJsonType the JSON type
@tparam BasicJsonType the JSON type
@tparam InputAdapterType the input adapter of the lexer that can be passed to
the constructor to record diagnostic positions; it
does not matter if no lexer is passed
*/
template<typename BasicJsonType, typename InputAdapterType>
template<typename BasicJsonType, typename InputAdapterType = string_input_adapter_type>
class json_sax_dom_parser
{
public:
@@ -518,7 +523,7 @@ class json_sax_dom_parser
lexer_t* m_lexer_ref = nullptr;
};
template<typename BasicJsonType, typename InputAdapterType>
template<typename BasicJsonType, typename InputAdapterType = string_input_adapter_type>
class json_sax_dom_callback_parser
{
public:
+592 -4
View File
@@ -10,7 +10,7 @@
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint32_t
#include <cstdint> // uint8_t, uint32_t
#include <cstdio> // snprintf
#include <initializer_list> // initializer_list
#include <string> // char_traits, string
@@ -222,9 +222,9 @@ class lexer : public lexer_base<BasicJsonType>
/////////////////////
/*!
@brief get codepoint from 4 hex characters following `\\u`
@brief get codepoint from 4 hex characters following `\u`
For input "\\u c1 c2 c3 c4" the codepoint is:
For input "\u c1 c2 c3 c4" the codepoint is:
(c1 * 0x1000) + (c2 * 0x0100) + (c3 * 0x0010) + c4
= (c1 << 12) + (c2 << 8) + (c3 << 4) + (c4 << 0)
@@ -439,8 +439,16 @@ class lexer : public lexer_base<BasicJsonType>
if (0xD800 <= codepoint1 && codepoint1 <= 0xDBFF)
{
// expect next \uxxxx entry
if (JSON_HEDLEY_LIKELY(get() == '\\' && get() == 'u'))
if (JSON_HEDLEY_LIKELY(get() == '\\'))
{
if (JSON_HEDLEY_UNLIKELY(get() != 'u'))
{
// current is the character escaped by the backslash
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
string_error_resume = resume_kind::escaped_character;
return token_type::parse_error;
}
const int codepoint2 = get_codepoint();
if (JSON_HEDLEY_UNLIKELY(codepoint2 == -1))
@@ -465,7 +473,11 @@ class lexer : public lexer_base<BasicJsonType>
}
else
{
// the second escape was read completely and is a
// code point of its own
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
string_error_resume = resume_kind::after_escape;
string_error_codepoint = codepoint2;
return token_type::parse_error;
}
}
@@ -479,7 +491,9 @@ class lexer : public lexer_base<BasicJsonType>
{
if (JSON_HEDLEY_UNLIKELY(0xDC00 <= codepoint1 && codepoint1 <= 0xDFFF))
{
// the escape was read completely
error_message = "invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF";
string_error_resume = resume_kind::after_escape;
return token_type::parse_error;
}
}
@@ -2095,6 +2109,573 @@ scan_number_done:
}
}
/////////////////////
// error recovery
/////////////////////
/*!
@brief make the best of the token that scan() rejected
Called by the parser after scan() returned token_type::parse_error and the
SAX parser asked to recover from the error (see #3989). Keeps what can be
read of the token and skips the rest:
- A string keeps its characters. An unknown escape stands for the escaped
character itself (as in JavaScript), an invalid `\u` escape and ill-formed
UTF-8 become U+FFFD, and a control character is kept. A line break or the
end of the input ends a string that lacks its closing quote.
- A number keeps its longest valid prefix, e.g. `1` for `1.` or `1e+`.
- A block comment that is not closed runs to the end of the input.
- Anything else is skipped.
The rest of an invalid token is skipped up to the next delimiter
(whitespace, a structural character, or a quote). A delimiter that the
invalid token consumed is returned to the input, so that the next scan()
reads it.
@return token_type::value_string or a number token type if a string or a
number could be read, token_type::end_of_input for a block comment
that is not closed, token_type::uninitialized otherwise
*/
token_type recover_token()
{
const resume_kind resume = string_error_resume;
const int codepoint = string_error_codepoint;
string_error_resume = resume_kind::character;
string_error_codepoint = -1;
if (error_message_starts_with("invalid string"))
{
return recover_string(resume, codepoint);
}
if (error_message_starts_with("invalid number"))
{
return recover_number();
}
if (error_message_starts_with("invalid comment; missing"))
{
// the comment runs to the end of the input
return token_type::end_of_input;
}
skip_to_delimiter();
return token_type::uninitialized;
}
/*!
@brief return the token that scan() read last to the input, so that the
next scan() reads it again
Called by the parser when recovering from an error. The token must be a
single character (',', ':', '[', ']', '{', or '}') or the end of the
input, and scan() must have read it last.
*/
void unget_token()
{
JSON_ASSERT(!next_unget);
unget();
}
/*!
@brief let the token string for the next error begin at the current character
The token string of an error reaches back to the beginning of the last
string or number. After an error, the parser calls this function so that
the next error does not report (and, with many errors, copy) everything
read since then.
*/
void restart_token_string()
{
restart_token_string_impl(std::integral_constant<bool, lazy_token_string> {});
}
private:
/// how recover_string() continues after the error scan_string() reported
enum class resume_kind : std::uint8_t
{
/// current is the next character of the string (or the end of input)
character,
/// current is the character escaped by the preceding backslash
escaped_character,
/// current is the last character of a complete escape
after_escape
};
/// whether error_message begins with @a prefix
bool error_message_starts_with(const char* prefix) const noexcept
{
const char* message = error_message;
while (*prefix != '\0')
{
if (*message++ != *prefix++)
{
return false;
}
}
return true;
}
/// whether current ends an invalid token (see recover_token())
bool current_is_delimiter() const noexcept
{
switch (current)
{
case ' ':
case '\t':
case '\n':
case '\r':
case '[':
case ']':
case '{':
case '}':
case ',':
case ':':
case '\"':
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
case char_traits<char_type>::eof():
return true;
case '/':
return ignore_comments;
default:
return false;
}
}
/// skip the rest of an invalid token and return its delimiter to the input
void skip_to_delimiter()
{
while (!current_is_delimiter())
{
get();
}
if (current != char_traits<char_type>::eof())
{
unget();
}
}
/// append U+FFFD REPLACEMENT CHARACTER to token_buffer
void add_replacement_character()
{
add(0xEF);
add(0xBF);
add(0xBD);
}
/// append the UTF-8 encoding of @a codepoint (not a surrogate) to token_buffer
void add_codepoint(const int codepoint)
{
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
const auto cp = static_cast<unsigned int>(codepoint);
if (cp < 0x80)
{
add(static_cast<char_int_type>(cp));
}
else if (cp <= 0x7FF)
{
add(static_cast<char_int_type>(0xC0u | (cp >> 6u)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
else if (cp <= 0xFFFF)
{
add(static_cast<char_int_type>(0xE0u | (cp >> 12u)));
add(static_cast<char_int_type>(0x80u | ((cp >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
else
{
add(static_cast<char_int_type>(0xF0u | (cp >> 18u)));
add(static_cast<char_int_type>(0x80u | ((cp >> 12u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | ((cp >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
}
/// append a code point read from a `\u` escape; a surrogate becomes U+FFFD
void add_escaped_codepoint(const int codepoint)
{
if (0xD800 <= codepoint && codepoint <= 0xDFFF)
{
add_replacement_character();
}
else
{
add_codepoint(codepoint);
}
}
/*!
@brief remove an incomplete UTF-8 sequence from the end of token_buffer
next_byte_in_range() adds the bytes of a sequence as it checks them, so
when it rejects a byte, the beginning of the sequence is already in
token_buffer, which otherwise holds only complete sequences.
@return whether an incomplete sequence was removed
*/
bool remove_incomplete_utf8_sequence()
{
std::size_t lead = token_buffer.size();
std::size_t continuation_bytes = 0;
while (lead > 0 && continuation_bytes < 3
&& (static_cast<unsigned char>(token_buffer[lead - 1]) & 0xC0u) == 0x80u)
{
--lead;
++continuation_bytes;
}
if (lead == 0)
{
return false;
}
const auto lead_byte = static_cast<unsigned char>(token_buffer[lead - 1]);
std::size_t expected = 0;
if (lead_byte >= 0xF0)
{
expected = 3;
}
else if (lead_byte >= 0xE0)
{
expected = 2;
}
else if (lead_byte >= 0xC0)
{
expected = 1;
}
if (continuation_bytes >= expected)
{
return false;
}
token_buffer.resize(lead - 1);
return true;
}
/*!
@brief read the UTF-8 sequence that begins with current, which is not ASCII
@return whether the next character must be read; false if current still
needs to be handled, because it does not belong to the sequence
*/
bool recover_utf8_sequence()
{
// the number of continuation bytes and the range of the first one;
// see the ranges in scan_string()
std::size_t count = 0;
char_int_type low = 0x80;
char_int_type high = 0xBF;
if (current >= 0xC2 && current <= 0xDF)
{
count = 1;
}
else if (current >= 0xE0 && current <= 0xEF)
{
count = 2;
low = (current == 0xE0) ? 0xA0 : 0x80;
high = (current == 0xED) ? 0x9F : 0xBF;
}
else if (current >= 0xF0 && current <= 0xF4)
{
count = 3;
low = (current == 0xF0) ? 0x90 : 0x80;
high = (current == 0xF4) ? 0x8F : 0xBF;
}
else
{
// an ill-formed byte
add_replacement_character();
return true;
}
const std::size_t start = token_buffer.size();
add(current);
for (std::size_t i = 0; i < count; ++i)
{
get();
if (current < low || current > high)
{
token_buffer.resize(start);
add_replacement_character();
return false;
}
add(current);
low = 0x80;
high = 0xBF;
}
return true;
}
/*!
@brief read the low surrogate that must follow the high surrogate @a high
@return whether the next character must be read; false if current still
needs to be handled
*/
bool recover_low_surrogate(int high)
{
while (true)
{
if (get() != '\\')
{
add_replacement_character();
return false;
}
if (get() != 'u')
{
add_replacement_character();
// not 'u', so this does not come back here
return recover_escape();
}
const int low = get_codepoint();
if (low == -1)
{
add_replacement_character();
return false;
}
if (0xDC00 <= low && low <= 0xDFFF)
{
add_codepoint(static_cast<int>((static_cast<unsigned int>(high) << 10u)
+ static_cast<unsigned int>(low) - 0x35FDC00u));
return true;
}
// high has no low surrogate
add_replacement_character();
if (low < 0xD800 || low > 0xDBFF)
{
add_codepoint(low);
return true;
}
// another high surrogate
high = low;
}
}
/*!
@brief read the escape whose backslash was read; current is the escaped character
@return whether the next character must be read; false if current still
needs to be handled
*/
bool recover_escape()
{
switch (current)
{
case '\"':
add('\"');
return true;
case '\\':
add('\\');
return true;
case '/':
add('/');
return true;
case 'b':
add('\b');
return true;
case 'f':
add('\f');
return true;
case 'n':
add('\n');
return true;
case 'r':
add('\r');
return true;
case 't':
add('\t');
return true;
case 'u':
{
const int codepoint = get_codepoint();
if (codepoint == -1)
{
add_replacement_character();
return false;
}
if (0xD800 <= codepoint && codepoint <= 0xDBFF)
{
return recover_low_surrogate(codepoint);
}
add_escaped_codepoint(codepoint);
return true;
}
// an unknown escape stands for the escaped character
default:
return false;
}
}
/*!
@brief read the rest of a string after scan_string() rejected it
token_buffer holds what scan_string() read before the error. See
recover_token() for how errors are repaired.
@param[in] resume how to continue, see resume_kind
@param[in] codepoint for a high surrogate followed by an escape of another
code point: that code point; -1 otherwise
*/
token_type recover_string(const resume_kind resume, const int codepoint)
{
// whether the next character must be read before it can be handled
bool fetch = false;
if (error_message_starts_with("invalid string: surrogate")
|| error_message_starts_with("invalid string: '\\u'")
|| (error_message_starts_with("invalid string: ill-formed UTF-8")
&& remove_incomplete_utf8_sequence()))
{
add_replacement_character();
}
switch (resume)
{
case resume_kind::escaped_character:
fetch = recover_escape();
break;
case resume_kind::after_escape:
if (0xD800 <= codepoint && codepoint <= 0xDBFF)
{
fetch = recover_low_surrogate(codepoint);
}
else
{
if (codepoint != -1)
{
add_escaped_codepoint(codepoint);
}
fetch = true;
}
break;
case resume_kind::character:
default:
break;
}
while (true)
{
if (fetch)
{
get();
}
fetch = true;
switch (current)
{
case '\"':
// a line break or the end of the input ends a string that
// lacks its closing quote
case '\n':
case '\r':
case char_traits<char_type>::eof():
return token_type::value_string;
#if !JSON_STRICT_NUL_HANDLING
case '\0':
// the end of the input, see scan()
unget();
return token_type::value_string;
#endif
case '\\':
get();
fetch = recover_escape();
break;
default:
if (current < 0x80)
{
// including control characters
add(current);
}
else
{
fetch = recover_utf8_sequence();
}
break;
}
}
}
/*!
@brief keep the longest valid prefix of a number that scan_number() rejected
token_buffer holds the characters scan_number() accepted before the error,
so the prefix ends at its last digit.
*/
token_type recover_number()
{
// only size(), operator[], and resize() are used, which every string
// type the library supports provides
std::size_t length = token_buffer.size();
while (length != 0 && (token_buffer[length - 1] < '0' || token_buffer[length - 1] > '9'))
{
--length;
}
token_buffer.resize(length);
if (length == 0)
{
skip_to_delimiter();
return token_type::uninitialized;
}
if (decimal_point_position >= length)
{
decimal_point_position = std::string::npos;
}
std::size_t exponent = std::string::npos;
for (std::size_t i = 0; i < length; ++i)
{
if (token_buffer[i] == 'e' || token_buffer[i] == 'E')
{
exponent = i;
break;
}
}
const std::size_t mantissa_end = (exponent == std::string::npos) ? length : exponent;
token_type number_type = token_type::value_unsigned;
if (decimal_point_position != std::string::npos || exponent != std::string::npos)
{
number_type = token_type::value_float;
}
else if (token_buffer[0] == '-')
{
number_type = token_type::value_integer;
}
const token_type result = convert_number(number_type, mantissa_end);
skip_to_delimiter();
return result;
}
/// seekable adapter: the token string begins at current, which was consumed
void restart_token_string_impl(std::true_type /*lazy*/) noexcept
{
const std::size_t consumed = ia.get_consumed_count();
token_string_start = (consumed > 0 && current != char_traits<char_type>::eof()) ? consumed - 1 : consumed;
}
/// streaming adapter: the token string begins at current; a character
/// that was put back is copied again when it is read again
void restart_token_string_impl(std::false_type /*lazy*/)
{
token_string.clear();
if (!next_unget && current != char_traits<char_type>::eof())
{
token_string.push_back(char_traits<char_type>::to_char_type(current));
}
}
private:
/// input adapter
InputAdapterType ia;
@@ -2135,6 +2716,13 @@ scan_number_done:
/// a description of occurred lexer errors
const char* error_message = "";
/// how recover_token() continues a string that scan_string() rejected;
/// set only on the error paths that need more than error_message
resume_kind string_error_resume = resume_kind::character;
/// the code point of the second escape when a high surrogate is followed
/// by an escape that is not a low surrogate; -1 otherwise
int string_error_codepoint = -1;
// number values
number_integer_t value_integer = 0;
number_unsigned_t value_unsigned = 0;
+662 -50
View File
@@ -138,26 +138,59 @@ class parser
bool accept(const bool strict = true)
{
json_sax_acceptor<BasicJsonType> sax_acceptor;
return sax_parse(&sax_acceptor, strict);
return sax_parse_impl<false>(&sax_acceptor, strict);
}
/*!
@brief public SAX interface
If the SAX parser's parse_error() returns true, the parser recovers from
the error: it repairs the input and continues (see #3989).
@param[in] sax the SAX parser
@param[in] strict whether to expect the last token to be EOF
@return whether the input was parsed without errors and no SAX event
returned false
*/
template<typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse(SAX* sax, const bool strict = true)
{
return sax_parse_impl<true>(sax, strict);
}
private:
/// what sax_parse_internal() does after an object key was expected
enum class next_step : std::uint8_t
{
/// stop parsing
stop,
/// parse a value that begins with last_token
parse_value,
/// evaluate the state of the innermost container, which reads
/// last_token again
evaluate_state
};
template<bool AllowRecovery, typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_impl(SAX* sax, const bool strict)
{
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
const bool result = sax_parse_internal(sax);
const bool result = sax_parse_internal<AllowRecovery>(sax);
if (result)
{
if (strict)
{
// strict mode: next byte must be EOF
if (get_token() != token_type::end_of_input)
// strict mode: next byte must be EOF; after recovering from an
// error, the end of the input may already have been read
if (last_token != token_type::end_of_input && get_token() != token_type::end_of_input)
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr));
// the value is complete, so there is nothing to recover
static_cast<void>(report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr),
std::integral_constant<bool, AllowRecovery> {}));
return false;
}
}
else
@@ -168,10 +201,9 @@ class parser
}
}
return result;
return result && !error_reported;
}
private:
/*!
@brief run a DOM SAX parser to completion and position the lexer
@@ -189,7 +221,7 @@ class parser
template<typename DomSax>
bool parse_dom(DomSax& sdp, const bool strict)
{
sax_parse_internal(&sdp);
sax_parse_internal<false>(&sdp);
if (strict)
{
@@ -212,10 +244,20 @@ class parser
return !sdp.is_errored();
}
template<typename SAX>
/*!
@brief parse a JSON value and pass it to a SAX parser
@tparam AllowRecovery whether to recover from an error if the SAX parser's
parse_error() returns true; false for the SAX parsers
of parse() and accept(), which never do, so that no
code for recovering is generated for them
*/
template<bool AllowRecovery, typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_internal(SAX* sax)
{
const std::integral_constant<bool, AllowRecovery> allow_recovery{};
// stack to remember the hierarchy of structured values we are parsing
// true = array; false = object
std::vector<bool> states;
@@ -246,12 +288,18 @@ class parser
break;
}
// parse key
// remember we are now inside an object
states.push_back(false);
// parse key (the steps of parse_key(), which are
// repeated here and below for speed)
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr));
if (!continue_after(key_error(sax, allow_recovery, false), skip_to_state_evaluation))
{
return false;
}
continue;
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
@@ -261,14 +309,13 @@ class parser
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr));
if (!continue_after(key_error(sax, allow_recovery, true), skip_to_state_evaluation))
{
return false;
}
continue;
}
// remember we are now inside an object
states.push_back(false);
// parse values
get_token();
continue;
@@ -304,9 +351,11 @@ class parser
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(res)))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr));
if (!overflow_error(sax, res, allow_recovery))
{
return false;
}
break;
}
if (JSON_HEDLEY_UNLIKELY(!sax->number_float(res, m_lexer.get_string())))
@@ -374,23 +423,63 @@ class parser
case token_type::parse_error:
{
// using "uninitialized" to avoid an "expected" message
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr));
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr), allow_recovery))
{
return false;
}
// recover: keep what can be read of the token
recover_token();
if (last_token != token_type::uninitialized)
{
// a string or a number
continue;
}
if (states.empty())
{
// look for the value after the garbage
if (!skip_to_value())
{
return false;
}
continue;
}
// nothing could be read
if (JSON_HEDLEY_UNLIKELY(!sax->null()))
{
return false;
}
break;
}
case token_type::end_of_input:
{
if (JSON_HEDLEY_UNLIKELY(m_lexer.get_position().chars_read_total == 1))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(),
"attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
// there is nothing to recover
static_cast<void>(report_error(sax, parse_error::create(101, m_lexer.get_position(),
"attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr), allow_recovery));
return false;
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr));
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr), allow_recovery))
{
return false;
}
// recover: the input ends where a value is missing
if (states.empty())
{
// there is no value
return false;
}
if (!recover_missing_value(sax, states))
{
return false;
}
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
continue;
}
case token_type::uninitialized:
case token_type::end_array:
@@ -400,9 +489,35 @@ class parser
case token_type::literal_or_value:
default: // the last token was unexpected
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr));
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr), allow_recovery))
{
return false;
}
// recover
if (states.empty())
{
// look for the value after the garbage
if (!skip_to_value())
{
return false;
}
continue;
}
if (last_token == token_type::name_separator)
{
// a stray ':'; the value may follow
get_token();
continue;
}
if (!recover_missing_value(sax, states))
{
return false;
}
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
continue;
}
}
}
@@ -453,9 +568,30 @@ class parser
continue;
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr));
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr), allow_recovery))
{
return false;
}
// recover
if (last_token == token_type::end_of_input)
{
// the input ends inside the array
return close_containers(sax, states);
}
if (last_token == token_type::end_object)
{
// a wrong closing bracket closes the innermost container
if (JSON_HEDLEY_UNLIKELY(!sax->end_array()))
{
return false;
}
states.pop_back();
skip_to_state_evaluation = true;
}
// otherwise, a missing ',' (or a stray ':', which value
// parsing drops): the next value begins here
continue;
}
// states.back() is false -> object
@@ -472,11 +608,12 @@ class parser
// parse key
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr));
if (!continue_after(key_error(sax, allow_recovery, false), skip_to_state_evaluation))
{
return false;
}
continue;
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return false;
@@ -485,9 +622,11 @@ class parser
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr));
if (!continue_after(key_error(sax, allow_recovery, true), skip_to_state_evaluation))
{
return false;
}
continue;
}
// parse values
@@ -515,12 +654,479 @@ class parser
continue;
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr));
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr), allow_recovery))
{
return false;
}
// recover
if (last_token == token_type::end_of_input)
{
// the input ends inside the object
return close_containers(sax, states);
}
if (last_token == token_type::end_array)
{
// a wrong closing bracket closes the innermost container
if (JSON_HEDLEY_UNLIKELY(!sax->end_object()))
{
return false;
}
states.pop_back();
skip_to_state_evaluation = true;
continue;
}
if (!continue_after(recover_member(sax, allow_recovery), skip_to_state_evaluation))
{
return false;
}
}
}
/*!
@brief continue sax_parse_internal() after a recovery
@return whether to continue parsing
*/
bool continue_after(const next_step step, bool& skip_to_state_evaluation)
{
if (step == next_step::evaluate_state)
{
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
}
return step != next_step::stop;
}
/// the parser for parse() and accept() never recovers: stop parsing
static std::false_type continue_after(std::false_type /*step*/, bool& /*skip_to_state_evaluation*/) noexcept
{
return {};
}
/*!
@brief parse an object key and the name separator (:) after it
last_token is the token where the key is expected. sax_parse_internal()
repeats these steps rather than calling this function, which is used
when recovering from an error.
@return next_step::parse_value if the value follows, with last_token its
first token; next_step::evaluate_state if the object's state is
to be evaluated after recovering from an error; next_step::stop
to stop parsing
*/
template<typename SAX>
next_step parse_key(SAX* sax)
{
const std::true_type allow_recovery{};
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return key_error(sax, allow_recovery, false);
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return next_step::stop;
}
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return key_error(sax, allow_recovery, true);
}
// the value begins with the next token
get_token();
return next_step::parse_value;
}
/*!
@brief report a number that is too large for number_float_t, and recover
from the error by passing the value on; the SAX parser gets the
number's text as well
This is a separate function, as reading other numbers is measurably
slower if the error is handled where they are read.
@param[in] sax the SAX parser
@param[in] value the value that is not finite
@return whether to continue parsing
*/
template<typename SAX, typename AllowRecovery>
bool overflow_error(SAX* sax, const number_float_t value, AllowRecovery allow_recovery)
{
if (!report_error(sax, out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr), allow_recovery))
{
return false;
}
return sax->number_float(value, m_lexer.get_string());
}
/*!
@brief report a missing key, or a missing name separator (:) after the
key; the parser for parse() and accept() never recovers
@param[in] key_read whether the key was read, so that the name separator
is missing
@return std::false_type, see report_error()
*/
template<typename SAX>
std::false_type key_error(SAX* sax, std::false_type allow_recovery, const bool key_read)
{
return report_error(sax, parse_error::create(101, m_lexer.get_position(), key_read
? exception_message(token_type::name_separator, "object separator")
: exception_message(token_type::value_string, "object key"), nullptr), allow_recovery);
}
/*!
@brief report a missing key, or a missing name separator (:) after the
key, and recover from it
@param[in] key_read whether the key was read, so that the name separator
is missing
*/
template<typename SAX>
next_step key_error(SAX* sax, std::true_type allow_recovery, const bool key_read)
{
if (!key_read)
{
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr), allow_recovery))
{
return next_step::stop;
}
return recover_key(sax);
}
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr), allow_recovery))
{
return next_step::stop;
}
return recover_name_separator(sax);
}
/////////////////////
// error recovery
/////////////////////
/*
The functions below repair an error after the SAX parser's parse_error()
returned true (see #3989). Each mistake is repaired by the smallest local
edit: a missing ',' or ':' is inserted, a stray token is removed, what can
be read of an invalid string or number is kept (see
lexer::recover_token()), a missing value becomes null, a wrong closing
bracket closes the innermost container, and the end of the input closes
all of them. The events stay balanced, and every key() is followed by
exactly one value.
A repair hands a token to the state evaluation, by returning it to the
lexer (lexer::unget_token()) so that the state evaluation reads it again,
only if it is ',', ']', '}', or the end of the input. The state evaluation
hands a token to value or key parsing only if it is none of them, so a
token is never handed back and forth. Every other step reads a token or
closes a container, so parsing always ends.
*/
/*!
@brief report an error to the SAX parser; the parser for parse() and
accept() never recovers
@return std::false_type rather than false: its value is known where the
function is called even if the call is not inlined, so the code
for recovering is not generated
*/
template<typename SAX, typename Exception>
std::false_type report_error(SAX* sax, const Exception& ex, std::false_type /*allow_recovery*/)
{
error_reported = true;
static_cast<void>(sax->parse_error(m_lexer.get_position(), m_lexer.get_token_string(), ex));
return {};
}
/*!
@brief report an error to the SAX parser
@return whether to recover from the error
*/
template<typename SAX, typename Exception>
bool report_error(SAX* sax, const Exception& ex, std::true_type /*allow_recovery*/)
{
const std::size_t position = m_lexer.get_position().chars_read_total;
if (error_reported && position == last_error_position && last_token == last_error_token)
{
// a repair handed on the token of the error it repaired; the
// token was reported already, and the SAX parser asked to recover
return true;
}
error_reported = true;
last_error_position = position;
last_error_token = last_token;
if (!sax->parse_error(m_lexer.get_position(), m_lexer.get_token_string(), ex))
{
return false;
}
// the token string of the next error begins here
m_lexer.restart_token_string();
return true;
}
/*!
@brief keep what can be read of the token that the lexer rejected
The error was reported for the rejected token, so it is not reported again
for the token it is repaired to (see lexer::recover_token()).
*/
token_type recover_token()
{
last_token = m_lexer.recover_token();
last_error_position = m_lexer.get_position().chars_read_total;
last_error_token = last_token;
return last_token;
}
/// pass the end events of all open containers
template<typename SAX>
bool close_containers(SAX* sax, std::vector<bool>& states)
{
while (!states.empty())
{
const bool is_array = states.back();
states.pop_back();
if (JSON_HEDLEY_UNLIKELY(is_array ? !sax->end_array() : !sax->end_object()))
{
return false;
}
}
return true;
}
/*!
@brief read tokens until one begins a value, skipping everything before
the top-level value
@return whether a value begins with last_token
*/
bool skip_to_value()
{
while (true)
{
switch (get_token())
{
case token_type::begin_array:
case token_type::begin_object:
case token_type::literal_false:
case token_type::literal_null:
case token_type::literal_true:
case token_type::value_float:
case token_type::value_integer:
case token_type::value_string:
case token_type::value_unsigned:
return true;
case token_type::end_of_input:
return false;
case token_type::parse_error:
recover_token();
if (last_token != token_type::uninitialized)
{
return true;
}
break;
case token_type::uninitialized:
case token_type::end_array:
case token_type::end_object:
case token_type::name_separator:
case token_type::value_separator:
case token_type::literal_or_value:
default:
break;
}
}
}
/*!
@brief skip the rest of an object member that cannot be read
Reads tokens, beginning with last_token, until a ',', '}', or ']' that is
not inside a container that begins in the skipped tokens, or the end of
the input.
*/
void skip_member()
{
std::size_t depth = 0;
while (true)
{
switch (last_token)
{
case token_type::begin_array:
case token_type::begin_object:
++depth;
break;
case token_type::end_array:
case token_type::end_object:
if (depth == 0)
{
return;
}
--depth;
break;
case token_type::value_separator:
if (depth == 0)
{
return;
}
break;
case token_type::end_of_input:
return;
case token_type::parse_error:
recover_token();
break;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::name_separator:
case token_type::literal_or_value:
default:
break;
}
get_token();
}
}
/*!
@brief pass a value where it is missing
last_token is ',', ']', '}', or the end of the input, where a value was
expected. In an object, the key gets null; in an array, a ',' where a
value is missing stands for null (as in JavaScript), while an array that
ends there just ends.
*/
template<typename SAX>
bool recover_missing_value(SAX* sax, const std::vector<bool>& states)
{
JSON_ASSERT(!states.empty());
if (!states.back() || last_token == token_type::value_separator)
{
return sax->null();
}
return true;
}
/// recover from a missing key; last_token is where it was expected
template<typename SAX>
next_step recover_key(SAX* sax)
{
switch (last_token)
{
case token_type::value_separator:
case token_type::end_object:
case token_type::end_array:
case token_type::end_of_input:
// no member: the object's state handles the token
return next_step::evaluate_state;
case token_type::parse_error:
recover_token();
if (last_token == token_type::value_string)
{
// a key that could be repaired
return parse_key(sax);
}
skip_member();
return next_step::evaluate_state;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::begin_array:
case token_type::begin_object:
case token_type::name_separator:
case token_type::literal_or_value:
default:
// a member without a key
skip_member();
return next_step::evaluate_state;
}
}
/// recover from a missing name separator (:) after the key; last_token
/// is where it was expected
template<typename SAX>
next_step recover_name_separator(SAX* sax)
{
switch (last_token)
{
case token_type::value_separator:
case token_type::end_object:
case token_type::end_array:
case token_type::end_of_input:
// the value is missing as well
return sax->null() ? next_step::evaluate_state : next_step::stop;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::begin_array:
case token_type::begin_object:
case token_type::name_separator:
case token_type::parse_error:
case token_type::literal_or_value:
default:
// a missing ':'; the value begins here
return next_step::parse_value;
}
}
/// recover from a token after an object member that is neither ',' nor
/// '}' (nor ']' or the end of the input, which the caller handles)
template<typename SAX>
next_step recover_member(SAX* sax, std::true_type /*allow_recovery*/)
{
if (last_token == token_type::parse_error)
{
recover_token();
}
if (last_token == token_type::value_string)
{
// a missing ','; the next key begins here
return parse_key(sax);
}
skip_member();
return next_step::evaluate_state;
}
/// the parser for parse() and accept() never recovers (and does not come
/// here, as report_error() returned false)
template<typename SAX>
std::false_type recover_member(SAX* /*sax*/, std::false_type /*allow_recovery*/) const noexcept
{
return {};
}
/// get next token from lexer
token_type get_token()
{
@@ -567,6 +1173,12 @@ class parser
const bool allow_exceptions = true;
/// whether trailing commas in objects and arrays should be ignored (true) or signaled as errors (false)
const bool ignore_trailing_commas = false;
/// whether an error was reported to the SAX parser
bool error_reported = false;
/// the position of the last reported error
std::size_t last_error_position = 0;
/// the token of the last reported error
token_type last_error_token = token_type::uninitialized;
};
} // namespace detail
@@ -35,7 +35,7 @@ This class implements a both iterators (iterator and const_iterator) for the
been set (e.g., by a constructor or a copy assignment). If the iterator is
default-constructed, it is *uninitialized* and most methods are undefined.
**The library uses assertions to detect calls on uninitialized iterators.**
This class satisfies the following concept requirements (REQ-JSON-01):
@requirement REQ-JSON-01 The class satisfies the following concept requirements:
-
[BidirectionalIterator](https://en.cppreference.com/w/cpp/named_req/BidirectionalIterator):
The iterator that can be moved can be moved in both directions (i.e.
@@ -213,11 +213,11 @@ namespace std
JSON_HEDLEY_PRAGMA(clang diagnostic ignored "-Wmismatched-tags")
#endif
template<typename IteratorType>
class tuple_size<::nlohmann::detail::iteration_proxy_value<IteratorType>> // NOLINT(cert-dcl58-cpp,bugprone-std-namespace-modification)
class tuple_size<::nlohmann::detail::iteration_proxy_value<IteratorType>> // NOLINT(cert-dcl58-cpp)
: public std::integral_constant<std::size_t, 2> {};
template<std::size_t N, typename IteratorType>
class tuple_element<N, ::nlohmann::detail::iteration_proxy_value<IteratorType >> // NOLINT(cert-dcl58-cpp,bugprone-std-namespace-modification)
class tuple_element<N, ::nlohmann::detail::iteration_proxy_value<IteratorType >> // NOLINT(cert-dcl58-cpp)
{
public:
using type = decltype(
@@ -29,7 +29,7 @@ namespace detail
iterator (to create @ref reverse_iterator) and @ref const_iterator (to
create @ref const_reverse_iterator).
This class satisfies the following concept requirements (REQ-JSON-02):
@requirement REQ-JSON-02 The class satisfies the following concept requirements:
-
[BidirectionalIterator](https://en.cppreference.com/w/cpp/named_req/BidirectionalIterator):
The iterator that can be moved can be moved in both directions (i.e.
+5 -5
View File
@@ -278,11 +278,11 @@ class json_pointer
JSON_THROW(detail::out_of_range::create(404, detail::concat("unresolved reference token '", s, "'"), nullptr));
}
// the index does not fit into size_type; on 64-bit platforms this is
// only SIZE_MAX itself (see #2203 and #5395)
// only triggered on special platforms (like 32bit), see also
// https://github.com/nlohmann/json/pull/2203
if (res >= static_cast<unsigned long long>((std::numeric_limits<size_type>::max)())) // NOLINT(runtime/int)
{
JSON_THROW(detail::out_of_range::create(410, detail::concat("array index ", s, " exceeds size_type"), nullptr));
JSON_THROW(detail::out_of_range::create(410, detail::concat("array index ", s, " exceeds size_type"), nullptr)); // LCOV_EXCL_LINE
}
return static_cast<size_type>(res);
@@ -316,7 +316,7 @@ class json_pointer
/*!
@brief create and return a reference to the pointed to value
Complexity: Linear in the number of reference tokens.
@complexity Linear in the number of reference tokens.
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if array index is not a number
@@ -403,7 +403,7 @@ class json_pointer
@return reference to the JSON value pointed to by the JSON pointer
Complexity: Linear in the length of the JSON pointer.
@complexity Linear in the length of the JSON pointer.
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if an array index was not a number
+11 -4
View File
@@ -195,6 +195,13 @@
#define JSON_NO_THREAD_LOCAL 1
#endif
// disable documentation warnings on clang
#if defined(__clang__)
#pragma clang diagnostic push
#pragma clang diagnostic ignored "-Wdocumentation"
#pragma clang diagnostic ignored "-Wdocumentation-unknown-command"
#endif
// allow disabling exceptions
#if (defined(__cpp_exceptions) || defined(__EXCEPTIONS) || defined(_CPPUNWIND)) && !defined(JSON_NOEXCEPTION)
#define JSON_THROW(exception) throw exception
@@ -253,7 +260,7 @@
{ \
/* NOLINTNEXTLINE(modernize-type-traits) we use C++11 */ \
static_assert(std::is_enum<ENUM_TYPE>::value, #ENUM_TYPE " must be an enum!"); \
/* NOLINTNEXTLINE(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) we don't want to depend on <array> */ \
/* NOLINTNEXTLINE(modernize-avoid-c-arrays) we don't want to depend on <array> */ \
static const std::pair<ENUM_TYPE, BasicJsonType> m[] = __VA_ARGS__; \
auto it = std::find_if(std::begin(m), std::end(m), \
[e](const std::pair<ENUM_TYPE, BasicJsonType>& ej_pair) -> bool \
@@ -267,7 +274,7 @@
{ \
/* NOLINTNEXTLINE(modernize-type-traits) we use C++11 */ \
static_assert(std::is_enum<ENUM_TYPE>::value, #ENUM_TYPE " must be an enum!"); \
/* NOLINTNEXTLINE(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) we don't want to depend on <array> */ \
/* NOLINTNEXTLINE(modernize-avoid-c-arrays) we don't want to depend on <array> */ \
static const std::pair<ENUM_TYPE, BasicJsonType> m[] = __VA_ARGS__; \
auto it = std::find_if(std::begin(m), std::end(m), \
[&j](const std::pair<ENUM_TYPE, BasicJsonType>& ej_pair) -> bool \
@@ -306,7 +313,7 @@ void templated_json_throw(ExceptionType exception)
{ \
/* NOLINTNEXTLINE(modernize-type-traits) we use C++11 */ \
static_assert(std::is_enum<ENUM_TYPE>::value, #ENUM_TYPE " must be an enum!"); \
/* NOLINTNEXTLINE(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) we don't want to depend on <array> */ \
/* NOLINTNEXTLINE(modernize-avoid-c-arrays) we don't want to depend on <array> */ \
static const std::pair<ENUM_TYPE, BasicJsonType> m[] = __VA_ARGS__; \
auto it = std::find_if(std::begin(m), std::end(m), \
[e](const std::pair<ENUM_TYPE, BasicJsonType>& ej_pair) -> bool \
@@ -321,7 +328,7 @@ void templated_json_throw(ExceptionType exception)
{ \
/* NOLINTNEXTLINE(modernize-type-traits) we use C++11 */ \
static_assert(std::is_enum<ENUM_TYPE>::value, #ENUM_TYPE " must be an enum!"); \
/* NOLINTNEXTLINE(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) we don't want to depend on <array> */ \
/* NOLINTNEXTLINE(modernize-avoid-c-arrays) we don't want to depend on <array> */ \
static const std::pair<ENUM_TYPE, BasicJsonType> m[] = __VA_ARGS__; \
auto it = std::find_if(std::begin(m), std::end(m), \
[&j](const std::pair<ENUM_TYPE, BasicJsonType>& ej_pair) -> bool \
@@ -8,6 +8,11 @@
#pragma once
// restore clang diagnostic settings
#if defined(__clang__)
#pragma clang diagnostic pop
#endif
// clean up
#undef JSON_ASSERT
#undef JSON_INTERNAL_CATCH
+26 -167
View File
@@ -26,7 +26,6 @@
#include <nlohmann/detail/input/binary_reader.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_utils.hpp>
@@ -94,12 +93,8 @@ class binary_writer
@param[in] sink output sink to write to (a value-type sink such as
output_vector_sink, or output_adapter_sink wrapping a
type-erased output adapter)
@param[in] error_handler_ how to treat a string value or object key that
is not valid UTF-8 (CBOR, UBJSON, BJData, and BSON only; never
consulted by @ref write_msgpack or @ref write_bon8)
*/
explicit binary_writer(OutputSinkType sink, const error_handler_t error_handler_ = error_handler_t::strict)
: oa(std::move(sink)), error_handler(error_handler_)
explicit binary_writer(OutputSinkType sink) : oa(std::move(sink))
{}
/*!
@@ -112,20 +107,14 @@ class binary_writer
from one.
@param[in] adapter output adapter to write to
@param[in] error_handler_ how to treat a string value or object key that
is not valid UTF-8 (CBOR, UBJSON, BJData, and BSON only; never
consulted by @ref write_msgpack or @ref write_bon8)
*/
template < typename SinkType = OutputSinkType,
typename std::enable_if < std::is_constructible<SinkType, output_adapter_t<CharType>>::value, int >::type = 0 >
explicit binary_writer(output_adapter_t<CharType> adapter, const error_handler_t error_handler_ = error_handler_t::strict)
: oa(SinkType(std::move(adapter))), error_handler(error_handler_)
explicit binary_writer(output_adapter_t<CharType> adapter) : oa(SinkType(std::move(adapter)))
{}
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
@throw type_error.317 if @a j is not an object
*/
void write_bson(const BasicJsonType& j)
@@ -156,8 +145,6 @@ class binary_writer
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_cbor(const BasicJsonType& j)
{
@@ -224,16 +211,13 @@ class binary_writer
case value_t::string:
{
string_t storage;
const string_t& value = sanitize_utf8_for_write(*j.m_data.m_value.string, j, storage);
// step 1: write control byte and the string length
write_cbor_head(0x60, value.size());
write_cbor_head(0x60, j.m_data.m_value.string->size());
// step 2: write the string
oa.write_characters(
reinterpret_cast<const CharType*>(value.data()),
value.size());
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
break;
}
@@ -303,17 +287,6 @@ class binary_writer
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object)
{
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead; for error_handler_t::keep
// and ::replace/::ignore the recursive write_cbor(el.first)
// call below handles the key like any other string, so no
// separate check is needed here for those
if (error_handler == error_handler_t::strict)
{
check_utf8(el.first, j);
}
write_cbor(el.first);
write_cbor(el.second);
}
@@ -656,8 +629,6 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true,
@@ -707,17 +678,14 @@ class binary_writer
case value_t::string:
{
string_t storage;
const string_t& value = sanitize_utf8_for_write(*j.m_data.m_value.string, j, storage);
if (add_prefix)
{
oa.write_character(to_char_type('S'));
}
write_number_with_ubjson_prefix(value.size(), true, use_bjdata);
write_number_with_ubjson_prefix(j.m_data.m_value.string->size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(value.data()),
value.size());
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
break;
}
@@ -872,12 +840,10 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
string_t storage;
const string_t& key = sanitize_utf8_for_write(el.first, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
key.size());
reinterpret_cast<const CharType*>(el.first.data()),
el.first.size());
write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
}
@@ -918,12 +884,8 @@ class binary_writer
/*!
@return The size of a BSON document entry header, including the id marker
and the entry name size (and its null-terminator).
@throw out_of_range.409 if @a name contains U+0000, before anything is
written
@throw type_error.316 if @a name is not valid UTF-8, before anything is
written
*/
std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{
const auto it = name.find(static_cast<typename string_t::value_type>(0));
if (JSON_HEDLEY_UNLIKELY(it != BasicJsonType::string_t::npos))
@@ -931,10 +893,8 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
}
string_t storage;
const string_t& sanitized = sanitize_utf8_for_write(name, j, storage);
return /*id*/ 1ul + sanitized.size() + /*zero-terminator*/1u;
static_cast<void>(j);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
}
/*!
@@ -954,28 +914,14 @@ class binary_writer
/*!
@brief Writes the given @a element_type and @a name to the output adapter
@a name has already been validated (and, for @ref error_handler_t::strict,
found well-formed) by @ref calc_bson_entry_header_size during the earlier
size pass, so only @ref error_handler_t::replace / @ref
error_handler_t::ignore need to sanitize it again here, to actually write
the bytes that size was computed from.
*/
void write_bson_entry_header(const string_t& name,
const std::uint8_t element_type)
{
oa.write_character(to_char_type(element_type));
if (error_handler == error_handler_t::keep || error_handler == error_handler_t::strict || is_valid_utf8(name))
{
oa.write_characters(reinterpret_cast<const CharType*>(name.data()), name.size());
}
else
{
const string_t sanitized = sanitize_utf8(name, error_handler);
oa.write_characters(reinterpret_cast<const CharType*>(sanitized.data()), sanitized.size());
}
oa.write_characters(
reinterpret_cast<const CharType*>(name.data()),
name.size());
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa.write_character(to_char_type(0x00));
@@ -1003,50 +949,24 @@ class binary_writer
/*!
@return The size of the BSON-encoded string in @a value
@throw type_error.316 if @a value is not valid UTF-8, before anything is
written
@note The UTF-8 check is skipped if @a value is already too long for the
32-bit BSON length field (@ref to_bson_length rejects it later, once
the size of the whole document is known); this also keeps the check
from reading past a StringType that reports a size larger than what
it actually holds.
*/
std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j)
static std::size_t calc_bson_string_size(const string_t& value)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
string_t storage;
const string_t& sanitized = sanitize_utf8_for_write(value, j, storage);
return sizeof(std::int32_t) + sanitized.size() + 1ul;
}
return sizeof(std::int32_t) + value.size() + 1ul;
}
/*!
@brief Writes a BSON element with key @a name and string value @a value
@a value has already been validated (and, for @ref error_handler_t::strict,
found well-formed) by @ref calc_bson_string_size during the earlier size
pass, so only @ref error_handler_t::replace / @ref error_handler_t::ignore
need to sanitize it again here, to actually write the bytes that size was
computed from.
*/
void write_bson_string(const string_t& name,
const string_t& value)
{
write_bson_entry_header(name, 0x02);
const bool sanitize = error_handler != error_handler_t::keep
&& error_handler != error_handler_t::strict
&& !is_valid_utf8(value);
const string_t sanitized = sanitize ? sanitize_utf8(value, error_handler) : string_t{};
const string_t& written = sanitize ? sanitized : value;
write_number<std::int32_t>(to_bson_length(written.size() + 1ul), true);
write_number<std::int32_t>(to_bson_length(value.size() + 1ul), true);
oa.write_characters(
reinterpret_cast<const CharType*>(written.data()),
written.size());
reinterpret_cast<const CharType*>(value.data()),
value.size());
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa.write_character(to_char_type(0x00));
@@ -1160,10 +1080,8 @@ class binary_writer
is neither an object nor an array
@throw out_of_range.415 if @a j is binary with a subtype that does not fit
into a byte, before anything is written
@throw type_error.316 if @a j is a string that is not valid UTF-8, before
anything is written
*/
std::size_t calc_bson_value_size(const BasicJsonType& j)
static std::size_t calc_bson_value_size(const BasicJsonType& j)
{
switch (j.type())
{
@@ -1183,7 +1101,7 @@ class binary_writer
return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned);
case value_t::string:
return calc_bson_string_size(*j.m_data.m_value.string, j);
return calc_bson_string_size(*j.m_data.m_value.string);
case value_t::null:
return 0ul;
@@ -1296,10 +1214,8 @@ class binary_writer
written
@throw out_of_range.415 if a binary value's subtype does not fit into a
byte, before anything is written
@throw type_error.316 if a string value or a key is not valid UTF-8,
before anything is written
*/
std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{
// the object or array whose entries are being sized, and the ones it
// is in; nothing is allocated unless the document nests
@@ -2176,7 +2092,7 @@ class binary_writer
*/
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{
check_utf8(s, context);
check_bon8_utf8(s, context);
// a string that follows another string terminates it
if (string_open)
@@ -2206,7 +2122,7 @@ class binary_writer
@throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence
*/
static void check_utf8(const string_t& s, const BasicJsonType& context)
static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
{
static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data());
@@ -2217,59 +2133,6 @@ class binary_writer
}
}
/*!
@brief return @a s as it should be written, honoring @ref error_handler
Used by @ref write_cbor, @ref write_ubjson (and so @ref write_bjdata), and
the BSON writing functions for string values and object keys; never by
@ref write_msgpack or @ref write_bon8, which do not take an @ref
error_handler (MessagePack's spec allows a str object to contain
ill-formed UTF-8, and BON8 always validates, since UTF-8 lead bytes are
structural there).
- @ref error_handler_t::keep: @a s is returned unchanged, without even
checking it (the behavior of release 3.12.0 and earlier).
- @ref error_handler_t::strict: @ref check_utf8 is called, which throws
type_error.316 if @a s is not valid UTF-8.
- @ref error_handler_t::replace / @ref error_handler_t::ignore: @a s is
sanitized into @a storage with exactly the rules @ref
serializer::dump_escaped_impl uses, so that parsing what @ref
basic_json::dump produces for the same string and the same handler
yields the same result.
Well-formed input is never copied: this returns a reference to @a s
itself in every case but a sanitized `replace`/`ignore` one, so @a
storage must outlive the returned reference only then.
@param[in] s the string (value or object key) to write
@param[in] context the value @a s belongs to (for diagnostics)
@param[out] storage backing storage for a sanitized copy
@return a reference to @a s, or to @a storage once it holds a sanitized copy
*/
const string_t& sanitize_utf8_for_write(const string_t& s, const BasicJsonType& context, string_t& storage) const
{
switch (error_handler)
{
case error_handler_t::keep:
return s;
case error_handler_t::strict:
check_utf8(s, context);
return s;
case error_handler_t::replace:
case error_handler_t::ignore:
default:
if (is_valid_utf8(s))
{
return s;
}
storage = sanitize_utf8(s, error_handler);
return storage;
}
}
/*!
@brief write an integer in the shortest encoding
@@ -2598,10 +2461,6 @@ class binary_writer
/// the output
OutputSinkType oa;
/// how to treat a string value or object key that is not valid UTF-8
/// (CBOR, UBJSON, BJData, and BSON only)
const error_handler_t error_handler = error_handler_t::strict;
};
} // namespace detail
@@ -1,38 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/// how to treat decoding errors
///
/// @ref basic_json::dump uses this to decide what to do with ill-formed
/// UTF-8 while escaping a string, and the binary writers (@ref
/// basic_json::to_cbor, @ref basic_json::to_ubjson, @ref
/// basic_json::to_bjdata, @ref basic_json::to_bson) use it the same way for
/// string values and object keys. The binary readers (@ref
/// basic_json::from_cbor, @ref basic_json::from_msgpack, @ref
/// basic_json::from_ubjson, @ref basic_json::from_bjdata, @ref
/// basic_json::from_bson) use it to decide whether to check text strings
/// and object keys for well-formed UTF-8 at all, since none of those
/// formats requires a decoder to do so.
enum class error_handler_t
{
strict, ///< throw a type_error/parse_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences unchanged
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -11,7 +11,6 @@
#include <cstddef> // size_t
#include <memory> // shared_ptr, make_shared
#include <string> // basic_string
#include <type_traits> // conditional, integral_constant, is_same
#include <utility> // move
#include <vector> // vector
@@ -119,13 +118,11 @@ class output_stream_adapter : public output_adapter_protocol<CharType>
: stream(s)
{}
// NOLINTNEXTLINE(portability-template-virtual-member-function)
void write_character(CharType c) override
{
stream.put(c);
}
// NOLINTNEXTLINE(portability-template-virtual-member-function)
void write_characters(const CharType* s, std::size_t length) override
{
stream.write(s, static_cast<std::streamsize>(length));
@@ -192,82 +189,7 @@ class output_adapter_sink
output_adapter_t<CharType> oa;
};
/// @brief whether std::basic_string<CharType> has a non-deprecated std::char_traits
/// specialization, and is therefore usable as output_adapter's default StringType
///
/// std::char_traits is only guaranteed (and, on some standard libraries, only
/// implemented without a deprecation warning) for the character types listed
/// below; std::char_traits<T> for any other T (e.g. std::uint8_t, as used by the
/// binary writers) is a non-standard extension some standard libraries deprecate.
/// See https://github.com/nlohmann/json/issues/5725 item 2.
template<typename CharType>
struct is_output_adapter_string_char_type : std::integral_constant < bool,
std::is_same<CharType, char>::value ||
std::is_same<CharType, wchar_t>::value ||
std::is_same<CharType, char16_t>::value ||
std::is_same<CharType, char32_t>::value
#if defined(__cpp_lib_char8_t) && (__cpp_lib_char8_t >= 201907L)
|| std::is_same<CharType, char8_t>::value
#endif
> {};
/// @brief placeholder type for output_adapter's StringType and (with JSON_NO_IO
/// undefined) its std::basic_ostream constructor parameter, for CharType
/// with no non-deprecated std::char_traits specialization
///
/// Never actually used: the StringType- and std::basic_ostream-based
/// output_adapter constructors are neither documented nor tested for such
/// CharType (only the std::vector-based constructor is used for them, by the
/// binary writers). Naming std::basic_string<CharType> or
/// std::basic_ostream<CharType> anywhere such a constructor would otherwise be
/// declared - even as an unused default template argument or an unused,
/// never-called overload - instantiates std::char_traits<CharType> merely to
/// name the type, which is exactly what triggers the deprecation warning this
/// placeholder avoids.
template<typename CharType>
struct output_adapter_no_string_type {};
// Select output_adapter's default StringType (and, below, its ostream
// constructor's parameter type) via partial specialization, not
// std::conditional: std::conditional<B, T, F> requires both T and F to be named
// as template arguments up front, which would still instantiate (and thus name)
// std::basic_string<CharType> / std::basic_ostream<CharType> for every CharType,
// defeating the point. A bool non-type parameter with two specializations only
// ever names the type that is actually selected.
template<typename CharType, bool = is_output_adapter_string_char_type<CharType>::value>
struct output_adapter_default_string_type
{
using type = output_adapter_no_string_type<CharType>;
};
template<typename CharType>
struct output_adapter_default_string_type<CharType, true>
{
using type = std::basic_string<CharType>;
};
#ifndef JSON_NO_IO
/// distinct from output_adapter_no_string_type, so the placeholder overloads of
/// output_adapter's constructor (used when CharType is not a character type)
/// stay distinct overloads instead of colliding into a single redeclaration
template<typename CharType>
struct output_adapter_no_ostream_type {};
template<typename CharType, bool = is_output_adapter_string_char_type<CharType>::value>
struct output_adapter_ostream_type
{
using type = output_adapter_no_ostream_type<CharType>;
};
template<typename CharType>
struct output_adapter_ostream_type<CharType, true>
{
using type = std::basic_ostream<CharType>;
};
#endif // JSON_NO_IO
template < typename CharType, typename StringType =
typename output_adapter_default_string_type<CharType>::type >
template<typename CharType, typename StringType = std::basic_string<CharType>>
class output_adapter
{
public:
@@ -276,7 +198,7 @@ class output_adapter
: oa(std::make_shared<output_vector_adapter<CharType, AllocatorType>>(vec)) {}
#ifndef JSON_NO_IO
output_adapter(typename output_adapter_ostream_type<CharType>::type& s)
output_adapter(std::basic_ostream<CharType>& s)
: oa(std::make_shared<output_stream_adapter<CharType>>(s)) {}
#endif // JSON_NO_IO
+14 -66
View File
@@ -27,7 +27,6 @@
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/cpp_future.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/recursion_depth_limit.hpp>
#include <nlohmann/detail/string_concat.hpp>
@@ -42,6 +41,14 @@ namespace detail
// serialization //
///////////////////
/// how to treat decoding errors
enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore ///< ignore invalid UTF-8 sequences
};
template<typename BasicJsonType>
class serializer
{
@@ -58,7 +65,7 @@ class serializer
@param[in] ichar indentation character to use
@param[in] pretty_print_ whether the output shall be pretty-printed
@param[in] ensure_ascii_ If @a ensure_ascii_ is true, all non-ASCII
characters in the output are escaped with `\\uXXXX` sequences, and the
characters in the output are escaped with `\uXXXX` sequences, and the
result consists of ASCII characters only.
@param[in] indent_step_ the indent level
@param[in] error_handler_ how to react on decoding errors
@@ -683,7 +690,7 @@ class serializer
@param[in] s the string to escape
Complexity: Linear in the length of string @a s.
@complexity Linear in the length of string @a s.
*/
void dump_escaped(const string_t& s)
{
@@ -832,16 +839,6 @@ class serializer
// EnsureAscii parameter is used, non-ASCII characters
if ((codepoint <= 0x1F) || (EnsureAscii && (codepoint >= 0x7F)))
{
if (EnsureAscii && error_handler == error_handler_t::keep)
{
// this character was buffered as raw bytes
// below in case it turned out to be part of
// an ill-formed sequence (which is kept as
// is); now that it decoded to a well-formed
// code point, undo that and \u-escape it
// like any other character instead
bytes = bytes_after_last_accept;
}
if (codepoint <= 0xFFFF)
{
write_u_escape(bytes, static_cast<std::uint16_t>(codepoint));
@@ -940,44 +937,6 @@ class serializer
break;
}
case error_handler_t::keep:
{
// the bytes of this (now abandoned) ill-formed
// sequence seen so far are already buffered below
// and are kept unchanged in the output
if (undumped_chars > 0)
{
// the byte that ended the sequence may be OK
// for itself (e.g., a quote that must still be
// escaped, or the lead byte of a well-formed
// code point), so read it again
--i;
}
else
{
// a byte that cannot start a sequence (e.g.,
// 0xFF or a stray continuation byte) is kept
// as well
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -986,12 +945,9 @@ class serializer
default: // decode found yet incomplete multibyte code point
{
if (!EnsureAscii || error_handler == error_handler_t::keep)
if (!EnsureAscii)
{
// code point will not be escaped (or will be kept as
// is if it turns out to be ill-formed) - copy byte to
// buffer; dropped again above if it decodes to a
// well-formed code point that needs \u-escaping
// code point will not be escaped - copy byte to buffer
string_buffer[bytes++] = s[i];
}
++undumped_chars;
@@ -1042,14 +998,6 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write the ill-formed trailing bytes as is; they were
// buffered above regardless of EnsureAscii
put_buffer(string_buffer, bytes);
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -1246,7 +1194,7 @@ class serializer
}
/*!
* @brief write a lowercase "\\uXXXX" escape sequence into @a string_buffer
* @brief write a lowercase "\uXXXX" escape sequence into @a string_buffer
*
* Branch-free replacement for `snprintf(buf, 7, "\\u%04x", codeunit)` in the
* string escaping hot path. It writes exactly six characters ('\\', 'u' and
@@ -1596,7 +1544,7 @@ class serializer
/// whether to pretty-print the output
const bool pretty_print;
/// whether to escape non-ASCII characters with \\uXXXX sequences
/// whether to escape non-ASCII characters with \uXXXX sequences
const bool ensure_ascii;
/// the indent level
+2 -1
View File
@@ -62,7 +62,8 @@ inline StringType escape(const StringType& s)
/*!
* @brief string unescaping as described in RFC 6901 (Sect. 4)
* @param[in,out] s string to unescape in place
* @param[in] s string to unescape
* @return unescaped string
*
* Note the order of escaping "~1" to "/" and "~0" to "~" is important.
*
+61 -85
View File
@@ -13,10 +13,10 @@
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <string> // string, to_string
#include <utility> // move
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -118,14 +118,13 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in three
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
places, which differ in speed, diagnostics, and how they read the input:
- decode() below: the serializer, to escape and, in strict mode, reject
ill-formed UTF-8 when dumping a string. The CBOR, MessagePack, BSON,
UBJSON and BJData readers do not use it: none of those specs requires a
decoder to reject ill-formed UTF-8 in text strings, so the readers keep
the bytes as is and leave the check to dump() and the binary writers.
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
@@ -181,19 +180,19 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
}
/*!
@brief check a string for well-formed UTF-8 (RFC 3629, section 4)
@brief check whether a string consists solely of valid UTF-8
Used by the binary readers (CBOR, MessagePack, UBJSON, BJData, BSON) when an
@ref error_handler_t other than `keep` is requested for a text string value
or object key: none of those formats requires a decoder to reject ill-formed
UTF-8 on its own, so the check is opt-in there, unlike the JSON lexer and the
serializer's @ref decode -based escaping, which always run it.
Used by the CBOR/MessagePack/BSON/UBJSON binary readers to reject text
strings that are not valid UTF-8 at decode time (RFC 8949 §3.1 and the
MessagePack/BSON specifications all require text strings to be UTF-8), so
that malformed input is caught immediately instead of only surfacing later
as a type_error.316 when the resulting value is dumped.
@param[in] s the string to check
@param[in] first the index to start checking at
@return whether `s.substr(first)` is well-formed UTF-8
@sa @ref decode
@param[in] first index of the first byte to check; the bytes before it are
assumed to have been validated already and to end on a
code point boundary
@return whether @a s (from index @a first on) is valid UTF-8
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
@@ -214,99 +213,76 @@ inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noex
}
/*!
@brief sanitize a string with ill-formed UTF-8 for @ref error_handler_t::replace or @ref error_handler_t::ignore
Replaces every maximal ill-formed subsequence with U+FFFD (`replace`) or
drops it (`ignore`), using exactly the same boundaries @ref
serializer::dump_escaped_impl uses while escaping a string: a byte that does
not extend the sequence started by the previous byte(s) is reread as the
start of a new one, instead of being swallowed along with them.
@pre @a error_handler is @ref error_handler_t::replace or @ref error_handler_t::ignore
@note Well-formed input is copied through unchanged, including bytes (e.g.
control characters or quotes) that @ref serializer::dump_escaped_impl
would itself escape; this function only concerns itself with
well-formedness, not with producing valid JSON text.
@param[in] s the string to sanitize
@param[in] error_handler @ref error_handler_t::replace or @ref error_handler_t::ignore
@return @a s with every ill-formed subsequence replaced or removed
@sa @ref decode
@brief append U+FFFD REPLACEMENT CHARACTER, encoded in UTF-8
@param[in,out] s the string to append to
*/
template<typename StringType>
inline StringType sanitize_utf8(const StringType& s, const error_handler_t error_handler)
inline void append_replacement_character(StringType& s)
{
JSON_ASSERT(error_handler == error_handler_t::replace || error_handler == error_handler_t::ignore);
s.push_back(static_cast<typename StringType::value_type>(0xEFu));
s.push_back(static_cast<typename StringType::value_type>(0xBFu));
s.push_back(static_cast<typename StringType::value_type>(0xBDu));
}
StringType result;
result.reserve(s.size());
/*!
@brief replace ill-formed UTF-8 with U+FFFD REPLACEMENT CHARACTER
Each maximal subpart of an ill-formed sequence becomes one U+FFFD, as the
Unicode Standard recommends (Section 3.9, "U+FFFD Substitution of Maximal
Subparts"), and as the parser for JSON text does when it recovers from errors.
@param[in,out] s the string to repair
@param[in] first index of the first byte to repair; the bytes before it are
assumed to be valid UTF-8 that ends on a code point boundary
*/
template<typename StringType>
inline void replace_invalid_utf8(StringType& s, const std::size_t first = 0)
{
StringType result = s;
result.resize(first);
std::uint32_t codepoint = 0;
std::uint8_t state = UTF8_ACCEPT;
// length of result after the last accepted code point
std::size_t result_len_after_last_accept = 0;
// whether bytes of an as yet unresolved sequence were already appended
bool pending = false;
std::uint32_t codepoint = 0;
// the first byte of the sequence being decoded
std::size_t sequence_start = first;
for (std::size_t i = 0; i < s.size(); ++i)
std::size_t i = first;
while (i < s.size())
{
switch (decode(state, codepoint, static_cast<std::uint8_t>(s[i])))
{
case UTF8_ACCEPT: // decode found a well-formed code point
{
result.push_back(s[i]);
result_len_after_last_accept = result.size();
pending = false;
case UTF8_ACCEPT:
for (++i; sequence_start < i; ++sequence_start)
{
result.push_back(s[sequence_start]);
}
break;
}
case UTF8_REJECT: // decode found an ill-formed byte
{
// in case we saw this byte for the first time, read it again,
// because it may be fine for itself, just not for the
// sequence that came before it
if (pending)
case UTF8_REJECT:
append_replacement_character(result);
// the byte that made the sequence ill-formed begins the next
// one, unless it began this one
if (i == sequence_start)
{
--i;
++i;
}
// drop the bytes of the ill-formed sequence buffered below
result.resize(result_len_after_last_accept);
if (error_handler == error_handler_t::replace)
{
result.append("\xEF\xBF\xBD");
result_len_after_last_accept = result.size();
}
pending = false;
state = UTF8_ACCEPT;
sequence_start = i;
break;
}
default: // decode found yet incomplete multibyte code point
{
result.push_back(s[i]);
pending = true;
default: // in the middle of a sequence
++i;
break;
}
}
}
// the string ended with an incomplete sequence
// a sequence that the string ends in the middle of
if (state != UTF8_ACCEPT)
{
result.resize(result_len_after_last_accept);
if (error_handler == error_handler_t::replace)
{
result.append("\xEF\xBF\xBD");
}
append_replacement_character(result);
}
return result;
s = std::move(result);
}
} // namespace detail
+85 -109
View File
@@ -18,6 +18,16 @@
#ifndef INCLUDE_NLOHMANN_JSON_HPP_
#define INCLUDE_NLOHMANN_JSON_HPP_
// Workaround for GCC template redefinition errors in C++ modules
// When nlohmann/json.hpp is included in a C++20 module preamble after
// other module imports, GCC may report spurious redefinition errors for
// STL templates. These pragmas suppress those false positives.
// See: https://github.com/nlohmann/json/issues/5103
#if defined(__GNUC__) && !defined(__clang__) && __cplusplus >= 202002L
#pragma GCC diagnostic push
#pragma GCC diagnostic ignored "-Wignored-attributes"
#endif
#include <algorithm> // all_of, find, for_each, none_of
#include <cmath> // isnan
#include <cstddef> // nullptr_t, ptrdiff_t, size_t
@@ -148,7 +158,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
friend class ::nlohmann::detail::iter_impl;
template<typename BasicJsonType, typename CharType, typename OutputSinkType>
friend class ::nlohmann::detail::binary_writer;
template<typename BasicJsonType, typename InputType, typename SAX>
template<typename BasicJsonType, typename InputType, typename SAX, bool AllowRecovery>
friend class ::nlohmann::detail::binary_reader;
template<typename BasicJsonType, typename InputAdapterType>
friend class ::nlohmann::detail::json_sax_dom_parser;
@@ -201,10 +211,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// used by the vector-returning to_* overloads
template<typename CharType> using vector_binary_writer =
::nlohmann::detail::binary_writer<basic_json, CharType, ::nlohmann::detail::output_vector_sink<CharType>>;
template<typename CharType> static vector_binary_writer<CharType> vector_writer(
std::vector<CharType>& v, const ::nlohmann::detail::error_handler_t error_handler = ::nlohmann::detail::error_handler_t::strict)
template<typename CharType> static vector_binary_writer<CharType> vector_writer(std::vector<CharType>& v)
{
return vector_binary_writer<CharType>(::nlohmann::detail::output_vector_sink<CharType>(v), error_handler);
return vector_binary_writer<CharType>(::nlohmann::detail::output_vector_sink<CharType>(v));
}
JSON_PRIVATE_UNLESS_TESTED:
@@ -2542,12 +2551,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
@throw what @ref json_serializer<ValueType> `from_json()` method throws
The example below shows several conversions from JSON values
@liveexample{The example below shows several conversions from JSON values
to other types. There a few things to note: (1) Floating-point numbers can
be converted to integers, (2) A JSON array can be converted to a standard
`std::vector<short>`, (3) A JSON object can be converted to C++
associative containers such as `std::unordered_map<std::string,
json>`.
be converted to integers\, (2) A JSON array can be converted to a standard
`std::vector<short>`\, (3) A JSON object can be converted to C++
associative containers such as `std::unordered_map<std::string\,
json>`.,get__ValueType_const}
@since version 2.1.0
*/
@@ -2614,7 +2623,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
@return a copy of *this, converted into @a BasicJsonType
Complexity: Depending on the implementation of the called `from_json()`
@complexity Depending on the implementation of the called `from_json()`
method.
@since version 3.2.0
@@ -2638,7 +2647,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
@return a copy of *this
Complexity: Constant.
@complexity Constant.
@since version 2.1.0
*/
@@ -2684,7 +2693,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
@tparam ValueTypeCV the provided value type
@tparam ValueType the returned value type
@return copy of the JSON value, converted to @a ValueType if necessary
@return copy of the JSON value, converted to @tparam ValueType if necessary
@throw what @ref json_serializer<ValueType> `from_json()` method throws if conversion is required
@@ -2722,12 +2731,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
@return pointer to the internally stored JSON value if the requested
pointer type @a PointerType fits to the JSON value; `nullptr` otherwise
Complexity: Constant.
@complexity Constant.
The example below shows how pointers to internal values of a
@liveexample{The example below shows how pointers to internal values of a
JSON value can be requested. Note that no type conversions are made and a
`nullptr` is returned if the value and the requested pointer type does not
match.
match.,get__PointerType}
@sa see @ref get_ptr() for explicit pointer-member access
@@ -2821,14 +2830,14 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
to the JSON value type (e.g., the JSON value is of type boolean, but a
string is requested); see example below
Complexity: Linear in the size of the JSON value.
@complexity Linear in the size of the JSON value.
The example below shows several conversions from JSON values
@liveexample{The example below shows several conversions from JSON values
to other types. There a few things to note: (1) Floating-point numbers can
be converted to integers, (2) A JSON array can be converted to a standard
`std::vector<short>`, (3) A JSON object can be converted to C++
associative containers such as `std::unordered_map<std::string,
json>`.
be converted to integers\, (2) A JSON array can be converted to a standard
`std::vector<short>`\, (3) A JSON object can be converted to C++
associative containers such as `std::unordered_map<std::string\,
json>`.,operator__ValueType}
@since version 1.0.0
*/
@@ -5258,7 +5267,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(sax, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(sax, strict);
}
/// @brief generate SAX events (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -5275,7 +5284,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(sax, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(sax, strict);
}
/// @brief generate SAX events
@@ -5283,19 +5292,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @deprecated This function is deprecated since 3.8.0 and will be removed in
/// version 4.0.0 of the library. Please use
/// sax_parse(ptr, ptr + len) instead.
//
// Clang reports "declaration is marked with '@deprecated' command but does
// not have a deprecation attribute" for this overload even though
// JSON_HEDLEY_DEPRECATED_FOR below does expand to __attribute__((deprecated));
// isolated reproductions of this exact declaration shape (doc comment,
// template<>, two stacked __attribute__ lines, an overload set of the same
// name) do not reproduce it, so this looks like a Clang comment/declaration
// association quirk specific to this overload within basic_json, not a
// genuine documentation bug. See #5725 item 2.
#if defined(__clang__)
#pragma clang diagnostic push
#pragma clang diagnostic ignored "-Wdocumentation-deprecated-sync"
#endif
template <typename SAX>
JSON_HEDLEY_DEPRECATED_FOR(3.8.0, sax_parse(ptr, ptr + len, ...))
JSON_HEDLEY_NON_NULL(2)
@@ -5310,11 +5306,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(sax, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(sax, strict);
}
#if defined(__clang__)
#pragma clang diagnostic pop
#endif
#ifndef JSON_NO_IO
/// @brief deserialize from stream
/// @sa https://json.nlohmann.me/api/basic_json/operator_gtgt/
@@ -5446,29 +5439,26 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
public:
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static std::vector<std::uint8_t> to_cbor(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict)
static std::vector<std::uint8_t> to_cbor(const basic_json& j)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result, error_handler).write_cbor(j);
vector_writer(result).write_cbor(j);
return result;
}
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict)
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o)
{
binary_writer<std::uint8_t>(o, error_handler).write_cbor(j);
binary_writer<std::uint8_t>(o).write_cbor(j);
}
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static void to_cbor(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict)
static void to_cbor(const basic_json& j, detail::output_adapter<char> o)
{
binary_writer<char>(o, error_handler).write_cbor(j);
binary_writer<char>(o).write_cbor(j);
}
/// @brief create a MessagePack serialization of a given JSON value
@@ -5499,31 +5489,28 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static std::vector<std::uint8_t> to_ubjson(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
const bool use_type = false)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result, error_handler).write_ubjson(j, use_size, use_type);
vector_writer(result).write_ubjson(j, use_size, use_type);
return result;
}
/// @brief create a UBJSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static void to_ubjson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
const bool use_size = false, const bool use_type = false)
{
binary_writer<std::uint8_t>(o, error_handler).write_ubjson(j, use_size, use_type);
binary_writer<std::uint8_t>(o).write_ubjson(j, use_size, use_type);
}
/// @brief create a UBJSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static void to_ubjson(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
const bool use_size = false, const bool use_type = false)
{
binary_writer<char>(o, error_handler).write_ubjson(j, use_size, use_type);
binary_writer<char>(o).write_ubjson(j, use_size, use_type);
}
/// @brief create a BJData serialization of a given JSON value
@@ -5531,12 +5518,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static std::vector<std::uint8_t> to_bjdata(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
const bjdata_version_t version = bjdata_version_t::draft2)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
vector_writer(result).write_ubjson(j, use_size, use_type, true, true, version);
return result;
}
@@ -5544,47 +5530,42 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/to_bjdata/
static void to_bjdata(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
const bjdata_version_t version = bjdata_version_t::draft2)
{
binary_writer<std::uint8_t>(o, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
binary_writer<std::uint8_t>(o).write_ubjson(j, use_size, use_type, true, true, version);
}
/// @brief create a BJData serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bjdata/
static void to_bjdata(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
const bjdata_version_t version = bjdata_version_t::draft2)
{
binary_writer<char>(o, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
binary_writer<char>(o).write_ubjson(j, use_size, use_type, true, true, version);
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static std::vector<std::uint8_t> to_bson(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict)
static std::vector<std::uint8_t> to_bson(const basic_json& j)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result, error_handler).write_bson(j);
vector_writer(result).write_bson(j);
return result;
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict)
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o)
{
binary_writer<std::uint8_t>(o, error_handler).write_bson(j);
binary_writer<std::uint8_t>(o).write_bson(j);
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static void to_bson(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict)
static void to_bson(const basic_json& j, detail::output_adapter<char> o)
{
binary_writer<char>(o, error_handler).write_bson(j);
binary_writer<char>(o).write_bson(j);
}
/// @brief create a BON8 serialization of a given JSON value
@@ -5618,13 +5599,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json from_cbor(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep)
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor, error_handler).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5639,13 +5619,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json from_cbor(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep)
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor, error_handler).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5687,13 +5666,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_msgpack(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5707,13 +5685,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_msgpack(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5753,13 +5730,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_ubjson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5773,13 +5749,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_ubjson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5819,13 +5794,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bjdata(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5839,13 +5813,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bjdata(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5895,13 +5868,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5915,13 +5887,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
const bool allow_exceptions = true)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -6999,7 +6970,7 @@ namespace std // NOLINT(cert-dcl58-cpp)
/// @brief hash value for JSON objects
/// @sa https://json.nlohmann.me/api/basic_json/std_hash/
NLOHMANN_BASIC_JSON_TPL_DECLARATION
struct hash<nlohmann::NLOHMANN_BASIC_JSON_TPL> // NOLINT(cert-dcl58-cpp,bugprone-std-namespace-modification)
struct hash<nlohmann::NLOHMANN_BASIC_JSON_TPL> // NOLINT(cert-dcl58-cpp)
{
std::size_t operator()(const nlohmann::NLOHMANN_BASIC_JSON_TPL& j) const
{
@@ -7032,7 +7003,7 @@ struct less< ::nlohmann::detail::value_t> // do not remove the space after '<',
/// @brief exchanges the values of two JSON objects
/// @sa https://json.nlohmann.me/api/basic_json/std_swap/
NLOHMANN_BASIC_JSON_TPL_DECLARATION
inline void swap(nlohmann::NLOHMANN_BASIC_JSON_TPL& j1, nlohmann::NLOHMANN_BASIC_JSON_TPL& j2) noexcept( // NOLINT(readability-inconsistent-declaration-parameter-name, cert-dcl58-cpp,bugprone-std-namespace-modification)
inline void swap(nlohmann::NLOHMANN_BASIC_JSON_TPL& j1, nlohmann::NLOHMANN_BASIC_JSON_TPL& j2) noexcept( // NOLINT(readability-inconsistent-declaration-parameter-name, cert-dcl58-cpp)
is_nothrow_move_constructible<nlohmann::NLOHMANN_BASIC_JSON_TPL>::value&& // NOLINT(misc-redundant-expression,cppcoreguidelines-noexcept-swap,performance-noexcept-swap)
is_nothrow_move_assignable<nlohmann::NLOHMANN_BASIC_JSON_TPL>::value)
{
@@ -7046,7 +7017,7 @@ inline void swap(nlohmann::NLOHMANN_BASIC_JSON_TPL& j1, nlohmann::NLOHMANN_BASIC
/// @brief std::formatter specialization for JSON values
/// @sa https://json.nlohmann.me/api/basic_json/std_formatter/
NLOHMANN_BASIC_JSON_TPL_DECLARATION
struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-cpp,bugprone-std-namespace-modification)
struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-cpp)
{
// -1 means compact output (dump()); any value >= 0 means pretty-printed
// output with that many spaces (or indent_char) per level (dump(indent, indent_char)).
@@ -7126,6 +7097,11 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
// unit that includes this header.
#include <nlohmann/detail/macro_unscope.hpp> // IWYU pragma: keep
// End of GCC diagnostic pragmas for C++ modules support
#if defined(__GNUC__) && !defined(__clang__) && __cplusplus >= 202002L
#pragma GCC diagnostic pop
#endif
// The user-defined string literals are in a separate header, because their
// bodies instantiate the parser in every translation unit that includes them.
// Define JSON_NO_AUTOMATIC_UDLS to include <nlohmann/json_literals.hpp> only
File diff suppressed because it is too large Load Diff
+1
View File
@@ -47,6 +47,7 @@ inline namespace json_literals
namespace detail
{
using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_callback_parser;
using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_parser;
using NLOHMANN_JSON_NAMESPACE::detail::unknown_size;
} // namespace detail
+2 -2
View File
@@ -89,11 +89,11 @@ target_compile_options(test_main PUBLIC
# https://github.com/nlohmann/json/pull/3229
$<$<CXX_COMPILER_ID:Intel>:-diag-disable=2196>
$<$<NOT:$<CXX_COMPILER_ID:MSVC>>:-Wno-deprecated;-Wno-float-equal>
$<$<CXX_COMPILER_ID:GNU>:-Wno-deprecated-declarations>
$<$<CXX_COMPILER_ID:Intel>:-diag-disable=1786>)
target_include_directories(test_main SYSTEM PUBLIC
thirdparty/doctest)
target_include_directories(test_main PUBLIC
thirdparty/doctest
${PROJECT_BINARY_DIR}/include)
target_link_libraries(test_main PUBLIC ${NLOHMANN_JSON_TARGET_NAME})
+1
View File
@@ -12,6 +12,7 @@ target_compile_options(abi_compat_common INTERFACE
# https://github.com/nlohmann/json/pull/3229
$<$<CXX_COMPILER_ID:Intel>:-diag-disable=2196>
$<$<NOT:$<CXX_COMPILER_ID:MSVC>>:-Wno-deprecated;-Wno-float-equal>
$<$<CXX_COMPILER_ID:GNU>:-Wno-deprecated-declarations>
$<$<CXX_COMPILER_ID:Intel>:-diag-disable=1786>)
target_include_directories(abi_compat_common SYSTEM INTERFACE
+1 -1
View File
@@ -28,7 +28,7 @@ std::string namespace_name(std::string ns, T* /*unused*/ = nullptr) // NOLINT(pe
std::smatch m;
// extract the true namespace name from the function signature
CAPTURE(ns)
CAPTURE(ns);
CHECK(std::regex_search(ns, m, std::regex("nlohmann(::[a-zA-Z0-9_]+)*::basic_json")));
return m.str();
+12
View File
@@ -45,6 +45,10 @@ dumps is stable under exactly the same values that break operator==.
The unit tests run the same checks on a fixed corpus (see the "BJData round-trip
invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bjdata() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -57,6 +61,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// value-stable comparison for the round-trip checks below; see the note
@@ -69,11 +75,15 @@ static bool is_value_stable(const json& lhs, const json& rhs)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bjdata).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bjdata(vec1);
assert(recovered_without_errors);
try
{
@@ -107,6 +117,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -115,6 +126,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -19,6 +19,10 @@ It also checks that reading the data from a stream, which reads strings byte by
byte, gives the same value or error as reading it from contiguous memory, which
copies strings in bulk.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bon8() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -32,6 +36,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
namespace
@@ -55,6 +61,9 @@ std::string read_bon8(InputType&& input)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bon8).errors == 0;
// contiguous and stream input must be read alike
{
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
@@ -66,6 +75,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bon8(vec1);
assert(recovered_without_errors);
try
{
@@ -87,6 +97,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -95,6 +106,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_bson(vec)
- assert(to_bson(j2) == vec)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -27,16 +31,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bson).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bson(vec1);
assert(recovered_without_errors);
try
{
@@ -58,6 +68,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -66,6 +77,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_cbor(vec)
- assert(to_cbor(j2) == vec)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_cbor() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -27,16 +31,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::cbor).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_cbor(vec1);
assert(recovered_without_errors);
try
{
@@ -58,6 +68,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -66,6 +77,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+13
View File
@@ -16,6 +16,10 @@ array data, it performs the following steps:
- s2 = serialize(j2)
- assert(s1 == s2)
Furthermore, it parses data with a SAX parser that recovers from every error
and checks that the events are balanced, that parsing ends, and that valid
input is parsed without errors (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -28,11 +32,20 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
{
const auto checker = check_recovering_parse(data, size, json::input_format_t::json);
assert(checker.events <= (4 * size) + 4);
assert((checker.errors == 0) == json::accept(data, data + size));
}
try
{
// step 1: parse input
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_msgpack(vec)
- assert(to_msgpack(j2) == vec)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_msgpack() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -27,16 +31,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::msgpack).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_msgpack(vec1);
assert(recovered_without_errors);
try
{
@@ -58,6 +68,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -66,6 +77,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -24,6 +24,10 @@ array data, it performs the following steps:
The unit tests run the same checks on a fixed corpus (see the "UBJSON round-trip
invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_ubjson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -36,16 +40,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::ubjson).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_ubjson(vec1);
assert(recovered_without_errors);
try
{
@@ -77,6 +87,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -85,6 +96,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+154
View File
@@ -0,0 +1,154 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cassert>
#include <cstddef>
#include <cstdint>
#include <sstream>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
namespace
{
// a SAX parser that recovers from every error and checks that the events are
// balanced and that every key is followed by exactly one value
class recovering_checker : public nlohmann::json_sax<nlohmann::json>
{
public:
bool null() override
{
return value();
}
bool boolean(bool /*val*/) override
{
return value();
}
bool number_integer(number_integer_t /*val*/) override
{
return value();
}
bool number_unsigned(number_unsigned_t /*val*/) override
{
return value();
}
bool number_float(number_float_t /*val*/, const string_t& /*s*/) override
{
return value();
}
bool string(string_t& /*val*/) override
{
return value();
}
bool binary(binary_t& /*val*/) override
{
return value();
}
bool start_object(std::size_t /*elements*/) override
{
value();
stack.push_back('o');
return true;
}
bool key(string_t& /*val*/) override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.back() = 'v';
return true;
}
bool end_object() override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.pop_back();
return true;
}
bool start_array(std::size_t /*elements*/) override
{
value();
stack.push_back('a');
return true;
}
bool end_array() override
{
++events;
assert(!stack.empty() && stack.back() == 'a');
stack.pop_back();
return true;
}
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const nlohmann::detail::exception& /*ex*/) override
{
++errors;
return true;
}
bool complete() const
{
return stack.empty();
}
std::size_t events = 0;
std::size_t errors = 0;
private:
bool value()
{
++events;
if (!stack.empty())
{
// an array element, or the value of a key
assert(stack.back() != 'o');
if (stack.back() == 'v')
{
stack.back() = 'o';
}
}
return true;
}
// 'a' for an array, 'o' for an object that expects a key, 'v' for an
// object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
};
/// parses @a data with a recovering_checker from memory and from a stream,
/// checks that both see the same, that the events are balanced, and that the
/// number of errors is bounded, and returns the checker (see #3989)
inline recovering_checker check_recovering_parse(const std::uint8_t* data, const std::size_t size, const nlohmann::json::input_format_t format)
{
recovering_checker checker;
const bool ok = nlohmann::json::sax_parse(data, data + size, &checker, format);
assert(checker.complete());
assert(checker.errors <= size + 1);
assert(ok == (checker.errors == 0));
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
recovering_checker stream_checker;
assert(nlohmann::json::sax_parse(stream, &stream_checker, format) == ok);
assert(stream_checker.complete());
assert(stream_checker.events == checker.events);
assert(stream_checker.errors == checker.errors);
return checker;
}
} // namespace
+3 -3
View File
@@ -239,7 +239,7 @@ TEST_CASE("controlled bad_alloc")
// iterative path instead, part-way through its worklist.
const auto check_deep_copy = [](bool objects)
{
CAPTURE(objects)
CAPTURE(objects);
next_construct_fails = false;
@@ -315,7 +315,7 @@ struct nth_alloc_fails_allocator : std::allocator<T>
template<class BasicJsonType>
void check_deep_copy_survives_failing_allocation(bool nest_objects)
{
CAPTURE(nest_objects)
CAPTURE(nest_objects);
fail_at_alloc_call = -1;
@@ -352,7 +352,7 @@ void check_deep_copy_survives_failing_allocation(bool nest_objects)
// must come out exactly as it went in
for (std::size_t n = 0; n < total_allocations; ++n)
{
CAPTURE(n)
CAPTURE(n);
alloc_call_count = 0;
fail_at_alloc_call = static_cast<long>(n);
+40
View File
@@ -419,6 +419,46 @@ TEST_CASE("alternative string type")
CHECK(j2.dump() == R"({"/foo/0":"bar","/foo/1":"baz"})");
}
SECTION("error recovery")
{
// a SAX parser that recovers from every error (see #3989)
struct recovering_parser : nlohmann::detail::json_sax_dom_parser<alt_json>
{
explicit recovering_parser(alt_json& j)
: nlohmann::detail::json_sax_dom_parser<alt_json>(j, false)
{}
// sax_parse() calls the SAX parser's own parse_error(), so hiding
// the one of the base class is what recovering takes
// NOLINTNEXTLINE(bugprone-derived-method-shadowing-base-method)
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const nlohmann::detail::exception& /*unused*/)
{
++errors;
return true;
}
std::size_t errors = 0;
};
alt_json j;
recovering_parser sax(j);
// not inside CHECK(): MSVC reads the escape in a stringized raw string
const std::string input = R"([1., "a\qb", tru, {"k" 2}])";
CHECK(!alt_json::sax_parse(input, &sax));
CHECK(sax.errors == 4);
CHECK(j.dump() == R"([1,"aqb",null,{"k":2}])");
// a UBJSON high-precision number, a CBOR key that is not a string
alt_json u;
recovering_parser ubjson_sax(u);
CHECK(!alt_json::sax_parse(std::vector<std::uint8_t> {'[', 'H', 'i', 2, '1', '.', ']'}, &ubjson_sax, alt_json::input_format_t::ubjson));
CHECK(u.dump() == "[1]");
alt_json c;
recovering_parser cbor_sax(c);
CHECK(!alt_json::sax_parse(std::vector<std::uint8_t> {0xA2, 0x01, 0x02, 0x61, 'a', 0x03}, &cbor_sax, alt_json::input_format_t::cbor));
CHECK(c.dump() == R"({"a":3})");
}
SECTION("strict enum")
{
// regression test for #5667: NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json
+1 -1
View File
@@ -18,7 +18,7 @@ DOCTEST_CLANG_SUPPRESS_WARNING("-Wstrict-overflow")
static int assert_counter;
/// set failure variable to true instead of calling assert(x)
#define JSON_ASSERT(x) do { if (!(x)) { ++assert_counter; } } while (false)
#define JSON_ASSERT(x) {if (!(x)) ++assert_counter; }
#include <nlohmann/json.hpp>
using nlohmann::json;
@@ -1,346 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <string>
#include <vector>
namespace
{
struct ill_formed_case
{
const char* name;
std::string bytes;
};
// RFC 3629 ill-formed sequences used throughout this file, plus one
// well-formed sequence for contrast
const std::vector<ill_formed_case> ill_formed_cases =
{
{"overlong", "\xC0\xAE"},
{"lone_0xFF", "\xFF"},
{"truncated", "\xE2\x82"},
{"surrogate", "\xED\xA0\x80"},
};
const std::string valid_sequence = "\xC3\xA9"; // U+00E9, "é"
using eh = json::error_handler_t;
const std::vector<eh> all_handlers = {eh::strict, eh::replace, eh::ignore, eh::keep};
// what dump()+parse() produces for a sanitizing error_handler; this is the
// ground truth every binary writer/reader is checked against
std::string dump_and_parse(const std::string& raw, eh error_handler)
{
return json::parse(json(raw).dump(-1, ' ', false, error_handler)).get<std::string>();
}
} // namespace
TEST_CASE("UTF-8 error_handler for the binary readers and writers")
{
SECTION("writers: string value")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
const json jval = c.bytes;
CHECK_THROWS_AS(json::to_cbor(jval, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jval, false, false, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::strict), json::type_error&);
{
json jobj;
jobj["k"] = jval;
CHECK_THROWS_AS(json::to_bson(jobj, eh::strict), json::type_error&);
}
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(json::to_cbor(jval, h)).get<std::string>() == expected);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, h)).get<std::string>() == expected);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, h)).get<std::string>() == expected);
{
json jobj;
jobj["k"] = jval;
const auto bytes = json::to_bson(jobj, h);
CHECK(json::from_bson(bytes)["k"].get<std::string>() == expected);
}
}
// keep: the writer passes the ill-formed bytes through unchanged,
// exactly as every binary writer did before this parameter existed
CHECK(json::from_cbor(json::to_cbor(jval, eh::keep)).get<std::string>() == c.bytes);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, eh::keep)).get<std::string>() == c.bytes);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::keep)).get<std::string>() == c.bytes);
{
json jobj;
jobj["k"] = jval;
const auto bytes = json::to_bson(jobj, eh::keep);
CHECK(json::from_bson(bytes)["k"].get<std::string>() == c.bytes);
}
}
}
SECTION("writers: object key")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
json jobj;
jobj[c.bytes] = 1;
CHECK_THROWS_AS(json::to_cbor(jobj, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jobj, false, false, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bson(jobj, eh::strict), json::type_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(json::to_cbor(jobj, h)).begin().key() == expected);
CHECK(json::from_ubjson(json::to_ubjson(jobj, false, false, h)).begin().key() == expected);
CHECK(json::from_bjdata(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, h)).begin().key() == expected);
CHECK(json::from_bson(json::to_bson(jobj, h)).begin().key() == expected);
}
// keep: object keys round-trip unchanged too
CHECK(json::from_cbor(json::to_cbor(jobj, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_ubjson(json::to_ubjson(jobj, false, false, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_bjdata(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_bson(json::to_bson(jobj, eh::keep)).begin().key() == c.bytes);
}
}
SECTION("readers: string value")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
// bytes produced the lenient (keep) way, as any binary reader
// accepted them before this parameter existed
const auto cbor_bytes = json::to_cbor(json(c.bytes), eh::keep);
const auto msgpack_bytes = json::to_msgpack(json(c.bytes)); // to_msgpack has no error_handler; always pass-through
const auto ubjson_bytes = json::to_ubjson(json(c.bytes), false, false, eh::keep);
const auto bjdata_bytes = json::to_bjdata(json(c.bytes), false, false, json::bjdata_version_t::draft2, eh::keep);
const auto bson_bytes = [&c]
{
json jobj;
jobj["k"] = c.bytes;
return json::to_bson(jobj, eh::keep);
}();
// keep (the default): bytes are kept unchanged
CHECK(json::from_cbor(cbor_bytes).get<std::string>() == c.bytes);
CHECK(json::from_msgpack(msgpack_bytes).get<std::string>() == c.bytes);
CHECK(json::from_ubjson(ubjson_bytes).get<std::string>() == c.bytes);
CHECK(json::from_bjdata(bjdata_bytes).get<std::string>() == c.bytes);
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == c.bytes);
// strict: parse_error.113, discarded (not thrown) when allow_exceptions is false
CHECK_THROWS_AS(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, eh::strict), json::parse_error&);
CHECK(json::from_cbor(cbor_bytes, true, false, json::cbor_tag_handler_t::error, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_msgpack(msgpack_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_msgpack(msgpack_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_ubjson(ubjson_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_ubjson(ubjson_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_bjdata(bjdata_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_bjdata(bjdata_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_bson(bson_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_bson(bson_bytes, true, false, eh::strict).is_discarded());
// replace / ignore: match what dump() would have sanitized the same bytes to
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, h).get<std::string>() == expected);
CHECK(json::from_msgpack(msgpack_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_ubjson(ubjson_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_bjdata(bjdata_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_bson(bson_bytes, true, true, h)["k"].get<std::string>() == expected);
}
}
}
SECTION("readers: object key")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
json jobj;
jobj[c.bytes] = 1;
const auto cbor_bytes = json::to_cbor(jobj, eh::keep);
const auto msgpack_bytes = json::to_msgpack(jobj);
const auto ubjson_bytes = json::to_ubjson(jobj, false, false, eh::keep);
const auto bjdata_bytes = json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::keep);
const auto bson_bytes = json::to_bson(jobj, eh::keep);
CHECK(json::from_cbor(cbor_bytes).begin().key() == c.bytes);
CHECK(json::from_msgpack(msgpack_bytes).begin().key() == c.bytes);
CHECK(json::from_ubjson(ubjson_bytes).begin().key() == c.bytes);
CHECK(json::from_bjdata(bjdata_bytes).begin().key() == c.bytes);
CHECK(json::from_bson(bson_bytes).begin().key() == c.bytes);
CHECK_THROWS_AS(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_msgpack(msgpack_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_ubjson(ubjson_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_bjdata(bjdata_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_bson(bson_bytes, true, true, eh::strict), json::parse_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, h).begin().key() == expected);
CHECK(json::from_msgpack(msgpack_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_ubjson(ubjson_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_bjdata(bjdata_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_bson(bson_bytes, true, true, h).begin().key() == expected);
}
}
}
SECTION("well-formed UTF-8 is unaffected by error_handler")
{
const json jval = valid_sequence;
json jobj;
jobj[valid_sequence] = valid_sequence;
for (const auto h : all_handlers)
{
CAPTURE(static_cast<int>(h));
CHECK(json::from_cbor(json::to_cbor(jval, h)).get<std::string>() == valid_sequence);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, h)).get<std::string>() == valid_sequence);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, h)).get<std::string>() == valid_sequence);
CHECK(json::from_bson(json::to_bson(jobj, h)).begin().key() == valid_sequence);
CHECK(json::from_cbor(json::to_cbor(jval, eh::keep), true, true, json::cbor_tag_handler_t::error, h).get<std::string>() == valid_sequence);
CHECK(json::from_msgpack(json::to_msgpack(jval), true, true, h).get<std::string>() == valid_sequence);
}
}
SECTION("dump() with error_handler_t::keep writes raw bytes as is")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
const json jval = c.bytes;
const std::string dumped = jval.dump(-1, ' ', false, eh::keep);
CHECK(dumped.find(c.bytes) != std::string::npos);
// even with ensure_ascii, the ill-formed bytes are written as is
const std::string dumped_ascii = jval.dump(-1, ' ', true, eh::keep);
CHECK(dumped_ascii.find(c.bytes) != std::string::npos);
}
// well-formed characters around an ill-formed sequence are still
// escaped as usual under ensure_ascii
const json mixed = valid_sequence + ill_formed_cases[1].bytes; // "é" + lone 0xFF
const std::string dumped_mixed = mixed.dump(-1, ' ', true, eh::keep);
CHECK(dumped_mixed.find("\\u00e9") != std::string::npos);
CHECK(dumped_mixed.find(ill_formed_cases[1].bytes) != std::string::npos);
// the byte that ends an ill-formed sequence is read again, so a quote,
// a backslash, or a control character after it is still escaped, and
// a well-formed code point after it is escaped under ensure_ascii
for (const bool ensure_ascii :
{
false, true
})
{
CAPTURE(ensure_ascii);
CHECK(json("\xC3\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\\"\"");
CHECK(json("\xC3\\").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\\\\"");
CHECK(json("\xC3\n").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\n\"");
CHECK(json("\xE2\x82\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xE2\x82\\\"\"");
CHECK(json("\xFF\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xFF\\\"\"");
CHECK(json("a\xE2\x82").dump(-1, ' ', ensure_ascii, eh::keep) == "\"a\xE2\x82\"");
}
CHECK(json("\xC3\xC3\xA9").dump(-1, ' ', false, eh::keep) == "\"\xC3\xC3\xA9\"");
CHECK(json("\xC3\xC3\xA9").dump(-1, ' ', true, eh::keep) == "\"\xC3\\u00e9\"");
}
SECTION("to_msgpack and to_bon8 are not affected by error_handler")
{
const json jval = ill_formed_cases[1].bytes; // lone 0xFF
// to_msgpack has no error_handler parameter; the bytes are always
// passed through, as MessagePack's spec allows
CHECK(json::from_msgpack(json::to_msgpack(jval)).get<std::string>() == ill_formed_cases[1].bytes);
// to_bon8 has no error_handler parameter either; UTF-8 is structural
// for BON8, so it always rejects ill-formed input
CHECK_THROWS_AS(json::to_bon8(jval), json::type_error&);
}
SECTION("allow_exceptions=false with error_handler_t::strict discards the value")
{
const auto bytes = json::to_cbor(json(ill_formed_cases[0].bytes), eh::keep);
const json result = json::from_cbor(bytes, true, false, json::cbor_tag_handler_t::error, eh::strict);
CHECK(result.is_discarded());
}
SECTION("default parameters are unchanged")
{
const json jval = ill_formed_cases[0].bytes;
// to_*: the default error_handler is strict, so ill-formed input still throws
CHECK_THROWS_AS(json::to_cbor(jval), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jval), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jval), json::type_error&);
{
json jobj;
jobj["k"] = jval;
CHECK_THROWS_AS(json::to_bson(jobj), json::type_error&);
}
// from_*: the default error_handler is keep, so ill-formed bytes are
// still accepted unchanged, exactly as in release 3.12.0
const auto cbor_bytes = json::to_cbor(jval, eh::keep);
CHECK(json::from_cbor(cbor_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto ubjson_bytes = json::to_ubjson(jval, false, false, eh::keep);
CHECK(json::from_ubjson(ubjson_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto bjdata_bytes = json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::keep);
CHECK(json::from_bjdata(bjdata_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto msgpack_bytes = json::to_msgpack(jval);
CHECK(json::from_msgpack(msgpack_bytes).get<std::string>() == ill_formed_cases[0].bytes);
json bson_obj;
bson_obj["k"] = jval;
const auto bson_bytes = json::to_bson(bson_obj, eh::keep);
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == ill_formed_cases[0].bytes);
}
}
+7 -7
View File
@@ -89,7 +89,7 @@ TEST_CASE("binary writer output sinks")
// the first iteration
for (const auto& j : test_values())
{
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace))
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace));
std::vector<std::uint8_t> cbor;
json::to_cbor(j, cbor);
@@ -120,8 +120,8 @@ TEST_CASE("binary writer output sinks")
{
continue; // not a supported combination
}
CAPTURE(use_size)
CAPTURE(use_type)
CAPTURE(use_size);
CAPTURE(use_type);
std::vector<std::uint8_t> ubjson;
json::to_ubjson(j, ubjson, use_size, use_type);
CHECK(json::to_ubjson(j, use_size, use_type) == ubjson);
@@ -141,7 +141,7 @@ TEST_CASE("binary writer output sinks")
for (const auto& j : bson_values())
{
CAPTURE(j.dump())
CAPTURE(j.dump());
std::vector<std::uint8_t> bson;
json::to_bson(j, bson);
CHECK(json::to_bson(j) == bson);
@@ -152,7 +152,7 @@ TEST_CASE("binary writer output sinks")
{
for (const auto& j : test_values())
{
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace))
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace));
const std::vector<std::uint8_t> expected = json::to_cbor(j);
std::vector<char> as_char;
@@ -177,7 +177,7 @@ TEST_CASE("binary_reserve_hint never over-reserves")
{
for (const auto& j : test_values())
{
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace))
CAPTURE(j.dump(-1, ' ', false, json::error_handler_t::replace));
const std::size_t hint = nlohmann::detail::binary_reserve_hint(j);
@@ -194,7 +194,7 @@ TEST_CASE("binary_reserve_hint never over-reserves")
for (const auto& j : bson_values())
{
CAPTURE(j.dump())
CAPTURE(j.dump());
CHECK(nlohmann::detail::binary_reserve_hint(j) <= json::to_bson(j).size());
}
+6 -41
View File
@@ -2523,7 +2523,7 @@ TEST_CASE("BJData")
{"uint8", "int8", "uint16", "int16", "uint32", "int32", "uint64", "int64", "char"
})
{
CAPTURE(type)
CAPTURE(type);
const std::string text = std::string(R"({"_ArrayType_":")") + type +
R"(","_ArraySize_":[2,3],"_ArrayData_":[1,2,3,4,5,6]})";
const auto from_text = json::to_bjdata(json::parse(text));
@@ -2831,7 +2831,7 @@ TEST_CASE("BJData")
R"({"_ArrayType_":"int16","_ArraySize_":[0,2],"_ArrayData_":[]})"
})
{
CAPTURE(text)
CAPTURE(text);
const json j = json::parse(text);
for (const bool use_size :
{
@@ -2865,7 +2865,7 @@ TEST_CASE("BJData")
R"({"_ArrayType_":"int16","_ArraySize_":[],"_ArrayData_":null})"
})
{
CAPTURE(text)
CAPTURE(text);
const json j = json::parse(text);
const auto out = json::to_bjdata(j);
CHECK(out.at(0) == '{');
@@ -3907,41 +3907,6 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_bjdata(j) == v);
CHECK(json::from_bjdata(v) == j);
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_bjdata() is
// strict, so such a value cannot be written back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_bjdata(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(j), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_bjdata(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_bjdata(j_key), json::type_error&);
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_bjdata(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")
@@ -4234,7 +4199,7 @@ TEST_CASE("BJData and UBJSON can be written to a string")
for (const auto& j : values)
{
CAPTURE(j.dump())
CAPTURE(j.dump());
for (const bool use_size :
{
false, true
@@ -4249,8 +4214,8 @@ TEST_CASE("BJData and UBJSON can be written to a string")
{
continue;
}
CAPTURE(use_size)
CAPTURE(use_type)
CAPTURE(use_size);
CAPTURE(use_type);
const auto bjdata = json::to_bjdata(j, use_size, use_type);
std::string bjdata_string;
+2 -47
View File
@@ -154,51 +154,6 @@ TEST_CASE("BSON")
#endif
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// a BSON document {"s": "\xC0\xAE"} (0xC0 0xAE is an overlong
// encoding of '.'); the BSON spec does not require a decoder to
// reject ill-formed UTF-8 in a string value, so the reader hands the
// bytes back unchanged
const std::vector<uint8_t> v =
{
0x0F, 0x00, 0x00, 0x00, // document length
0x02, 's', 0x00, // type 0x02 (string), key "s"
0x03, 0x00, 0x00, 0x00, // string length (including null)
0xc0, 0xae, 0x00, // string content and its null terminator
0x00 // document terminator
};
json j;
CHECK_NOTHROW(j = json::from_bson(v));
REQUIRE(j.is_object());
REQUIRE(j.contains("s"));
CHECK(j["s"].get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value
CHECK_THROWS_AS(j.dump(), json::type_error&);
// to_bson() is strict as well, so the value cannot be written back
CHECK_THROWS_AS(json::to_bson(j), json::type_error&);
// to_bson() rejects the same kind of ill-formed string value, before
// any bytes reach the output adapter (the BSON document length
// prefix must be known up front, so nothing is written incrementally)
std::vector<std::uint8_t> out{0x42}; // a sentinel byte the writer must not touch
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
CHECK(out == std::vector<std::uint8_t> {0x42});
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC3"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected as well; unlike
// the reader (which never validates element names), the writer
// checks both string values and object keys
CHECK_THROWS_WITH_AS(json::to_bson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
{
// out_of_range.412 is thrown from a single shared helper
@@ -1797,7 +1752,7 @@ TEST_CASE("BSON: deeply nested values")
json value = "leaf";
for (std::size_t depth = 0; depth <= 300; ++depth)
{
CAPTURE(depth)
CAPTURE(depth);
const json document = {{"value", value}, {"n", depth}};
CHECK(json::from_bson(json::to_bson(document)) == document);
@@ -1845,7 +1800,7 @@ value = depth % 2 == 0 ? json{{"a", std::move(value)}, {"b", {1, "x"}}} :
false, true
})
{
CAPTURE(objects)
CAPTURE(objects);
std::string text = "{\"a\":";
for (std::size_t i = 0; i < depth; ++i)
{
+30 -30
View File
@@ -17,7 +17,7 @@ TEST_CASE("capacity")
{
SECTION("boolean")
{
json j = true;
json j = true; // NOLINT(misc-const-correctness)
const json j_const = true;
SECTION("result of empty")
@@ -35,7 +35,7 @@ TEST_CASE("capacity")
SECTION("string")
{
json j = "hello world";
json j = "hello world"; // NOLINT(misc-const-correctness)
const json j_const = "hello world";
SECTION("result of empty")
@@ -55,7 +55,7 @@ TEST_CASE("capacity")
{
SECTION("empty array")
{
json j = json::array();
json j = json::array(); // NOLINT(misc-const-correctness)
const json j_const = json::array();
SECTION("result of empty")
@@ -73,7 +73,7 @@ TEST_CASE("capacity")
SECTION("filled array")
{
json j = {1, 2, 3};
json j = {1, 2, 3}; // NOLINT(misc-const-correctness)
const json j_const = {1, 2, 3};
SECTION("result of empty")
@@ -94,7 +94,7 @@ TEST_CASE("capacity")
{
SECTION("empty object")
{
json j = json::object();
json j = json::object(); // NOLINT(misc-const-correctness)
const json j_const = json::object();
SECTION("result of empty")
@@ -112,7 +112,7 @@ TEST_CASE("capacity")
SECTION("filled object")
{
json j = {{"one", 1}, {"two", 2}, {"three", 3}};
json j = {{"one", 1}, {"two", 2}, {"three", 3}}; // NOLINT(misc-const-correctness)
const json j_const = {{"one", 1}, {"two", 2}, {"three", 3}};
SECTION("result of empty")
@@ -131,7 +131,7 @@ TEST_CASE("capacity")
SECTION("number (integer)")
{
json j = -23;
json j = -23; // NOLINT(misc-const-correctness)
const json j_const = -23;
SECTION("result of empty")
@@ -149,7 +149,7 @@ TEST_CASE("capacity")
SECTION("number (unsigned)")
{
json j = 23u;
json j = 23u; // NOLINT(misc-const-correctness)
const json j_const = 23u;
SECTION("result of empty")
@@ -167,7 +167,7 @@ TEST_CASE("capacity")
SECTION("number (float)")
{
json j = 23.42;
json j = 23.42; // NOLINT(misc-const-correctness)
const json j_const = 23.42;
SECTION("result of empty")
@@ -185,7 +185,7 @@ TEST_CASE("capacity")
SECTION("null")
{
json j = nullptr;
json j = nullptr; // NOLINT(misc-const-correctness)
const json j_const = nullptr;
SECTION("result of empty")
@@ -206,7 +206,7 @@ TEST_CASE("capacity")
{
SECTION("boolean")
{
json j = true;
json j = true; // NOLINT(misc-const-correctness)
const json j_const = true;
SECTION("result of size")
@@ -226,7 +226,7 @@ TEST_CASE("capacity")
SECTION("string")
{
json j = "hello world";
json j = "hello world"; // NOLINT(misc-const-correctness)
const json j_const = "hello world";
SECTION("result of size")
@@ -248,7 +248,7 @@ TEST_CASE("capacity")
{
SECTION("empty array")
{
json j = json::array();
json j = json::array(); // NOLINT(misc-const-correctness)
const json j_const = json::array();
SECTION("result of size")
@@ -268,7 +268,7 @@ TEST_CASE("capacity")
SECTION("filled array")
{
json j = {1, 2, 3};
json j = {1, 2, 3}; // NOLINT(misc-const-correctness)
const json j_const = {1, 2, 3};
SECTION("result of size")
@@ -291,7 +291,7 @@ TEST_CASE("capacity")
{
SECTION("empty object")
{
json j = json::object();
json j = json::object(); // NOLINT(misc-const-correctness)
const json j_const = json::object();
SECTION("result of size")
@@ -311,7 +311,7 @@ TEST_CASE("capacity")
SECTION("filled object")
{
json j = {{"one", 1}, {"two", 2}, {"three", 3}};
json j = {{"one", 1}, {"two", 2}, {"three", 3}}; // NOLINT(misc-const-correctness)
const json j_const = {{"one", 1}, {"two", 2}, {"three", 3}};
SECTION("result of size")
@@ -332,7 +332,7 @@ TEST_CASE("capacity")
SECTION("number (integer)")
{
json j = -23;
json j = -23; // NOLINT(misc-const-correctness)
const json j_const = -23;
SECTION("result of size")
@@ -352,7 +352,7 @@ TEST_CASE("capacity")
SECTION("number (unsigned)")
{
json j = 23u;
json j = 23u; // NOLINT(misc-const-correctness)
const json j_const = 23u;
SECTION("result of size")
@@ -372,7 +372,7 @@ TEST_CASE("capacity")
SECTION("number (float)")
{
json j = 23.42;
json j = 23.42; // NOLINT(misc-const-correctness)
const json j_const = 23.42;
SECTION("result of size")
@@ -392,7 +392,7 @@ TEST_CASE("capacity")
SECTION("null")
{
json j = nullptr;
json j = nullptr; // NOLINT(misc-const-correctness)
const json j_const = nullptr;
SECTION("result of size")
@@ -415,7 +415,7 @@ TEST_CASE("capacity")
{
SECTION("boolean")
{
json j = true;
json j = true; // NOLINT(misc-const-correctness)
const json j_const = true;
SECTION("result of max_size")
@@ -427,7 +427,7 @@ TEST_CASE("capacity")
SECTION("string")
{
json j = "hello world";
json j = "hello world"; // NOLINT(misc-const-correctness)
const json j_const = "hello world";
SECTION("result of max_size")
@@ -441,7 +441,7 @@ TEST_CASE("capacity")
{
SECTION("empty array")
{
json j = json::array();
json j = json::array(); // NOLINT(misc-const-correctness)
const json j_const = json::array();
SECTION("result of max_size")
@@ -453,7 +453,7 @@ TEST_CASE("capacity")
SECTION("filled array")
{
json j = {1, 2, 3};
json j = {1, 2, 3}; // NOLINT(misc-const-correctness)
const json j_const = {1, 2, 3};
SECTION("result of max_size")
@@ -468,7 +468,7 @@ TEST_CASE("capacity")
{
SECTION("empty object")
{
json j = json::object();
json j = json::object(); // NOLINT(misc-const-correctness)
const json j_const = json::object();
SECTION("result of max_size")
@@ -480,7 +480,7 @@ TEST_CASE("capacity")
SECTION("filled object")
{
json j = {{"one", 1}, {"two", 2}, {"three", 3}};
json j = {{"one", 1}, {"two", 2}, {"three", 3}}; // NOLINT(misc-const-correctness)
const json j_const = {{"one", 1}, {"two", 2}, {"three", 3}};
SECTION("result of max_size")
@@ -493,7 +493,7 @@ TEST_CASE("capacity")
SECTION("number (integer)")
{
json j = -23;
json j = -23; // NOLINT(misc-const-correctness)
const json j_const = -23;
SECTION("result of max_size")
@@ -505,7 +505,7 @@ TEST_CASE("capacity")
SECTION("number (unsigned)")
{
json j = 23u;
json j = 23u; // NOLINT(misc-const-correctness)
const json j_const = 23u;
SECTION("result of max_size")
@@ -517,7 +517,7 @@ TEST_CASE("capacity")
SECTION("number (float)")
{
json j = 23.42;
json j = 23.42; // NOLINT(misc-const-correctness)
const json j_const = 23.42;
SECTION("result of max_size")
@@ -529,7 +529,7 @@ TEST_CASE("capacity")
SECTION("null")
{
json j = nullptr;
json j = nullptr; // NOLINT(misc-const-correctness)
const json j_const = nullptr;
SECTION("result of max_size")
+20 -67
View File
@@ -1801,40 +1801,19 @@ TEST_CASE("CBOR")
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
SECTION("invalid UTF-8 in string (see #5529)")
{
// RFC 8949 §3.1 leaves it up to the decoder whether to reject
// ill-formed UTF-8 in a text string; this library does not, and
// hands the original bytes back unchanged, matching the
// MessagePack reader and the behavior before #5185/#5531 (not in
// any release)
// a two-character text string (major type 3) whose bytes are not
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') round-trips
// byte for byte as a string value
const std::vector<uint8_t> ill_formed_value = {0x62, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_cbor(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// to_cbor() is strict as well, so the value cannot be written back
CHECK_THROWS_AS(json::to_cbor(j_value), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0xa1, 0x62, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_cbor(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_cbor(j_key), json::type_error&);
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') must be
// rejected at decode time, matching every other kind of
// malformed binary input, rather than only failing later when
// the resulting value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae}), true, false).is_discarded());
// a CBOR byte string (major type 2) with the very same bytes is
// NOT text and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x42, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
@@ -1843,46 +1822,17 @@ TEST_CASE("CBOR")
CHECK(json::from_cbor(json::to_cbor(j)) == j);
}
SECTION("to_cbor rejects ill-formed UTF-8 (see #5651)")
{
// to_cbor() must reject the same ill-formed strings from_cbor()
// rejects, so a value it accepts can always be read back
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_cbor(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
}
SECTION("ill-formed UTF-8 in indefinite-length string")
SECTION("invalid UTF-8 in indefinite-length string")
{
json _;
// the chunks are concatenated as is, without checking that each
// chunk is valid UTF-8 on its own (RFC 8949, Section 3.2.3), so
// a code point split across two chunks yields a valid string
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})));
CHECK(_ == "\xc3\xa9");
CHECK(_.dump() == "\"\xc3\xa9\"");
// every chunk must be valid UTF-8 on its own (RFC 8949, Section
// 3.2.3), so a code point split across two chunks is rejected
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff}), true, false).is_discarded());
// a truncated code point is kept as is
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0xff})));
CHECK(_ == "\xc3");
CHECK_THROWS_AS(_.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_cbor(_), json::type_error&);
// an ill-formed later chunk is kept after valid ones
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})));
CHECK(_ == "\xc3\xa9\xc0\xae");
CHECK_THROWS_AS(_.dump(), json::type_error&);
// an ill-formed later chunk is rejected after valid ones
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})), "[json.exception.parse_error.113] parse error at byte 7: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
// valid multi-byte chunks are accepted
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc3, 0xb6, 0xff})) == "\xc3\xa9\xc3\xb6");
@@ -1890,6 +1840,9 @@ TEST_CASE("CBOR")
SECTION("many chunks in indefinite-length string")
{
// only the newly read chunk is validated, not the whole string
// collected so far; validating the latter made this input take
// quadratic time (about ten seconds for 100000 chunks)
constexpr std::size_t chunks = 100000;
std::vector<uint8_t> v{0x7f};
for (std::size_t i = 0; i < chunks; ++i)
@@ -2925,7 +2878,7 @@ TEST_CASE("Tagged values")
0xD5, 0xD6, 0xD7
})
{
CAPTURE(b)
CAPTURE(b);
// add tag to value
auto v_tagged = v;
@@ -3265,7 +3218,7 @@ TEST_CASE("CBOR large strings and binaries (chunked reader)")
std::size_t{4097}, std::size_t{8192}, std::size_t{100000}
})
{
CAPTURE(len)
CAPTURE(len);
// text string
const json j_string = std::string(len, 'x');
+12 -12
View File
@@ -272,7 +272,7 @@ TEST_CASE("lexer number fast path")
std::stringstream ss(doc);
const json b = json::parse(ss);
CAPTURE(n)
CAPTURE(n);
CHECK(a == b);
CHECK(a.dump() == b.dump());
CHECK(a[0].type() == b[0].type());
@@ -310,7 +310,7 @@ TEST_CASE("lexer number fast path")
for (const auto& n : numbers)
{
CAPTURE(n)
CAPTURE(n);
const std::string doc = "[" + n + "]";
const json a = json::parse(doc); // contiguous fast path
@@ -347,7 +347,7 @@ TEST_CASE("lexer number fast path")
{"-", "1.", "1e", "1e+", "1.2e", "01", "-01", "1..2", "1.2.3"
})
{
CAPTURE(bad)
CAPTURE(bad);
// the contiguous fast path must decline and let the byte path report
const std::string doc = std::string("[") + bad + "]";
CHECK_FALSE(json::accept(doc));
@@ -415,7 +415,7 @@ TEST_CASE("lexer number fast path")
// 7 + 49 + 343 + 2401 tokens
CHECK(tokens.size() == 2401);
CAPTURE(mismatches)
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
@@ -459,7 +459,7 @@ TEST_CASE("lexer number fast path")
"[1 \n2]", "[\n1\n2]", "1\n2", "[01\r\n]", "[1e\n]", "[-\n]"
})
{
CAPTURE(bad)
CAPTURE(bad);
const std::string doc = bad;
const std::string contiguous_what = contiguous_error(doc);
@@ -576,7 +576,7 @@ TEST_CASE("lexer string fast path")
// 13 + 169 + 2197 tokens, each at two offsets
CHECK(tokens.size() == 2197);
CAPTURE(mismatches)
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
@@ -604,7 +604,7 @@ TEST_CASE("lexer string fast path")
}
}
}
CAPTURE(mismatches)
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
#endif
@@ -649,10 +649,10 @@ TEST_CASE("lexer string fast path")
for (const auto& test_case : cases)
{
CAPTURE(test_case.description)
CAPTURE(test_case.description);
for (const std::size_t offset : offsets)
{
CAPTURE(offset)
CAPTURE(offset);
const std::string doc = "[\"" + std::string(offset, 'a') + test_case.sequence + "\"]";
CHECK(json::accept(doc) == test_case.valid);
#if !defined(JSON_NOEXCEPTION)
@@ -1236,7 +1236,7 @@ TEST_CASE("Eisel-Lemire float conversion")
for (const auto& c : known)
{
CAPTURE(c.first)
CAPTURE(c.first);
double out = 0;
if (eisel_lemire(c.first, out))
{
@@ -1277,7 +1277,7 @@ TEST_CASE("Eisel-Lemire float conversion")
std::array<char, 64> buffer{};
const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), d);
const std::string token(buffer.data(), static_cast<std::size_t>(end - buffer.data()));
CAPTURE(token)
CAPTURE(token);
double out = 0;
REQUIRE(eisel_lemire(token, out));
CHECK(bits_of(out) == b);
@@ -1289,7 +1289,7 @@ TEST_CASE("Eisel-Lemire float conversion")
const std::size_t dot = longer.find('.');
const std::string extra = dot == std::string::npos ? ".000000000000000000001" : "000000000000000000001";
longer.insert(e == std::string::npos ? longer.size() : e, extra);
CAPTURE(longer)
CAPTURE(longer);
if (eisel_lemire(longer, out))
{
CHECK(bits_of(out) == b);
+585 -3
View File
@@ -143,11 +143,13 @@ class SaxEventLogger
{
errored = true;
events.push_back("parse_error(" + std::to_string(position) + ")");
return false;
return recover;
}
std::vector<std::string> events {}; // NOLINT(readability-redundant-member-init)
bool errored = false;
/// whether parse_error() asks the parser to recover from the error (see #3989)
bool recover = false;
};
class SaxCountdown : public nlohmann::json::json_sax_t
@@ -2561,7 +2563,7 @@ TEST_CASE("last-read diagnostics are identical across input adapters")
for (const auto& s : inputs)
{
CAPTURE(s)
CAPTURE(s);
// reference: contiguous std::string -> seekable (lazy) path
const std::string reference = parse_error_message(s);
@@ -2645,7 +2647,7 @@ TEST_CASE("diagnostic positions: value lifetime, input adapters, and SAX")
SECTION("move constructor resets the moved-from value to npos")
{
// basic_json(basic_json&&) (json.hpp, around line 1951) copies
// basic_json(basic_json&&) (json.hpp, around line 1265) copies
// other's start_position/end_position into *this and then resets
// other's to npos (see the cppcheck-suppress[accessForwarded]
// annotation there, which flags this reset as worth a second
@@ -2937,3 +2939,583 @@ TEST_CASE("diagnostic positions: value lifetime, input adapters, and SAX")
}
}
#endif
namespace
{
/// builds a value like json::parse(), but asks the parser to recover from
/// errors (see #3989), and checks that the events it receives are balanced
class RecoveringDomParser
{
public:
explicit RecoveringDomParser(json& j, std::size_t max_errors_ = static_cast<std::size_t>(-1))
: dom(j, false)
, max_errors(max_errors_)
{}
bool null()
{
value();
return dom.null();
}
bool boolean(bool val)
{
value();
return dom.boolean(val);
}
bool number_integer(json::number_integer_t val)
{
value();
return dom.number_integer(val);
}
bool number_unsigned(json::number_unsigned_t val)
{
value();
return dom.number_unsigned(val);
}
bool number_float(json::number_float_t val, const std::string& s)
{
value();
return dom.number_float(val, s);
}
bool string(std::string& val)
{
value();
return dom.string(val);
}
bool binary(json::binary_t& val)
{
value();
return dom.binary(val);
}
bool start_object(std::size_t elements)
{
value();
stack.push_back('o');
return dom.start_object(elements);
}
bool key(std::string& val)
{
++events;
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.back() = 'v';
return dom.key(val);
}
bool end_object()
{
++events;
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_object();
}
bool start_array(std::size_t elements)
{
value();
stack.push_back('a');
return dom.start_array(elements);
}
bool end_array()
{
++events;
if (stack.empty() || stack.back() != 'a')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
errors.emplace_back(ex.what());
return errors.size() < max_errors;
}
/// whether the events were balanced and every key was followed by a value
bool balanced() const
{
return well_formed && stack.empty();
}
/// builds the value
nlohmann::detail::json_sax_dom_parser<json> dom;
std::vector<std::string> errors {}; // NOLINT(readability-redundant-member-init)
std::size_t events = 0;
/// the open containers: 'a' for an array, 'o' for an object that expects
/// a key, 'v' for an object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
std::size_t max_errors;
private:
/// a value is passed: it is an array element, or the value of a key
void value()
{
++events;
if (!stack.empty())
{
if (stack.back() == 'v')
{
stack.back() = 'o';
}
else if (stack.back() == 'o')
{
// a value without a key
well_formed = false;
}
}
}
};
struct RecoveryResult
{
json value;
std::vector<std::string> errors;
std::size_t events;
bool ok;
bool balanced;
};
template<typename InputType>
RecoveryResult parse_recovering(InputType&& input, const bool strict = true,
const bool ignore_comments = false, const bool ignore_trailing_commas = false)
{
json j;
RecoveringDomParser sax(j);
const bool ok = json::sax_parse(std::forward<InputType>(input), &sax, json::input_format_t::json,
strict, ignore_comments, ignore_trailing_commas);
return {j, sax.errors, sax.events, ok, sax.balanced()};
}
/// stops after a number of events, but recovers from errors
class RecoveringCountdown : public SaxCountdown
{
public:
using SaxCountdown::SaxCountdown;
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const json::exception& /*ex*/) override
{
return true;
}
};
/// a repaired input: the value it is repaired to, and the number of errors
struct Repair
{
const char* input;
const char* expected;
std::size_t errors;
};
} // namespace
TEST_CASE("parser error recovery (#3989)")
{
SECTION("repairs")
{
const std::vector<Repair> repairs =
{
// a missing separator is inserted
{"[1 2]", "[1,2]", 1},
{R"({"a":1 "b":2})", R"({"a":1,"b":2})", 1},
{R"({"a" 1})", R"({"a":1})", 1},
{"[1 tru 2]", "[1,null,2]", 2},
{R"({"a" "b": 1})", R"({"a":"b"})", 2},
// a missing value is null in an object; in an array, a ',' stands
// for null, while an array that ends there just ends
{R"({"a":})", R"({"a":null})", 1},
{R"({"a"})", R"({"a":null})", 1},
{R"({"a","b":1})", R"({"a":null,"b":1})", 1},
{"[1,,2]", "[1,null,2]", 1},
{"[,1]", "[null,1]", 1},
{"[1,]", "[1]", 1},
{"[1,2,3,]", "[1,2,3]", 1},
{R"({"a":1,})", R"({"a":1})", 1},
// a broken string keeps what can be read
{R"(["a\qb"])", R"(["aqb"])", 1},
{R"({"na\me":1})", R"({"name":1})", 1},
{"[\"\xFF\"]", R"(["\uFFFD"])", 1},
{"[\"a\xC3(\"]", R"(["a\uFFFD("])", 1},
{"[\"\xE2\x82\"]", R"(["\uFFFD"])", 1},
{"[\"\xC3\\\\\", 1]", R"(["\uFFFD\\",1])", 1},
{R"(["\u12"])", R"(["\uFFFD"])", 1},
{R"(["\u12G4"])", R"(["\uFFFDG4"])", 1},
{R"(["\uDC00x"])", R"(["\uFFFDx"])", 1},
{R"(["\uD800x"])", R"(["\uFFFDx"])", 1},
{R"(["\uD800\u0041"])", R"(["\uFFFDA"])", 1},
{R"(["\uD800\uD800\uDC00"])", R"(["\uFFFD\uD800\uDC00"])", 1},
{R"(["\uD800\uD800\uD800x"])", R"(["\uFFFD\uFFFD\uFFFDx"])", 1},
{
R"(["\uD800\"x", 1])", R"(["\uFFFD\"x",1])", 1
},
{R"(["\uD800\q"])", R"(["\uFFFDq"])", 1},
{"[\"a\tb\"]", R"(["a\tb"])", 1},
{R"(["a\qb\u0041\x"])", R"(["aqbAx"])", 1},
// a broken number keeps its longest valid prefix
{"[1.]", "[1]", 1},
{"[-2.]", "[-2]", 1},
{"[1.5e]", "[1.5]", 1},
{"[1e+]", "[1]", 1},
{"[1.x2, 3]", "[1,3]", 1},
// what cannot be read at all is null
{"[1,NaN,3]", "[1,null,3]", 1},
{"[tru]", "[null]", 1},
{"[-]", "[null]", 1},
{R"({"a":Infinity})", R"({"a":null})", 1},
// a stray token is dropped
{"[:1]", "[1]", 1},
{R"(["a":1])", R"(["a",1])", 1},
{R"({"a"::1})", R"({"a":1})", 1},
// a member that cannot be read is skipped
{R"({1:2,"b":3})", R"({"b":3})", 1},
{R"({"a":1 2})", R"({"a":1})", 1},
{R"({,"a":1})", R"({"a":1})", 1},
{R"({"a":1,,"b":2})", R"({"a":1,"b":2})", 1},
{"{a:1}", "{}", 1},
{R"({"a":1 [1,{"b":2}], "c":3})", R"({"a":1,"c":3})", 1},
{R"([{1}, "a"])", R"([{},"a"])", 1},
// a wrong closing bracket closes the innermost container
{R"({"a":[1,2}, "b":3})", R"({"a":[1,2],"b":3})", 1},
{R"([{"a":1], 2])", R"([{"a":1},2])", 1},
{"{]", "{}", 1},
{"[}", "[]", 1},
// the end of the input closes all containers
{R"({"a":[1,2)", R"({"a":[1,2]})", 1},
{"[", "[]", 1},
{"{", "{}", 1},
{R"({"a")", R"({"a":null})", 1},
{R"({"a":)", R"({"a":null})", 1},
{"[1,", "[1]", 1},
{"[[[1", "[[[1]]]", 1},
{
R"(["abc)", R"(["abc"])", 2
},
{"[1,tr", "[1,null]", 2},
{"\"abc", "\"abc\"", 1},
{"[\"ab\ncd\"]", R"(["ab",null,"]"])", 4},
// what comes before the top-level value is skipped
{")]}'\n{\"a\":1}", R"({"a":1})", 1},
{R"(data: {"a":1})", R"({"a":1})", 1},
{"\xEF\xBB[1]", "[1]", 1},
// what comes after it is an error that ends parsing
{R"({"a":1}})", R"({"a":1})", 1},
{"[1}]", "[1]", 2},
{"[1] [2]", "[1]", 1},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.input);
const auto result = parse_recovering(std::string(repair.input));
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.value == json::parse(repair.expected));
CHECK(result.errors.size() == repair.errors);
}
}
SECTION("number overflow")
{
const auto result = parse_recovering(std::string("[1e999,-1e999]"));
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors.size() == 2);
CHECK(result.errors[0] == "[json.exception.out_of_range.406] number overflow parsing '1e999'");
REQUIRE(result.value.size() == 2);
CHECK(result.value[0].is_number_float());
CHECK(result.value[0].get<double>() == std::numeric_limits<double>::infinity());
CHECK(result.value[1].get<double>() == -std::numeric_limits<double>::infinity());
// the SAX parser gets the number's text
SaxEventLogger logger;
logger.recover = true;
CHECK(!json::sax_parse("1e999", &logger));
CHECK(logger.events == std::vector<std::string>({"parse_error(5)", "number_float(1e999)"}));
}
SECTION("nothing to recover")
{
for (const std::string s :
{
"", " ", "]", "tru", "NaN", ",:", "/* comment"
})
{
CAPTURE(s);
const auto result = parse_recovering(s, true, true);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.events == 0);
CHECK(result.value == nullptr);
CHECK(result.errors.size() == 1);
}
}
SECTION("error messages")
{
// the first error is reported as without recovery
for (const std::string s :
{
"[1 2]", R"({"a":1 "b":2})", R"({"a" 1})", R"({"a":})", "[1,]", "[1.]",
R"(["a\qb"])", "[1e999]", "{1:2}", R"({"a":[1,2}})", "[1,", "[1] [2]", "{a:1}"
})
{
CAPTURE(s);
const auto result = parse_recovering(s);
REQUIRE(!result.errors.empty());
json _;
CHECK_THROWS_WITH_STD_STR(_ = json::parse(s), result.errors.front());
}
// the token of an error begins where the previous error was
const auto result = parse_recovering(std::string("[tru, fals, nul]"));
CHECK(result.errors == std::vector<std::string>(
{
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: '[tru,'",
"[json.exception.parse_error.101] parse error at line 1, column 11: syntax error while parsing value - invalid literal; last read: ', fals,'",
"[json.exception.parse_error.101] parse error at line 1, column 16: syntax error while parsing value - invalid literal; last read: ', nul]'"
}));
CHECK(result.value == json::parse("[null,null,null]"));
}
SECTION("events")
{
// see #4522
SaxEventLogger logger;
logger.recover = true;
CHECK(!json::sax_parse(R"([{1}, "a"])", &logger));
CHECK(logger.events == std::vector<std::string>(
{
"start_array()", "start_object()", "parse_error(3)", "end_object()", "string(a)", "end_array()"
}));
}
SECTION("options")
{
SECTION("strict")
{
const auto result = parse_recovering(std::string("[1 2] [3]"), false);
CHECK(!result.ok);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
}
SECTION("ignore_trailing_commas")
{
for (const std::string s :
{
"[1,]", R"({"a":1,})", "[[1,],]"
})
{
CAPTURE(s);
const auto result = parse_recovering(s, true, false, true);
CHECK(result.ok);
CHECK(result.errors.empty());
}
auto result = parse_recovering(std::string("[1,,]"), true, false, true);
CHECK(result.value == json::parse("[1,null]"));
CHECK(result.errors.size() == 1);
result = parse_recovering(std::string(R"({"a":1,,})"), true, false, true);
CHECK(result.value == json::parse(R"({"a":1})"));
CHECK(result.errors.size() == 1);
}
SECTION("ignore_comments")
{
auto result = parse_recovering(std::string("[1 /* one */ 2]"), true, true);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
// a comment that is not closed runs to the end of the input, which
// is not reported again
result = parse_recovering(std::string("[1, 2 /* unterminated"), true, true);
CHECK(result.balanced);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
// a '/' that does not begin a comment is garbage
result = parse_recovering(std::string("[1, /x, 2]"), true, true);
CHECK(result.balanced);
CHECK(result.value == json::parse("[1,null,2]"));
CHECK(result.errors.size() == 1);
}
}
SECTION("null bytes")
{
// a null byte ends the input, unless JSON_STRICT_NUL_HANDLING is set
const auto result = parse_recovering(std::string("[1,\0x", 5));
CHECK(result.balanced);
CHECK(!result.ok);
#ifdef JSON_TEST_STRICT_NUL_HANDLING_ENABLED
CHECK(result.value == json::parse("[1,null]"));
#else
CHECK(result.value == json::parse("[1]"));
CHECK(result.errors.size() == 1);
#endif
const auto in_string = parse_recovering(std::string("[\"a\0b\"]", 7));
CHECK(in_string.balanced);
#ifdef JSON_TEST_STRICT_NUL_HANDLING_ENABLED
CHECK(in_string.value == json::array({std::string("a\0b", 3)}));
#else
CHECK(in_string.value == json::parse(R"(["a"])"));
#endif
}
SECTION("the SAX parser stops recovering")
{
json j;
RecoveringDomParser sax(j, 2);
CHECK(!json::sax_parse("[1 2 3 4 5]", &sax));
CHECK(sax.errors.size() == 2);
// an error at a delimiter that an invalid token consumed is reported
// to the SAX parser, too
json j2;
RecoveringDomParser sax2(j2, 2);
CHECK(!json::sax_parse("[tru}, 1]", &sax2));
CHECK(sax2.errors.size() == 2);
}
SECTION("an event stops parsing during a repair")
{
// start_object() and key() are passed, then null() for the missing
// value returns false
RecoveringCountdown countdown(2);
CHECK(!json::sax_parse(R"({"a":})", &countdown));
// the end of the input: end_array() for the second array returns false
RecoveringCountdown countdown2(4);
CHECK(!json::sax_parse("[[1", &countdown2));
}
SECTION("input adapters")
{
// the lexer reads contiguous and streaming input differently, and it
// puts back a character that ended an invalid token
for (const std::string s :
{
"[1 2]", "[tru}, 1]", R"({"a" "b\q", "c":[1.x, 2}})", "[\"\xFF\xC3(\", -, 1e+]", "{a:1,\"b\":2", ")]}' [1]"
})
{
CAPTURE(s);
const auto reference = parse_recovering(s);
CHECK(reference.balanced);
const auto from_c_string = parse_recovering(s.c_str());
CHECK(from_c_string.value == reference.value);
CHECK(from_c_string.errors == reference.errors);
const std::list<char> l(s.begin(), s.end());
json j;
RecoveringDomParser sax(j);
CHECK(!json::sax_parse(l.begin(), l.end(), &sax));
CHECK(j == reference.value);
CHECK(sax.errors == reference.errors);
std::istringstream ss(s);
const auto from_stream = parse_recovering(ss);
CHECK(from_stream.value == reference.value);
CHECK(from_stream.errors == reference.errors);
}
}
SECTION("long runs of errors")
{
// no error may copy all the input read before it
const auto closing = parse_recovering("[" + std::string(100000, '}'));
CHECK(closing.balanced);
CHECK(closing.value == json::array());
const auto garbage = parse_recovering("[" + std::string(100000, 'x') + "]");
CHECK(garbage.balanced);
CHECK(garbage.errors.size() == 1);
const auto commas = parse_recovering("{" + std::string(100000, ',') + "}");
CHECK(commas.balanced);
CHECK(commas.value == json::object());
}
SECTION("mutations of valid input")
{
// whatever the input, the events are balanced, every error is reported
// at most once, and valid input is parsed as usual
const std::vector<std::string> documents =
{
R"({"name": "value", "list": [1, -2.5, true, null, {"x": [[]]}], "e": "\u00e9"})",
R"([{"a": [1, 2, {"b": "c"}]}, [], {}, "\ud83d\ude00", 1e10])",
"{\"\xC3\xA9\": \"\xF0\x9F\x98\x80\"}",
R"( {"k" : [ "v" , 0 ] } )",
};
// each character that can be inserted, including a null byte
const std::string insertions("[]{},:\"x\\\0\xFF", 11);
std::vector<std::string> inputs;
for (const auto& doc : documents)
{
for (std::size_t i = 0; i <= doc.size(); ++i)
{
inputs.push_back(doc.substr(0, i));
if (i < doc.size())
{
inputs.push_back(doc.substr(0, i) + doc.substr(i + 1));
}
for (const char c : insertions)
{
inputs.push_back(doc.substr(0, i) + c + doc.substr(i));
}
}
}
for (const auto& s : inputs)
{
CAPTURE(s);
const auto result = parse_recovering(s);
CHECK(result.balanced);
CHECK(result.errors.size() <= s.size() + 1);
CHECK(result.events <= (4 * s.size()) + 4);
if (json::accept(s))
{
CHECK(result.ok);
CHECK(result.errors.empty());
CHECK(result.value == json::parse(s));
}
else
{
CHECK(!result.ok);
CHECK(!result.errors.empty());
}
}
}
}
+3 -3
View File
@@ -857,7 +857,7 @@ TEST_CASE("equality of objects whose entries have no fixed order")
for (const std::size_t depth : std::vector<std::size_t> {0, 200})
{
CAPTURE(depth)
CAPTURE(depth);
const unordered_json descending = nest(make_unordered_object(true), depth);
const unordered_json ascending = nest(make_unordered_object(false), depth);
@@ -909,7 +909,7 @@ TEST_CASE("equality of an object whose comparator treats different keys as equiv
for (const std::size_t depth : std::vector<std::size_t> {0, 127, 128, 200})
{
CAPTURE(depth)
CAPTURE(depth);
const ci_json x = nest(a, depth);
const ci_json y = nest(b, depth);
@@ -935,7 +935,7 @@ TEST_CASE("containers are compared element by element")
for (const std::size_t depth : std::vector<std::size_t> {0, 200})
{
CAPTURE(depth)
CAPTURE(depth);
// objects with different keys
{
+1 -1
View File
@@ -85,7 +85,7 @@ TEST_CASE("other constructors and destructor")
CHECK(j.type() == json::value_t::object);
const json k(std::move(j));
CHECK(k.type() == json::value_t::object);
CHECK(j.type() == json::value_t::null); // NOLINT(bugprone-use-after-move,hicpp-invalid-access-moved) access after move is OK here
CHECK(j.type() == json::value_t::null); // NOLINT: access after move is OK here
}
SECTION("copy assignment")
+4 -4
View File
@@ -1640,7 +1640,7 @@ TEST_CASE("value conversion")
enum class cards {kreuz, pik, herz, karo};
// NOLINTNEXTLINE(misc-use-internal-linkage,misc-const-correctness) - false positive
// NOLINTNEXTLINE(misc-use-internal-linkage,misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) - false positive
NLOHMANN_JSON_SERIALIZE_ENUM(cards,
{
{cards::kreuz, "kreuz"},
@@ -1658,7 +1658,7 @@ enum TaskState // NOLINT(cert-int09-c,readability-enum-initial-value,cppcoreguid
TS_INVALID = -1,
};
// NOLINTNEXTLINE(misc-const-correctness,misc-use-internal-linkage) - false positive
// NOLINTNEXTLINE(misc-const-correctness,misc-use-internal-linkage,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) - false positive
NLOHMANN_JSON_SERIALIZE_ENUM(TaskState,
{
{TS_INVALID, nullptr},
@@ -1708,7 +1708,7 @@ TEST_CASE("JSON to enum mapping")
enum class strict_cards {kreuz, pik, herz, karo, andere}; // andere not included in mapping
// NOLINTNEXTLINE(misc-use-internal-linkage,misc-const-correctness) - false positive
// NOLINTNEXTLINE(misc-use-internal-linkage,misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) - false positive
NLOHMANN_JSON_SERIALIZE_ENUM_STRICT(strict_cards,
{
{strict_cards::kreuz, "kreuz"},
@@ -1727,7 +1727,7 @@ enum StrictTaskState // NOLINT(cert-int09-c,readability-enum-initial-value,cppco
STRICT_TS_INVALID = -1,
};
// NOLINTNEXTLINE(misc-const-correctness,misc-use-internal-linkage) - false positive
// NOLINTNEXTLINE(misc-const-correctness,misc-use-internal-linkage,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) - false positive
NLOHMANN_JSON_SERIALIZE_ENUM_STRICT(StrictTaskState,
{
{STRICT_TS_INVALID, nullptr},
+2 -2
View File
@@ -191,7 +191,7 @@ TEST_CASE("hash of deeply nested values")
// every depth on either side of where the iterative path takes over
for (std::size_t depth = 0; depth <= (2 * nlohmann::detail::recursion_depth_limit()) + 10; ++depth)
{
CAPTURE(depth)
CAPTURE(depth);
const auto arrays = nested<json>(depth, false);
const auto objects = nested<json>(depth, true);
const auto ordered = nested<ordered_json>(depth, true);
@@ -212,7 +212,7 @@ TEST_CASE("hash of deeply nested values")
false, true
})
{
CAPTURE(objects)
CAPTURE(objects);
const auto text = nested_text(depth, objects);
const auto a = json::parse(text);
const auto b = json::parse(text);
+6 -6
View File
@@ -1829,13 +1829,13 @@ TEST_CASE("JSON patch: diff of deeply nested values")
for (const auto depth : depths)
{
CAPTURE(depth)
CAPTURE(depth);
for (int from = 0; from < 3; ++from)
{
for (int to = 0; to < 3; ++to)
{
CAPTURE(from)
CAPTURE(to)
CAPTURE(from);
CAPTURE(to);
const auto source = nested<json>(depth, from);
const auto target = nested<json>(depth, to);
const auto patch = json::diff(source, target);
@@ -1854,7 +1854,7 @@ TEST_CASE("JSON patch: diff of deeply nested values")
{
for (std::size_t depth = 0; depth <= 300; ++depth)
{
CAPTURE(depth)
CAPTURE(depth);
json source = 1;
json target = 2;
for (std::size_t i = 0; i < depth; ++i)
@@ -1878,7 +1878,7 @@ TEST_CASE("JSON patch: diff of deeply nested values")
false, true
})
{
CAPTURE(objects)
CAPTURE(objects);
std::string source_text;
std::string target_text;
std::string equal_text;
@@ -2046,7 +2046,7 @@ TEST_CASE("JSON patch - every operation on ordered_json")
};
for (const auto& target : targets)
{
CAPTURE(target.dump())
CAPTURE(target.dump());
CHECK(source.patch(ordered_json::diff(source, target)) == target);
}
}
+6 -6
View File
@@ -135,7 +135,7 @@ TEST_CASE("tests on deeply nested JSONs")
// are known to meet cleanly - wherever the bound is set.
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d)
CAPTURE(d);
const json array = json::parse(std::string(d, '[') + '0' + std::string(d, ']'));
const json array_copy(array); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
@@ -252,7 +252,7 @@ TEST_CASE("tests on deeply nested JSONs")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern)
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const json j = json::parse(text);
@@ -265,7 +265,7 @@ TEST_CASE("tests on deeply nested JSONs")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern)
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const nlohmann::ordered_json o = nlohmann::ordered_json::parse(text);
@@ -278,7 +278,7 @@ TEST_CASE("tests on deeply nested JSONs")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern)
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const json j = json::parse(text);
@@ -290,10 +290,10 @@ TEST_CASE("tests on deeply nested JSONs")
{
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d)
CAPTURE(d);
for (const auto& pattern : patterns)
{
CAPTURE(pattern)
CAPTURE(pattern);
const std::string text = nested_text(d, pattern);
const json j = json::parse(text);
+3 -3
View File
@@ -289,8 +289,8 @@ TEST_CASE("locale changes between lexer construction and number conversion (#519
for (const auto& transition : transitions)
{
CAPTURE(transition.first)
CAPTURE(transition.second)
CAPTURE(transition.first);
CAPTURE(transition.second);
if (std::setlocale(LC_NUMERIC, transition.first) == nullptr)
{
@@ -368,7 +368,7 @@ TEST_CASE("locale with a multi-byte decimal point")
{
continue;
}
CAPTURE(name)
CAPTURE(name);
tested = true;
// too many significant digits for Clinger's fast path, and an underflow
+3 -3
View File
@@ -305,10 +305,10 @@ TEST_CASE("JSON Merge Patch on deeply nested values")
// over (detail::recursion_depth_limit(), 128)
for (std::size_t depth = 0; depth <= 300; ++depth)
{
CAPTURE(depth)
CAPTURE(depth);
for (int variant = 0; variant < 3; ++variant)
{
CAPTURE(variant)
CAPTURE(variant);
const json patch = json::parse(nested_objects(depth, variant));
json result = json::parse(nested_objects(depth, (variant + 1) % 3));
@@ -403,7 +403,7 @@ TEST_CASE("merge_patch() with an argument that aliases *this (#5641)")
std::size_t{0}, std::size_t{127}, std::size_t{128}, std::size_t{300}
})
{
CAPTURE(depth)
CAPTURE(depth);
json j = json::parse(nested_objects(depth, 0));
const json expected = j;
j.merge_patch(j);
+3 -3
View File
@@ -1084,10 +1084,10 @@ TEST_CASE("update() on deeply nested values")
// over (detail::recursion_depth_limit(), 128)
for (std::size_t depth = 0; depth <= 300; ++depth)
{
CAPTURE(depth)
CAPTURE(depth);
for (int variant = 0; variant < 3; ++variant)
{
CAPTURE(variant)
CAPTURE(variant);
const json source = json::parse(nested_objects(depth, variant));
json result = json::parse(nested_objects(depth, (variant + 1) % 3));
json expected = result;
@@ -1177,7 +1177,7 @@ TEST_CASE("update() with an argument that aliases *this (#5641)")
std::size_t{0}, std::size_t{127}, std::size_t{128}, std::size_t{300}
})
{
CAPTURE(depth)
CAPTURE(depth);
json j = json::parse(nested_objects(depth, 0));
const json expected = j;
j.update(j, true);
+8 -28
View File
@@ -1540,39 +1540,19 @@ TEST_CASE("MessagePack")
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
SECTION("invalid UTF-8 in string (see #5529)")
{
// the MessagePack specification explicitly allows a str object to
// contain a byte sequence that is not valid UTF-8 and expects a
// deserializer to hand the original bytes back unchanged; this
// library follows that, unlike CBOR/UBJSON/BJData/BSON, whose
// specifications require text strings to be valid UTF-8
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
// (0xC0 0xAE is an overlong encoding of '.') round-trips byte for
// byte as a string value
const std::vector<uint8_t> ill_formed_value = {0xa2, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_msgpack(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK(json::from_msgpack(json::to_msgpack(j_value)) == j_value);
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0x81, 0xa2, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_msgpack(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_msgpack(json::to_msgpack(j_key)) == j_key);
// (0xC0 0xAE is an overlong encoding of '.') must be rejected at
// decode time, matching every other kind of malformed binary
// input, rather than only failing later when the resulting
// value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing MessagePack string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae}), true, false).is_discarded());
// a MessagePack bin8 blob with the very same bytes is NOT text
// and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_msgpack(std::vector<uint8_t>({0xc4, 0x02, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
+3 -3
View File
@@ -113,7 +113,7 @@ TEST_CASE("JSON_PRECISE_STREAM_POSITION")
for (const auto& test : tests)
{
CAPTURE(test.first)
CAPTURE(test.first);
std::istringstream ss(test.first);
json j;
ss >> j;
@@ -135,7 +135,7 @@ TEST_CASE("JSON_PRECISE_STREAM_POSITION")
for (const auto& test : tests)
{
CAPTURE(test.first)
CAPTURE(test.first);
std::istringstream ss(test.first);
json j;
ss >> j;
@@ -149,7 +149,7 @@ TEST_CASE("JSON_PRECISE_STREAM_POSITION")
{"1", "12", "-3.5e2", " 7 "
})
{
CAPTURE(s)
CAPTURE(s);
std::istringstream ss(s);
json j;
ss >> j;
+582 -1
View File
@@ -126,7 +126,7 @@ enum class for_1647
two
};
// NOLINTNEXTLINE(misc-const-correctness): this is a false positive
// NOLINTNEXTLINE(misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays): this is a false positive
NLOHMANN_JSON_SERIALIZE_ENUM(for_1647,
{
{for_1647::one, "one"},
@@ -866,4 +866,585 @@ TEST_CASE("regression test - excessive binary container size honors allow_except
CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9b, 0, 0, 0, 0, 0, 0, 0, 0x02}, true, false).is_discarded());
}
namespace
{
/// builds a value from SAX events, asks the parser to recover from its first
/// 100 errors, and checks that the events are balanced (see #3989)
class RecoveringParser
{
public:
explicit RecoveringParser(json& j)
: dom(j, false)
{}
bool null()
{
value();
return dom.null();
}
bool boolean(bool val)
{
value();
return dom.boolean(val);
}
bool number_integer(json::number_integer_t val)
{
value();
return dom.number_integer(val);
}
bool number_unsigned(json::number_unsigned_t val)
{
value();
return dom.number_unsigned(val);
}
bool number_float(json::number_float_t val, const std::string& s)
{
value();
return dom.number_float(val, s);
}
bool string(std::string& val)
{
value();
return dom.string(val);
}
bool binary(json::binary_t& val)
{
value();
return dom.binary(val);
}
bool start_object(std::size_t elements)
{
value();
stack.push_back('o');
return dom.start_object(elements);
}
bool key(std::string& val)
{
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.back() = 'v';
return dom.key(val);
}
bool end_object()
{
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_object();
}
bool start_array(std::size_t elements)
{
value();
stack.push_back('a');
return dom.start_array(elements);
}
bool end_array()
{
if (stack.empty() || stack.back() != 'a')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
messages.emplace_back(ex.what());
// a limit, so that a reader that does not stop fails the test
// instead of making it hang
return ++errors < 100;
}
/// whether the events were balanced and every key was followed by a value
bool balanced() const
{
return well_formed && stack.empty();
}
/// builds the value
nlohmann::detail::json_sax_dom_parser<json> dom;
std::size_t errors = 0;
std::vector<std::string> messages {}; // NOLINT(readability-redundant-member-init)
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
private:
void value()
{
if (!stack.empty())
{
if (stack.back() == 'v')
{
stack.back() = 'o';
}
else if (stack.back() == 'o')
{
well_formed = false;
}
}
}
};
struct BinaryParseResult
{
json value;
std::size_t errors;
std::vector<std::string> messages;
bool ok;
bool balanced;
};
BinaryParseResult parse_binary_recovering(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
json j;
RecoveringParser sax(j);
const bool ok = json::sax_parse(input, &sax, format);
return {j, sax.errors, sax.messages, ok, sax.balanced()};
}
#if !defined(JSON_NOEXCEPTION)
/// the message of the exception that reading @a input into a JSON value
/// throws, or an empty string if reading succeeds
std::string binary_error_message(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
try
{
json _;
switch (format)
{
case json::input_format_t::cbor:
_ = json::from_cbor(input);
break;
case json::input_format_t::msgpack:
_ = json::from_msgpack(input);
break;
case json::input_format_t::ubjson:
_ = json::from_ubjson(input);
break;
case json::input_format_t::bjdata:
_ = json::from_bjdata(input);
break;
case json::input_format_t::bson:
_ = json::from_bson(input);
break;
case json::input_format_t::bon8:
_ = json::from_bon8(input);
break;
case json::input_format_t::json:
default:
break;
}
}
catch (const json::exception& e)
{
return e.what();
}
return "";
}
#endif
/// a BSON element: its type, its name, and its value
std::vector<std::uint8_t> bson_element(const std::uint8_t type, const std::string& name, const std::vector<std::uint8_t>& value)
{
std::vector<std::uint8_t> result = {type};
result.insert(result.end(), name.begin(), name.end());
result.push_back(0x00);
result.insert(result.end(), value.begin(), value.end());
return result;
}
/// a BSON document of the given elements; @a size_offset is added to the
/// size it declares
std::vector<std::uint8_t> bson_document(const std::vector<std::vector<std::uint8_t>>& elements, const int size_offset = 0)
{
std::vector<std::uint8_t> body;
for (const auto& element : elements)
{
body.insert(body.end(), element.begin(), element.end());
}
const auto size = static_cast<std::uint32_t>(static_cast<int>(body.size()) + 5 + size_offset);
std::vector<std::uint8_t> result = {static_cast<std::uint8_t>(size & 0xFFu), static_cast<std::uint8_t>((size >> 8u) & 0xFFu),
static_cast<std::uint8_t>((size >> 16u) & 0xFFu), static_cast<std::uint8_t>((size >> 24u) & 0xFFu)
};
result.insert(result.end(), body.begin(), body.end());
result.push_back(0x00);
return result;
}
/// a BSON int32 value
std::vector<std::uint8_t> bson_int32(const std::int32_t value)
{
const auto u = static_cast<std::uint32_t>(value);
return {static_cast<std::uint8_t>(u & 0xFFu), static_cast<std::uint8_t>((u >> 8u) & 0xFFu),
static_cast<std::uint8_t>((u >> 16u) & 0xFFu), static_cast<std::uint8_t>((u >> 24u) & 0xFFu)};
}
/// a BSON string value, whose length is @a length_offset off
std::vector<std::uint8_t> bson_string(const std::string& value, const std::int32_t length_offset = 0)
{
auto result = bson_int32(static_cast<std::int32_t>(value.size() + 1) + length_offset);
result.insert(result.end(), value.begin(), value.end());
result.push_back(0x00);
return result;
}
/// @a count bytes of value 0xAB
std::vector<std::uint8_t> bytes(const std::size_t count)
{
return std::vector<std::uint8_t>(count, 0xAB);
}
template<typename... Parts>
std::vector<std::uint8_t> concatenated(const std::vector<std::uint8_t>& first, const Parts& ... rest)
{
std::vector<std::uint8_t> result = first;
for (const auto& part : std::initializer_list<std::vector<std::uint8_t>> {rest...})
{
result.insert(result.end(), part.begin(), part.end());
}
return result;
}
/// U+FFFD REPLACEMENT CHARACTER
std::string replacement_character()
{
return "\xEF\xBF\xBD";
}
} // namespace
TEST_CASE("regression test - #3989 SAX parse_error() returning true")
{
SECTION("binary formats complete what was read before the input ends")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& bytes = encoding.second;
CAPTURE(format);
// every prefix is truncated input
for (std::size_t length = 0; length < bytes.size(); ++length)
{
CAPTURE(length);
const auto result = parse_binary_recovering(std::vector<std::uint8_t>(bytes.begin(), bytes.begin() + static_cast<std::ptrdiff_t>(length)), format);
CHECK(!result.ok);
CHECK(result.errors == 1);
CHECK(result.balanced);
}
// the complete input is read as usual (binary values do not
// round-trip through every format, so compare with a plain parse)
json expected;
nlohmann::detail::json_sax_dom_parser<json> dom(expected);
CHECK(json::sax_parse(bytes, &dom, format));
const auto complete = parse_binary_recovering(bytes, format);
CHECK(complete.ok);
CHECK(complete.errors == 0);
CHECK(complete.value == expected);
// a byte after the value
auto trailing_bytes = bytes;
trailing_bytes.push_back(0x01);
const auto trailing = parse_binary_recovering(trailing_bytes, format);
CHECK(!trailing.ok);
CHECK(trailing.errors == 1);
CHECK(trailing.value == expected);
}
}
SECTION("containers without an end")
{
// these made the readers loop, or read on, after the error
const auto cbor_array = parse_binary_recovering({0x9F}, json::input_format_t::cbor);
CHECK(cbor_array.errors == 1);
CHECK(cbor_array.value == json::array());
const auto cbor_map = parse_binary_recovering({0xBF, 0x61, 'a'}, json::input_format_t::cbor);
CHECK(cbor_map.errors == 1);
CHECK(cbor_map.value == json({{"a", nullptr}}));
const auto msgpack_array = parse_binary_recovering({0xDD, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::msgpack);
CHECK(msgpack_array.errors == 1);
CHECK(msgpack_array.value == json::array());
const auto msgpack_map = parse_binary_recovering({0x81, 0xA1, 'a', 0x92, 0x01}, json::input_format_t::msgpack);
CHECK(msgpack_map.errors == 1);
CHECK(msgpack_map.value == json({{"a", {1}}}));
}
SECTION("BJData ndarray")
{
// a 2x3 int8 array with two of its six elements; the annotated array
// format opens an object and two arrays of its own
const auto result = parse_binary_recovering({'[', '$', 'i', '#', '[', '$', 'i', '#', 'i', 2, 2, 3, 1, 2}, json::input_format_t::bjdata);
CHECK(result.errors == 1);
CHECK(result.balanced);
CHECK(result.value == json({{"_ArrayType_", "int8"}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2}}}));
}
SECTION("binary formats repair items whose end is known")
{
struct Repair
{
json::input_format_t format;
std::vector<std::uint8_t> input;
json expected;
std::size_t errors;
};
const std::vector<Repair> repairs =
{
// CBOR: tags are ignored (here tag 1 and the self-describe tag 55799)
{json::input_format_t::cbor, {0x82, 0xC1, 0x05, 0xD9, 0xD9, 0xF7, 0x06}, {5, 6}, 2},
// CBOR: undefined and other simple values become null
{json::input_format_t::cbor, {0x84, 0xF7, 0xE0, 0xF8, 0x20, 0x01}, {nullptr, nullptr, nullptr, 1}, 3},
// CBOR: ill-formed UTF-8 becomes U+FFFD, also in keys
{json::input_format_t::cbor, {0xA1, 0x61, 0xFF, 0x62, 0xC3, 0x28}, {{replacement_character(), replacement_character() + "("}}, 2},
// CBOR: members whose key is not a string are skipped, whatever their key and value
{json::input_format_t::cbor, {0xA4, 0x01, 0x02, 0x82, 0x01, 0x02, 0xA1, 0x61, 'x', 0x9F, 0xFF, 0xC1, 0x01, 0x5F, 0x41, 0x00, 0xFF, 0x61, 'a', 0x03}, {{"a", 3}}, 3},
{json::input_format_t::cbor, {0xBF, 0xF5, 0xBF, 0x61, 'x', 0x7F, 0x61, 'y', 0xFF, 0xFF, 0x61, 'a', 0x03, 0xFF}, {{"a", 3}}, 1},
// MessagePack: members whose key is not a string are skipped
{json::input_format_t::msgpack, {0x84, 0x01, 0x02, 0x81, 0xA1, 'x', 0x01, 0x92, 0x01, 0x02, 0xD4, 0x01, 0x02, 0xC0, 0xA1, 'a', 0x04}, {{"a", 4}}, 3},
// MessagePack: ill-formed UTF-8 becomes U+FFFD
{json::input_format_t::msgpack, {0x92, 0xA2, 0xC3, 0x28, 0xA3, 0xE2, 0x82, 'x'}, {replacement_character() + "(", replacement_character() + "x"}, 2},
// UBJSON: a char that is not ASCII becomes U+FFFD
{json::input_format_t::ubjson, {'[', 'C', 0x80, 'C', 'A', ']'}, {replacement_character(), "A"}, 1},
// UBJSON: the longest beginning of a high-precision number is kept
{json::input_format_t::ubjson, {'[', 'H', 'i', 5, '1', '2', 'a', 'b', 'c', 'H', 'i', 2, '1', '.', 'H', 'i', 3, 'a', 'b', 'c', 'H', 'i', 3, '4', '.', '5', ']'}, {12, 1, nullptr, 4.5}, 3},
// BJData, too
{json::input_format_t::bjdata, {'[', 'C', 0xFF, 'H', 'i', 2, '-', '1', 'H', 'i', 2, '-', 'x', ']'}, {replacement_character(), -1, nullptr}, 2},
// BON8: members whose key is not a string are skipped
{json::input_format_t::bon8, {0x89, 0x91, 0x92, 0xC9, 0x40, 0x82, 0x91, 0x92, 0x61, 0x93}, {{"a", 3}}, 2},
{json::input_format_t::bon8, {0x8B, 0x91, 0x85, 0x91, 0xFE, 0xFA, 0x8B, 'x', 0x91, 0xFE, 0x61, 0x93, 0xFE}, {{"a", 3}}, 2},
// BSON: elements of types the library does not read become null
{
json::input_format_t::bson, bson_document(
{
bson_element(0x07, "_id", bytes(12)), // ObjectId
bson_element(0x09, "date", bytes(8)), // UTC datetime
bson_element(0x13, "decimal", bytes(16)), // 128-bit decimal
bson_element(0x0B, "regex", {'a', '+', 0, 'i', 0}), // regular expression
bson_element(0x0D, "code", bson_string("f()")), // JavaScript code
bson_element(0x0E, "symbol", bson_string("s")), // symbol
bson_element(0x0C, "pointer", concatenated(bson_string("c"), bytes(12))), // DBPointer
bson_element(0x0F, "scope", concatenated(bson_int32(15), bson_string("g"), bson_document({}))), // code with scope
bson_element(0x06, "undefined", {}), // undefined
bson_element(0xFF, "min", {}), // min key
bson_element(0x7F, "max", {}), // max key
bson_element(0x10, "z", bson_int32(7)),
}),
{{"_id", nullptr}, {"date", nullptr}, {"decimal", nullptr}, {"regex", nullptr}, {"code", nullptr}, {"symbol", nullptr}, {"pointer", nullptr}, {"scope", nullptr}, {"undefined", nullptr}, {"min", nullptr}, {"max", nullptr}, {"z", 7}},
11
},
// BSON: an element of an unknown type becomes null, and the rest of its document is skipped
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x04, "array", bson_document({bson_element(0x10, "0", bson_int32(1)), bson_element(0x42, "1", bytes(3))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"a", 1}, {"x", nullptr}}}, {"array", {1, nullptr}}, {"after", 3}},
2
},
// BSON: so does a string or byte array whose length cannot be right
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x02, "s", bson_string("abc", -10)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x03, "bin", bson_document({bson_element(0x05, "b", concatenated(bson_int32(-1), bytes(1))), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"s", nullptr}}}, {"bin", {{"b", nullptr}}}, {"after", 3}},
2
},
// BSON: a string without its terminator, and a document whose size does not match, are kept
{
json::input_format_t::bson, bson_document(
{
bson_element(0x02, "s", {2, 0, 0, 0, 'a', 'X'}),
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1))}, 1)),
}),
{{"s", "a"}, {"inner", {{"a", 1}}}},
2
},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.format);
CAPTURE(repair.input);
const auto result = parse_binary_recovering(repair.input, repair.format);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors == repair.errors);
CHECK(result.value == repair.expected);
REQUIRE(!result.messages.empty());
#if !defined(JSON_NOEXCEPTION)
// the first error is the one reported without recovering; under
// JSON_NOEXCEPTION, reading without recovering aborts instead of
// throwing, so there is no message to compare with
CHECK(result.messages.front() == binary_error_message(repair.input, repair.format));
#endif
}
}
SECTION("binary formats repair numbers that are out of range")
{
// CBOR: a negative integer below the range of number_integer_t
const auto cbor = parse_binary_recovering({0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value.is_number_float());
CHECK(cbor.value.get<double>() == -18446744073709551616.0);
// UBJSON: a high-precision number too large for number_float_t
const auto ubjson = parse_binary_recovering({'H', 'i', 5, '1', 'e', '9', '9', '9'}, json::input_format_t::ubjson);
CHECK(ubjson.errors == 1);
CHECK(ubjson.value.is_number_float());
CHECK(std::isinf(ubjson.value.get<double>()));
}
SECTION("binary formats stop where the end of an item is not known")
{
// a byte that begins no item
const auto cbor = parse_binary_recovering({0x82, 0x01, 0x1C, 0x02}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value == json({1}));
// a key that is no item: the unused MessagePack byte, a CBOR break
// in a map of known size, and the end of a BON8 container
const auto msgpack = parse_binary_recovering({0x82, 0xA1, 'a', 0x01, 0xC1, 0x02}, json::input_format_t::msgpack);
CHECK(msgpack.errors == 1);
CHECK(msgpack.value == json({{"a", 1}}));
const auto cbor_break = parse_binary_recovering({0xA2, 0x61, 'a', 0x01, 0xFF, 0x02}, json::input_format_t::cbor);
CHECK(cbor_break.errors == 1);
CHECK(cbor_break.value == json({{"a", 1}}));
const auto bon8 = parse_binary_recovering({0x88, 0x61, 0x91, 0xFE}, json::input_format_t::bon8);
CHECK(bon8.errors == 1);
CHECK(bon8.value == json({{"a", 1}}));
// a skipped member that the input ends in
const auto truncated = parse_binary_recovering({0xA2, 0x01, 0x82, 0x01}, json::input_format_t::cbor);
CHECK(truncated.errors == 2);
CHECK(truncated.balanced);
CHECK(truncated.value == json::object());
// a BSON element of an unknown type in a document whose size cannot be right
const auto bson = parse_binary_recovering(bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3))}, -10), json::input_format_t::bson);
CHECK(bson.errors == 1);
CHECK(bson.value == json({{"a", 1}, {"x", nullptr}}));
}
SECTION("changed bytes in binary input")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}, {"i", "\xC3\xA4"}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
const std::vector<std::uint8_t> replacements = {0x00, 0x01, 0x7F, 0x80, 0xC1, 0xD9, 0xE0, 0xF7, 0xFE, 0xFF};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& original = encoding.second;
CAPTURE(format);
std::vector<std::vector<std::uint8_t>> inputs;
for (std::size_t position = 0; position < original.size(); ++position)
{
for (const auto replacement : replacements)
{
auto changed = original;
changed[position] = replacement;
inputs.push_back(changed);
}
auto removed = original;
removed.erase(removed.begin() + static_cast<std::ptrdiff_t>(position));
inputs.push_back(removed);
}
for (const auto& input : inputs)
{
CAPTURE(input);
const auto result = parse_binary_recovering(input, format);
CHECK(result.balanced);
CHECK(result.errors <= input.size() + 1);
#if !defined(JSON_NOEXCEPTION)
// an error is reported exactly if reading into a JSON value
// fails, and the first one is the same (under JSON_NOEXCEPTION,
// that reading aborts instead of throwing)
const auto message = binary_error_message(input, format);
CHECK(result.ok == message.empty());
if (!result.ok && result.errors < 100)
{
CHECK(result.messages.front() == message);
}
#endif
}
}
}
SECTION("JSON text")
{
// the parser stopped, but reported success
json j;
RecoveringParser sax(j);
CHECK(!json::sax_parse("[1,2,3,]", &sax));
CHECK(sax.errors == 1);
CHECK(j == json({1, 2, 3}));
}
SECTION("the SAX parsers of the library stop")
{
json _;
CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9F}, true, false).is_discarded());
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<std::uint8_t> {0x9F}), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR value: unexpected end of input", json::parse_error&);
CHECK(json::parse("[1,2,3,]", nullptr, false).is_discarded());
CHECK(!json::accept("[1,2,3,]"));
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+2 -2
View File
@@ -849,13 +849,13 @@ TEST_CASE("issue #5338 - truncated CBOR tagged binary subtype is rejected")
for (const auto& data : truncated_tags)
{
CAPTURE(data)
CAPTURE(data);
for (const auto tag_handler :
{
json::cbor_tag_handler_t::ignore, json::cbor_tag_handler_t::store
})
{
CAPTURE(tag_handler)
CAPTURE(tag_handler);
const auto result = json::from_cbor(data, true, false, tag_handler);
CHECK(result.is_discarded());
}
+6 -6
View File
@@ -584,7 +584,7 @@ TEST_CASE("serialization of deeply nested values")
// value are known to meet cleanly - wherever the bound is set.
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d)
CAPTURE(d);
const std::string array_text = std::string(d, '[') + '7' + std::string(d, ']');
CHECK(json::parse(array_text).dump() == array_text);
@@ -604,7 +604,7 @@ TEST_CASE("serialization of deeply nested values")
{
for (std::size_t d = 120; d <= 140; ++d)
{
CAPTURE(d)
CAPTURE(d);
const json j = json::parse(std::string(d, '[') + '7' + std::string(d, ']'));
@@ -629,7 +629,7 @@ TEST_CASE("serialization of deeply nested values")
// so it must not gain a newline when it is reached iteratively
for (std::size_t d = 125; d <= 135; ++d)
{
CAPTURE(d)
CAPTURE(d);
const std::string compact = std::string(d, '[') + "[]" + std::string(d, ']');
CHECK(json::parse(compact).dump() == compact);
@@ -711,10 +711,10 @@ TEST_CASE("serialization of every kind of value below the bound of the descent")
for (const std::size_t depth : std::vector<std::size_t> {1, 200})
{
CAPTURE(depth)
CAPTURE(depth);
for (const auto& inner : values)
{
CAPTURE(inner.dump())
CAPTURE(inner.dump());
const json j = wrap_in_arrays(inner, depth);
CHECK(j.dump() == std::string(depth, '[') + inner.dump() + std::string(depth, ']'));
CHECK(j.dump(2) == expected_pretty_in_arrays(inner, depth));
@@ -725,7 +725,7 @@ TEST_CASE("serialization of every kind of value below the bound of the descent")
{
for (std::size_t d = 120; d <= 140; ++d)
{
CAPTURE(d)
CAPTURE(d);
// built from the inside out: {"k": <level below>, "n": <level>}
json j = 7;
-35
View File
@@ -2505,41 +2505,6 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_ubjson(j) == v);
CHECK(json::from_ubjson(v) == j);
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_ubjson() is
// strict, so such a value cannot be written back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_ubjson(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(j), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_ubjson(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_ubjson(j_key), json::type_error&);
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_ubjson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")
+3 -3
View File
@@ -350,7 +350,7 @@ TEST_CASE("std::counted_iterator reaches the contiguous fast paths")
for (const auto& text : diagnostic_docs)
{
CAPTURE(text)
CAPTURE(text);
const std::counted_iterator<const char*> it(text.data(), static_cast<std::iter_difference_t<const char*>>(text.size()));
std::string counted_message;
std::string string_message;
@@ -460,8 +460,8 @@ TEST_CASE("std::counted_iterator bulk scanning stops at the counted end")
for (const auto& tc : cases)
{
CAPTURE(tc.buffer)
CAPTURE(tc.count)
CAPTURE(tc.buffer);
CAPTURE(tc.count);
const std::string buffer = tc.buffer;
CHECK(via_counted(buffer, tc.count) == via_prefix(buffer, tc.count));
}
+4 -4
View File
@@ -63,10 +63,10 @@ TEST_CASE_TEMPLATE_DEFINE("value_in_range_of trait", T, value_in_range_of_test)
INFO("type := ", type_str);
CAPTURE(val_min)
CAPTURE(min_in_range)
CAPTURE(val_max)
CAPTURE(max_in_range)
CAPTURE(val_min);
CAPTURE(min_in_range);
CAPTURE(val_max);
CAPTURE(max_in_range);
if (min_in_range)
{