Compare commits

..
Author SHA1 Message Date
Niels Lohmann 5ab47ef138 Merge branch 'develop' into claude/compile-time-docs
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:22:12 +02:00
Niels Lohmann 633de8e44b Fix CI: clang-tidy and GCC -Wnoexcept in the locale test (#5613)
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:

- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
  calls whose result was discarded. Check the result, like the other
  resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
  callbacks, which cannot throw but were not declared noexcept.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:20:43 +02:00
Niels Lohmann afba254d39 Document options to reduce compile times
Add an integration page that collects the ways to reduce compile times
with measurements: json_fwd.hpp in headers, JSON_NO_AUTOMATIC_UDLS,
explicit instantiation with extern template, modules, and precompiled
headers, and notes that JSON_NO_IO and JSON_USE_GLOBAL_UDLS have no
measurable effect.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:02:18 +02:00
Niels Lohmann 1419aeff60 Move the user-defined string literals to <nlohmann/json_literals.hpp>
Following the review in #5294, the literals now live in their own header
instead of being removed entirely: <nlohmann/json.hpp> includes it at the
end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined,
so a project can opt out globally and include the header only where the
literals are used.

The header only uses public and standard macros, because the library's
internal macros are undefined at the end of json.hpp and the amalgamation
inlines macro_scope.hpp only once. For the same reason, the library no
longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition
is still visible to the header. The single-header copy is identical to
the multi-header one, as it only includes <nlohmann/json.hpp>. The module
always exports the literals.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 21:49:37 +02:00
Niels Lohmann e0ee07fe73 Mention JSON_NO_UDLS in the list of exported module symbols
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 21:24:40 +02:00
Niels Lohmann dc4a70e7b4 Add JSON_NO_UDLS to leave out the user-defined string literals
The bodies of operator""_json and operator""_json_pointer call the
parser, so every translation unit including the library instantiates it,
even if it never parses anything. Defining JSON_NO_UDLS leaves the
literals out entirely, which saves 15-35% compile time for such
translation units (#5294). Nothing changes if the macro is not defined.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 21:23:52 +02:00
Niels Lohmann fc03b9912e Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction

The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.

token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.

Fixes #5198

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the strtod retry loop when the decimal point is unchanged

convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.

Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Weffc++ errors in the #5198 locale test

GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:56:11 +02:00
Niels Lohmann 9e1a09eec0 Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys

CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:

  syntax error while parsing MessagePack object key: only string keys
  are supported, but found nil; last byte: 0xC0

The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.

Refs #2766, #3381

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point the MessagePack key note to the spec's profile section

The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:51:07 +02:00
50 changed files with 1520 additions and 547 deletions
+1
View File
@@ -67,6 +67,7 @@ jobs:
python3 $TOOL_DIR/amalgamate.py -c $TOOL_DIR/config_json.json -s . python3 $TOOL_DIR/amalgamate.py -c $TOOL_DIR/config_json.json -s .
python3 $TOOL_DIR/amalgamate.py -c $TOOL_DIR/config_json_fwd.json -s . python3 $TOOL_DIR/amalgamate.py -c $TOOL_DIR/config_json_fwd.json -s .
cp include/nlohmann/json_literals.hpp $INCLUDE_DIR/json_literals.hpp
# the header list of the Bazel "json" target must match the files in include/ # the header list of the Bazel "json" target must match the files in include/
cmake -P cmake/scripts/gen_bazel_build_file.cmake cmake -P cmake/scripts/gen_bazel_build_file.cmake
+1
View File
@@ -65,6 +65,7 @@ cc_library(
"include/nlohmann/detail/value_t.hpp", "include/nlohmann/detail/value_t.hpp",
"include/nlohmann/json.hpp", "include/nlohmann/json.hpp",
"include/nlohmann/json_fwd.hpp", "include/nlohmann/json_fwd.hpp",
"include/nlohmann/json_literals.hpp",
"include/nlohmann/ordered_map.hpp", "include/nlohmann/ordered_map.hpp",
"include/nlohmann/thirdparty/hedley/hedley.hpp", "include/nlohmann/thirdparty/hedley/hedley.hpp",
"include/nlohmann/thirdparty/hedley/hedley_undef.hpp", "include/nlohmann/thirdparty/hedley/hedley_undef.hpp",
+16 -5
View File
@@ -21,6 +21,8 @@ TESTS_SRCS=$(shell find tests -type f \( -name '*.hpp' -o -name '*.cpp' -o -name
# the single headers (amalgamated from the source files) # the single headers (amalgamated from the source files)
AMALGAMATED_FILE=single_include/nlohmann/json.hpp AMALGAMATED_FILE=single_include/nlohmann/json.hpp
AMALGAMATED_FWD_FILE=single_include/nlohmann/json_fwd.hpp AMALGAMATED_FWD_FILE=single_include/nlohmann/json_fwd.hpp
# json_literals.hpp only includes <nlohmann/json.hpp>, so it is copied verbatim
AMALGAMATED_LITERALS_FILE=single_include/nlohmann/json_literals.hpp
########################################################################## ##########################################################################
@@ -29,7 +31,7 @@ AMALGAMATED_FWD_FILE=single_include/nlohmann/json_fwd.hpp
# main target # main target
all: all:
@echo "amalgamate - amalgamate files single_include/nlohmann/json{,_fwd}.hpp from the include/nlohmann sources" @echo "amalgamate - amalgamate files single_include/nlohmann/json{,_fwd,_literals}.hpp from the include/nlohmann sources"
@echo "BUILD.bazel - regenerate the Bazel BUILD file from the include/nlohmann sources" @echo "BUILD.bazel - regenerate the Bazel BUILD file from the include/nlohmann sources"
@echo "ChangeLog.md - generate ChangeLog file" @echo "ChangeLog.md - generate ChangeLog file"
@echo "check-amalgamation - check whether sources have been amalgamated and BUILD.bazel is up to date" @echo "check-amalgamation - check whether sources have been amalgamated and BUILD.bazel is up to date"
@@ -154,14 +156,14 @@ install_astyle:
# call the Artistic Style pretty printer on all source files # call the Artistic Style pretty printer on all source files
pretty: install_astyle pretty: install_astyle
$(ASTYLE) --project=tools/astyle/.astylerc $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) docs/mkdocs/docs/examples/*.cpp $(ASTYLE) --project=tools/astyle/.astylerc $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE) docs/mkdocs/docs/examples/*.cpp
# call the Clang-Format on all source files # call the Clang-Format on all source files
pretty_format: pretty_format:
for FILE in $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) docs/mkdocs/docs/examples/*.cpp; do echo $$FILE; clang-format -i $$FILE; done for FILE in $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) docs/mkdocs/docs/examples/*.cpp; do echo $$FILE; clang-format -i $$FILE; done
# create single header files and pretty print # create single header files and pretty print
amalgamate: $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) amalgamate: $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE)
$(MAKE) pretty $(MAKE) pretty
# call the amalgamation tool for json.hpp # call the amalgamation tool for json.hpp
@@ -172,16 +174,23 @@ $(AMALGAMATED_FILE): $(SRCS)
$(AMALGAMATED_FWD_FILE): $(SRCS) $(AMALGAMATED_FWD_FILE): $(SRCS)
tools/amalgamate/amalgamate.py -c tools/amalgamate/config_json_fwd.json -s . --verbose=yes tools/amalgamate/amalgamate.py -c tools/amalgamate/config_json_fwd.json -s . --verbose=yes
# copy json_literals.hpp
$(AMALGAMATED_LITERALS_FILE): include/nlohmann/json_literals.hpp
cp include/nlohmann/json_literals.hpp $(AMALGAMATED_LITERALS_FILE)
# check if file single_include/nlohmann/json.hpp has been amalgamated from the nlohmann sources # check if file single_include/nlohmann/json.hpp has been amalgamated from the nlohmann sources
# Note: this target is called by Travis # Note: this target is called by Travis
check-amalgamation: check-amalgamation:
@mv $(AMALGAMATED_FILE) $(AMALGAMATED_FILE)~ @mv $(AMALGAMATED_FILE) $(AMALGAMATED_FILE)~
@mv $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_FWD_FILE)~ @mv $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_FWD_FILE)~
@mv $(AMALGAMATED_LITERALS_FILE) $(AMALGAMATED_LITERALS_FILE)~
@$(MAKE) amalgamate @$(MAKE) amalgamate
@diff $(AMALGAMATED_FILE) $(AMALGAMATED_FILE)~ || (echo "===================================================================\n Amalgamation required! Please read the contribution guidelines\n in file .github/CONTRIBUTING.md.\n===================================================================" ; mv $(AMALGAMATED_FILE)~ $(AMALGAMATED_FILE) ; false) @diff $(AMALGAMATED_FILE) $(AMALGAMATED_FILE)~ || (echo "===================================================================\n Amalgamation required! Please read the contribution guidelines\n in file .github/CONTRIBUTING.md.\n===================================================================" ; mv $(AMALGAMATED_FILE)~ $(AMALGAMATED_FILE) ; false)
@diff $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_FWD_FILE)~ || (echo "===================================================================\n Amalgamation required! Please read the contribution guidelines\n in file .github/CONTRIBUTING.md.\n===================================================================" ; mv $(AMALGAMATED_FWD_FILE)~ $(AMALGAMATED_FWD_FILE) ; false) @diff $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_FWD_FILE)~ || (echo "===================================================================\n Amalgamation required! Please read the contribution guidelines\n in file .github/CONTRIBUTING.md.\n===================================================================" ; mv $(AMALGAMATED_FWD_FILE)~ $(AMALGAMATED_FWD_FILE) ; false)
@diff $(AMALGAMATED_LITERALS_FILE) $(AMALGAMATED_LITERALS_FILE)~ || (echo "===================================================================\n Amalgamation required! Please read the contribution guidelines\n in file .github/CONTRIBUTING.md.\n===================================================================" ; mv $(AMALGAMATED_LITERALS_FILE)~ $(AMALGAMATED_LITERALS_FILE) ; false)
@mv $(AMALGAMATED_FILE)~ $(AMALGAMATED_FILE) @mv $(AMALGAMATED_FILE)~ $(AMALGAMATED_FILE)
@mv $(AMALGAMATED_FWD_FILE)~ $(AMALGAMATED_FWD_FILE) @mv $(AMALGAMATED_FWD_FILE)~ $(AMALGAMATED_FWD_FILE)
@mv $(AMALGAMATED_LITERALS_FILE)~ $(AMALGAMATED_LITERALS_FILE)
@mv BUILD.bazel BUILD.bazel~ @mv BUILD.bazel BUILD.bazel~
@$(MAKE) BUILD.bazel @$(MAKE) BUILD.bazel
@diff BUILD.bazel BUILD.bazel~ || (echo "===================================================================\n BUILD.bazel is out of date! Please run 'make BUILD.bazel'.\n===================================================================" ; mv BUILD.bazel~ BUILD.bazel ; false) @diff BUILD.bazel BUILD.bazel~ || (echo "===================================================================\n BUILD.bazel is out of date! Please run 'make BUILD.bazel'.\n===================================================================" ; mv BUILD.bazel~ BUILD.bazel ; false)
@@ -222,7 +231,7 @@ json.tar.xz:
# We use `-X` to make the resulting ZIP file reproducible, see # We use `-X` to make the resulting ZIP file reproducible, see
# <https://content.pivotal.io/blog/barriers-to-deterministic-reproducible-zip-files>. # <https://content.pivotal.io/blog/barriers-to-deterministic-reproducible-zip-files>.
include.zip: BUILD.bazel include.zip: BUILD.bazel
zip -9 --recurse-paths -X include.zip $(SRCS) $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) BUILD.bazel MODULE.bazel meson.build LICENSE.MIT zip -9 --recurse-paths -X include.zip $(SRCS) $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE) BUILD.bazel MODULE.bazel meson.build LICENSE.MIT
# Create the files for a release and add signatures and hashes. # Create the files for a release and add signatures and hashes.
release: include.zip json.tar.xz release: include.zip json.tar.xz
@@ -231,10 +240,12 @@ release: include.zip json.tar.xz
gpg --armor --detach-sig include.zip gpg --armor --detach-sig include.zip
gpg --armor --detach-sig $(AMALGAMATED_FILE) gpg --armor --detach-sig $(AMALGAMATED_FILE)
gpg --armor --detach-sig $(AMALGAMATED_FWD_FILE) gpg --armor --detach-sig $(AMALGAMATED_FWD_FILE)
gpg --armor --detach-sig $(AMALGAMATED_LITERALS_FILE)
gpg --armor --detach-sig json.tar.xz gpg --armor --detach-sig json.tar.xz
cp $(AMALGAMATED_FILE) release_files cp $(AMALGAMATED_FILE) release_files
cp $(AMALGAMATED_FWD_FILE) release_files cp $(AMALGAMATED_FWD_FILE) release_files
mv $(AMALGAMATED_FILE).asc $(AMALGAMATED_FWD_FILE).asc json.tar.xz json.tar.xz.asc include.zip include.zip.asc release_files cp $(AMALGAMATED_LITERALS_FILE) release_files
mv $(AMALGAMATED_FILE).asc $(AMALGAMATED_FWD_FILE).asc $(AMALGAMATED_LITERALS_FILE).asc json.tar.xz json.tar.xz.asc include.zip include.zip.asc release_files
cd release_files ; shasum -a 256 json.hpp include.zip json.tar.xz > hashes.txt cd release_files ; shasum -a 256 json.hpp include.zip json.tar.xz > hashes.txt
+2 -2
View File
@@ -363,7 +363,7 @@ std::cout << j_string << " == " << serialized_string << std::endl;
[`.dump()`](https://json.nlohmann.me/api/basic_json/dump/) returns the originally stored string value. [`.dump()`](https://json.nlohmann.me/api/basic_json/dump/) returns the originally stored string value.
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers. Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
#### To/from streams (e.g., files, string streams) #### To/from streams (e.g., files, string streams)
@@ -1915,7 +1915,7 @@ The library supports **Unicode input** as follows:
- [Unicode noncharacters](https://www.unicode.org/faq/private_use.html#nonchar1) will not be replaced by the library. - [Unicode noncharacters](https://www.unicode.org/faq/private_use.html#nonchar1) will not be replaced by the library.
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors. - Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs. - The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers. - When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
- To store wide strings (e.g., `std::wstring`), you need to convert them to a UTF-8 encoded `std::string` before, see [an example](https://json.nlohmann.me/home/faq/#wide-string-handling). - To store wide strings (e.g., `std::wstring`), you need to convert them to a UTF-8 encoded `std::string` before, see [an example](https://json.nlohmann.me/home/faq/#wide-string-handling).
### Comments in JSON ### Comments in JSON
+4 -1
View File
@@ -373,9 +373,10 @@ file(GLOB_RECURSE INDENT_FILES
set(include_dir ${PROJECT_SOURCE_DIR}/single_include/nlohmann) set(include_dir ${PROJECT_SOURCE_DIR}/single_include/nlohmann)
set(tool_dir ${PROJECT_SOURCE_DIR}/tools/amalgamate) set(tool_dir ${PROJECT_SOURCE_DIR}/tools/amalgamate)
add_custom_target(ci_test_amalgamation add_custom_target(ci_test_amalgamation
COMMAND rm -fr ${include_dir}/json.hpp~ ${include_dir}/json_fwd.hpp~ COMMAND rm -fr ${include_dir}/json.hpp~ ${include_dir}/json_fwd.hpp~ ${include_dir}/json_literals.hpp~
COMMAND cp ${include_dir}/json.hpp ${include_dir}/json.hpp~ COMMAND cp ${include_dir}/json.hpp ${include_dir}/json.hpp~
COMMAND cp ${include_dir}/json_fwd.hpp ${include_dir}/json_fwd.hpp~ COMMAND cp ${include_dir}/json_fwd.hpp ${include_dir}/json_fwd.hpp~
COMMAND cp ${include_dir}/json_literals.hpp ${include_dir}/json_literals.hpp~
COMMAND ${Python3_EXECUTABLE} -mvenv venv_astyle COMMAND ${Python3_EXECUTABLE} -mvenv venv_astyle
COMMAND venv_astyle/bin/pip3 --quiet install -r ${CMAKE_SOURCE_DIR}/tools/astyle/requirements.txt COMMAND venv_astyle/bin/pip3 --quiet install -r ${CMAKE_SOURCE_DIR}/tools/astyle/requirements.txt
@@ -383,10 +384,12 @@ add_custom_target(ci_test_amalgamation
COMMAND ${Python3_EXECUTABLE} ${tool_dir}/amalgamate.py -c ${tool_dir}/config_json.json -s . COMMAND ${Python3_EXECUTABLE} ${tool_dir}/amalgamate.py -c ${tool_dir}/config_json.json -s .
COMMAND ${Python3_EXECUTABLE} ${tool_dir}/amalgamate.py -c ${tool_dir}/config_json_fwd.json -s . COMMAND ${Python3_EXECUTABLE} ${tool_dir}/amalgamate.py -c ${tool_dir}/config_json_fwd.json -s .
COMMAND cp ${PROJECT_SOURCE_DIR}/include/nlohmann/json_literals.hpp ${include_dir}/json_literals.hpp
COMMAND venv_astyle/bin/astyle --project=tools/astyle/.astylerc --suffix=none ${include_dir}/json.hpp ${include_dir}/json_fwd.hpp COMMAND venv_astyle/bin/astyle --project=tools/astyle/.astylerc --suffix=none ${include_dir}/json.hpp ${include_dir}/json_fwd.hpp
COMMAND diff ${include_dir}/json.hpp~ ${include_dir}/json.hpp COMMAND diff ${include_dir}/json.hpp~ ${include_dir}/json.hpp
COMMAND diff ${include_dir}/json_fwd.hpp~ ${include_dir}/json_fwd.hpp COMMAND diff ${include_dir}/json_fwd.hpp~ ${include_dir}/json_fwd.hpp
COMMAND diff ${include_dir}/json_literals.hpp~ ${include_dir}/json_literals.hpp
COMMAND venv_astyle/bin/astyle --project=tools/astyle/.astylerc --suffix=orig ${INDENT_FILES} COMMAND venv_astyle/bin/astyle --project=tools/astyle/.astylerc --suffix=orig ${INDENT_FILES}
COMMAND for FILE in `find . -name '*.orig'`\; do false \; done COMMAND for FILE in `find . -name '*.orig'`\; do false \; done
+4 -10
View File
@@ -25,15 +25,10 @@ and `ensure_ascii` parameters.
result consists of ASCII characters only. result consists of ASCII characters only.
`error_handler` (in) `error_handler` (in)
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md)): : how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
- `strict`: throw a [`type_error`](../../home/exceptions.md#type-errors) exception in case a decoding error occurs with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
(default), output unchanged, and invalid bytes are dropped)).
- `replace`: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER),
- `ignore`: ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the output unchanged,
and invalid bytes are dropped, and
- `keep`: keep invalid UTF-8 sequences during serialization; all bytes are copied to the output unchanged, so the
result is not valid UTF-8.
## Return value ## Return value
@@ -99,4 +94,3 @@ Binary values are serialized as an object containing two keys:
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0. - Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
- Error handlers added in version 3.4.0. - Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0. - Serialization of binary values added in version 3.8.0.
- Error handler value `keep` added in version 3.13.0.
@@ -4,13 +4,12 @@
enum class error_handler_t { enum class error_handler_t {
strict, strict,
replace, replace,
ignore, ignore
keep
}; };
``` ```
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Four values are differentiated: `basic_json` value. Three values are differentiated:
strict strict
: throw a `type_error` exception in case of invalid UTF-8 : throw a `type_error` exception in case of invalid UTF-8
@@ -21,12 +20,6 @@ replace
ignore ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped : ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences; all bytes are copied to the output unchanged. Valid characters are still escaped as
usual (e.g., `"`, `\\`, and control characters), so the result has valid JSON syntax, but it is not valid UTF-8.
In particular, [`parse`](parse.md) rejects it, and with `ensure_ascii` set to `true`, the invalid bytes are the
only non-ASCII bytes of the output.
## Examples ## Examples
??? example ??? example
@@ -47,4 +40,3 @@ keep
## Version history ## Version history
- Added in version 3.4.0. - Added in version 3.4.0.
- Added value `keep` in version 3.13.0.
+2 -2
View File
@@ -80,8 +80,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were - Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key, - Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
but not found types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity ## Complexity
@@ -73,8 +73,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from - Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key, - Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
but not found types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity ## Complexity
+1
View File
@@ -28,6 +28,7 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_HAS_RANGES**](json_has_ranges.md) - control `std::ranges` support - [**JSON_HAS_RANGES**](json_has_ranges.md) - control `std::ranges` support
- [**JSON_HAS_STD_FORMAT**](json_has_std_format.md) - control `std::format`/`std::formatter` support - [**JSON_HAS_STD_FORMAT**](json_has_std_format.md) - control `std::format`/`std::formatter` support
- [**JSON_HAS_THREE_WAY_COMPARISON**](json_has_three_way_comparison.md) - control 3-way comparison support - [**JSON_HAS_THREE_WAY_COMPARISON**](json_has_three_way_comparison.md) - control 3-way comparison support
- [**JSON_NO_AUTOMATIC_UDLS**](json_no_automatic_udls.md) - do not include the user-defined string literals (UDLs) automatically
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers - [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
- [**JSON_NO_THREAD_LOCAL**](json_no_thread_local.md) - switch off the use of `thread_local` storage - [**JSON_NO_THREAD_LOCAL**](json_no_thread_local.md) - switch off the use of `thread_local` storage
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers - [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
@@ -0,0 +1,69 @@
# JSON_NO_AUTOMATIC_UDLS
```cpp
#define JSON_NO_AUTOMATIC_UDLS
```
When defined, `<nlohmann/json.hpp>` does not include `<nlohmann/json_literals.hpp>`, so the user-defined string
literals [`operator""_json`](../operator_literal_json.md) and
[`operator""_json_pointer`](../operator_literal_json_pointer.md) are not declared. Include
`<nlohmann/json_literals.hpp>` in the files that use them.
The literals are ordinary inline functions whose bodies call the parser, so every translation unit that includes them
instantiates the parser — even if it never parses anything itself. Defining `JSON_NO_AUTOMATIC_UDLS` for a whole project
avoids this cost in translation units that do not parse (e.g., ones that only define types and conversions or pass
`json` values around) and reduces their compile time.
## Default definition
By default, `#!cpp JSON_NO_AUTOMATIC_UDLS` is not defined, and `<nlohmann/json.hpp>` includes
`<nlohmann/json_literals.hpp>`.
```cpp
#undef JSON_NO_AUTOMATIC_UDLS
```
## Notes
!!! info "Header `<nlohmann/json_literals.hpp>`"
The header includes `<nlohmann/json.hpp>` itself and places the literals according to
[`JSON_USE_GLOBAL_UDLS`](json_use_global_udls.md). It is part of the multi-header sources (`include/nlohmann`)
and of the single-header sources (`single_include/nlohmann`), next to `json.hpp`.
!!! info "C++ modules"
The `nlohmann.json` [module](../../features/modules.md) always exports the literals, regardless of this macro.
## Examples
??? example
The code below includes the library without the literals and adds them in a single translation unit.
```cpp
// compiled with -DJSON_NO_AUTOMATIC_UDLS for the whole project
#include <nlohmann/json.hpp>
// this file uses the literals, so it includes them explicitly
#include <nlohmann/json_literals.hpp>
int main()
{
auto j = R"({"foo": 42})"_json;
return j.at("/foo"_json_pointer) == 42 ? 0 : 1;
}
```
Without the include of `<nlohmann/json_literals.hpp>`, the code would fail to compile.
## See also
- [`operator""_json`](../operator_literal_json.md)
- [`operator""_json_pointer`](../operator_literal_json_pointer.md)
- [`JSON_USE_GLOBAL_UDLS`](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [Compile times](../../integration/compile_times.md) - options to reduce compile times
## Version history
- Added in version 3.13.0.
@@ -15,7 +15,7 @@ The default value is `1`.
#define JSON_USE_GLOBAL_UDLS 1 #define JSON_USE_GLOBAL_UDLS 1
``` ```
When the macro is not defined, the library will define it to its default value. When the macro is not defined, the library behaves as if it were defined to its default value.
## Notes ## Notes
@@ -32,6 +32,11 @@ When the macro is not defined, the library will define it to its default value.
[`JSON_GlobalUDLs`](../../integration/cmake.md#json_globaludls) (`ON` by default) which defines [`JSON_GlobalUDLs`](../../integration/cmake.md#json_globaludls) (`ON` by default) which defines
`JSON_USE_GLOBAL_UDLS` accordingly. `JSON_USE_GLOBAL_UDLS` accordingly.
!!! info "Leaving out the literals"
If [`JSON_NO_AUTOMATIC_UDLS`](json_no_automatic_udls.md) is defined, the literals are only declared where
`<nlohmann/json_literals.hpp>` is included; this macro then applies to that header.
## Examples ## Examples
??? example "Example 1: Default behavior" ??? example "Example 1: Default behavior"
@@ -92,6 +97,7 @@ When the macro is not defined, the library will define it to its default value.
- [`operator""_json`](../operator_literal_json.md) - [`operator""_json`](../operator_literal_json.md)
- [`operator""_json_pointer`](../operator_literal_json_pointer.md) - [`operator""_json_pointer`](../operator_literal_json_pointer.md)
- [`JSON_NO_AUTOMATIC_UDLS`](json_no_automatic_udls.md) - do not include the user-defined string literals automatically
- [:simple-cmake: JSON_GlobalUDLs](../../integration/cmake.md#json_globaludls) - CMake option to control the macro - [:simple-cmake: JSON_GlobalUDLs](../../integration/cmake.md#json_globaludls) - CMake option to control the macro
## Version history ## Version history
@@ -18,7 +18,9 @@ using namespace nlohmann;
``` ```
This is suggested to ease migration to the next major version release of the library. See This is suggested to ease migration to the next major version release of the library. See
[`JSON_USE_GLOBAL_UDLS`](macros/json_use_global_udls.md#notes) for details. [`JSON_USE_GLOBAL_UDLS`](macros/json_use_global_udls.md#notes) for details. The operator is declared in header
`<nlohmann/json_literals.hpp>`, which `<nlohmann/json.hpp>` includes unless
[`JSON_NO_AUTOMATIC_UDLS`](macros/json_no_automatic_udls.md) is defined.
## Parameters ## Parameters
@@ -59,6 +61,8 @@ Linear.
## See also ## See also
- [Creating JSON values](../features/creating_values.md) - the article on creating JSON values - [Creating JSON values](../features/creating_values.md) - the article on creating JSON values
- [JSON_NO_AUTOMATIC_UDLS](macros/json_no_automatic_udls.md) - do not include the user-defined string literals
automatically
## Version history ## Version history
@@ -17,7 +17,9 @@ using namespace nlohmann::literals::json_literals;
using namespace nlohmann; using namespace nlohmann;
``` ```
This is suggested to ease migration to the next major version release of the library. See This is suggested to ease migration to the next major version release of the library. See
[`JSON_USE_GLOBAL_UDLS`](macros/json_use_global_udls.md#notes) for details. [`JSON_USE_GLOBAL_UDLS`](macros/json_use_global_udls.md#notes) for details. The operator is declared in header
`<nlohmann/json_literals.hpp>`, which `<nlohmann/json.hpp>` includes unless
[`JSON_NO_AUTOMATIC_UDLS`](macros/json_no_automatic_udls.md) is defined.
## Parameters ## Parameters
@@ -58,6 +60,8 @@ Linear.
## See also ## See also
- [json_pointer](json_pointer/index.md) - type to represent JSON Pointers - [json_pointer](json_pointer/index.md) - type to represent JSON Pointers
- [JSON_NO_AUTOMATIC_UDLS](macros/json_no_automatic_udls.md) - do not include the user-defined string literals
automatically
## Version history ## Version history
@@ -1,4 +1,3 @@
#include <iomanip>
#include <iostream> #include <iostream>
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
@@ -22,12 +21,4 @@ int main()
<< "\nstring with ignored invalid characters: " << "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore) << j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< '\n'; << '\n';
// the invalid byte is kept; print the result byte-wise to make it visible
std::cout << "string with kept invalid characters:";
for (const unsigned char c : j_invalid.dump(-1, ' ', false, json::error_handler_t::keep))
{
std::cout << ' ' << std::hex << std::setw(2) << std::setfill('0') << static_cast<int>(c);
}
std::cout << '\n';
} }
@@ -1,4 +1,3 @@
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9 [json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü" string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü" string with ignored invalid characters: "äü"
string with kept invalid characters: 22 c3 a4 a9 c3 bc 22
@@ -174,7 +174,20 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Object keys" !!! warning "Object keys"
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps with keys other than UTF-8 strings are rejected. CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps
with keys other than text strings (major type 3) are rejected with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set
to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. This is a deliberate restriction of the library's JSON value model, not an oversight: formats built on CBOR
maps with integer keys, such as COSE ([RFC 9052](https://www.rfc-editor.org/rfc/rfc9052.html)) or CWT
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings" !!! warning "UTF-8 validation of text strings"
@@ -138,6 +138,21 @@ The library maps MessagePack types to JSON value types as follows:
Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`. Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`.
!!! warning "Object keys"
MessagePack allows map keys of any type, whereas JSON only allows strings as keys in object values. Like the
JSON-compatible [profile](https://github.com/msgpack/msgpack/blob/master/spec.md#profile) sketched in the
MessagePack specification, this library restricts map keys to `str` values. Maps with keys of any other type are
rejected with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values" !!! warning "UTF-8 validation of string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8. The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
+9
View File
@@ -83,6 +83,15 @@ When defined, default parse and serialize functions for enums are excluded and h
See [full documentation of `JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md). See [full documentation of `JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md).
## `JSON_NO_AUTOMATIC_UDLS`
When defined, `<nlohmann/json.hpp>` does not include `<nlohmann/json_literals.hpp>` with the user-defined string literals
`operator""_json` and `operator""_json_pointer`. This reduces the compile time of translation units that do not use
them, because the literals instantiate the parser in every translation unit that includes them. Include
`<nlohmann/json_literals.hpp>` where the literals are needed.
See [full documentation of `JSON_NO_AUTOMATIC_UDLS`](../api/macros/json_no_automatic_udls.md).
## `JSON_NO_IO` ## `JSON_NO_IO`
When defined, headers `<cstdio>`, `<ios>`, `<iosfwd>`, `<istream>`, and `<ostream>` are not included and parse functions When defined, headers `<cstdio>`, `<ios>`, `<iosfwd>`, `<istream>`, and `<ostream>` are not included and parse functions
+3
View File
@@ -40,6 +40,9 @@ Only the following symbols are exported from `nlohmann.json`:
- `nlohmann::literals::json_literals::operator""_json` - `nlohmann::literals::json_literals::operator""_json`
- `nlohmann::literals::json_literals::operator""_json_pointer` - `nlohmann::literals::json_literals::operator""_json_pointer`
The module always exports the two user-defined string literals, even if
[`JSON_NO_AUTOMATIC_UDLS`](../api/macros/json_no_automatic_udls.md) is defined when building it.
Additionally, the following `nlohmann::detail` symbols are exported, solely to work around an MSVC compilation issue Additionally, the following `nlohmann::detail` symbols are exported, solely to work around an MSVC compilation issue
([#3970](https://github.com/nlohmann/json/issues/3970)). They are implementation details, not part of the public API, ([#3970](https://github.com/nlohmann/json/issues/3970)). They are implementation details, not part of the public API,
and should not be used directly: and should not be used directly:
@@ -64,7 +64,6 @@ serialization fails by default. The fourth argument of `dump` selects an
- `strict` (default) — throw a [`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) exception. - `strict` (default) — throw a [`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) exception.
- `replace` — replace invalid bytes with the Unicode replacement character U+FFFD (`�`). - `replace` — replace invalid bytes with the Unicode replacement character U+FFFD (`�`).
- `ignore` — silently drop invalid bytes. - `ignore` — silently drop invalid bytes.
- `keep` — copy invalid bytes to the output unchanged; the result is not valid UTF-8.
??? example ??? example
+9 -3
View File
@@ -343,13 +343,20 @@ A string could not be read from a [binary format](../features/binary_formats/ind
string was read where one was required (for instance as a map key), the string's length specification is invalid, or string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8. the string's bytes are not valid UTF-8.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
[CBOR](../features/binary_formats/cbor.md) and [MessagePack](../features/binary_formats/messagepack.md).
!!! failure "Example messages" !!! failure "Example messages"
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
``` ```
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C
``` ```
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82 [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82
@@ -755,7 +762,6 @@ The `dump()` function only works with UTF-8 encoded strings; that is, if you ass
- Pass an error handler as last parameter to the `dump()` function to avoid this exception: - Pass an error handler as last parameter to the `dump()` function to avoid this exception:
- `json::error_handler_t::replace` will replace invalid bytes sequences with `U+FFFD` - `json::error_handler_t::replace` will replace invalid bytes sequences with `U+FFFD`
- `json::error_handler_t::ignore` will silently ignore invalid byte sequences - `json::error_handler_t::ignore` will silently ignore invalid byte sequences
- `json::error_handler_t::keep` will copy invalid byte sequences to the output unchanged
### json.exception.type_error.317 ### json.exception.type_error.317
+1 -1
View File
@@ -85,7 +85,7 @@ The library supports **Unicode input** as follows:
- The library will not replace [Unicode noncharacters](http://www.unicode.org/faq/private_use.html#nonchar1). - The library will not replace [Unicode noncharacters](http://www.unicode.org/faq/private_use.html#nonchar1).
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors. - Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs. - The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers. - When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding. In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding.
@@ -0,0 +1,149 @@
# Compile times
The library is header-only and makes heavy use of templates, so every translation unit that includes
`<nlohmann/json.hpp>` pays for parsing the header and instantiating what it uses. This page lists the options to reduce
that cost, ordered by how much they typically save.
!!! info "Measurements"
The numbers below are medians of nine runs compiling a single translation unit with `-std=c++17 -c` against the
single-header version, with Apple clang and GCC 16 on macOS (Apple silicon). They show the order of magnitude to
expect; measure your own code before and after a change.
## Include `json_fwd.hpp` in headers
Header files that only need to *name* the `json` type — for function declarations, members held by pointer or
reference, or friend declarations — can include `<nlohmann/json_fwd.hpp>` instead of `<nlohmann/json.hpp>`. It only
forward-declares `basic_json`, `json`, `ordered_json`, `json_pointer`, and `adl_serializer`. The translation units that
actually use the values then include `<nlohmann/json.hpp>`.
```cpp title="person.hpp"
#pragma once
#include <nlohmann/json_fwd.hpp>
struct person;
void to_json(nlohmann::json& j, const person& p);
void from_json(const nlohmann::json& j, person& p);
```
```cpp title="person.cpp"
#include "person.hpp"
#include <nlohmann/json.hpp>
void to_json(nlohmann::json& j, const person& p) { /* ... */ }
void from_json(const nlohmann::json& j, person& p) { /* ... */ }
```
| Compiler | `json.hpp` (`-O0`) | `json_fwd.hpp` (`-O0`) | Change |
|-------------|-------------------:|-----------------------:|-------:|
| Apple clang | 704 ms | 329 ms | −53% |
| GCC 16 | 779 ms | 242 ms | −69% |
This is the most effective option, because it avoids the full header in every translation unit that includes
*your* headers.
## Opt out of the automatic user-defined string literals
The user-defined string literals [`operator""_json`](../api/operator_literal_json.md) and
[`operator""_json_pointer`](../api/operator_literal_json_pointer.md) are ordinary inline functions whose bodies call the
parser. As `<nlohmann/json.hpp>` includes them by default, every translation unit instantiates the parser, even if it
never parses anything itself.
Define [`JSON_NO_AUTOMATIC_UDLS`](../api/macros/json_no_automatic_udls.md) for the whole project and include
`<nlohmann/json_literals.hpp>` only in the files that use the literals:
```cmake
target_compile_definitions(my_target PRIVATE JSON_NO_AUTOMATIC_UDLS)
```
```cpp
#include <nlohmann/json.hpp>
#include <nlohmann/json_literals.hpp> // only where "..."_json is used
```
The saving applies to translation units that do not parse JSON, for example ones that define types and their
conversions or only pass `json` values around:
| Compiler | Translation unit | Default (`-O0` / `-O2`) | `JSON_NO_AUTOMATIC_UDLS` (`-O0` / `-O2`) | Change |
|-------------|------------------|------------------------:|-----------------------------------------:|------------:|
| Apple clang | model | 776 ms / 846 ms | 629 ms / 692 ms | −19% / −18% |
| GCC 16 | model | 1022 ms / 1120 ms | 882 ms / 965 ms | −14% / −14% |
| Apple clang | parsing | 992 ms / 1815 ms | 1006 ms / 1823 ms | +1% / 0% |
| GCC 16 | parsing | 2018 ms / 3420 ms | 1990 ms / 3454 ms | −1% / +1% |
Translation units that include only the header save up to a third. Translation units that parse anyway instantiate
the parser regardless and see no difference.
## Instantiate `basic_json` once
Each translation unit instantiates the member functions of `nlohmann::json` it uses. An explicit instantiation
declaration tells the compiler that the non-template members are instantiated elsewhere, so it can skip them:
```cpp title="json_instance.hpp"
#pragma once
#include <nlohmann/json.hpp>
extern template class nlohmann::basic_json<>;
```
```cpp title="json_instance.cpp"
#include "json_instance.hpp"
template class nlohmann::basic_json<>;
```
Include `json_instance.hpp` instead of `<nlohmann/json.hpp>` and compile and link `json_instance.cpp` once.
| Compiler | Translation unit | Default (`-O0` / `-O2`) | `extern template` (`-O0` / `-O2`) | Change |
|-------------|---------------------|------------------------:|----------------------------------:|------------:|
| Apple clang | parsing | 992 ms / 1815 ms | 953 ms / 1625 ms | −4% / −10% |
| GCC 16 | parsing | 2018 ms / 3420 ms | 1522 ms / 2728 ms | −25% / −20% |
| Apple clang | `json_instance.cpp` | — | 2166 ms / 4660 ms | — |
| GCC 16 | `json_instance.cpp` | — | 5085 ms / 10616 ms | — |
Notes:
- The saving grows with the number of translation units that use `json`, while the instantiation translation unit is
compiled only once (and is rarely recompiled, as it does not depend on your code).
- Member function templates (such as `get<T>()`, `parse(InputType&&)`, or `value(key, default)`) are not covered by
the explicit instantiation and are still instantiated where they are used.
- The declaration covers exactly `nlohmann::json`. Add the same lines for `nlohmann::ordered_json`
(`nlohmann::basic_json<nlohmann::ordered_map>`) or your own `basic_json` specializations if you use them.
## Use C++20 modules
With a toolchain that supports named modules, `import nlohmann.json;` compiles the library once into a module and
avoids parsing the header in every translation unit. See [Modules](../features/modules.md) for requirements and known
issues. Module support is experimental and currently depends heavily on the compiler version.
## Use precompiled headers
Build systems can precompile `<nlohmann/json.hpp>` together with other stable headers, for example with CMake's
[`target_precompile_headers`](https://cmake.org/cmake/help/latest/command/target_precompile_headers.html):
```cmake
target_precompile_headers(my_target PRIVATE <nlohmann/json.hpp>)
```
This removes the cost of parsing the header, but not of instantiating templates in each translation unit, so it
combines well with the options above.
## Options without effect on compile times
Some configuration macros change what the library declares, but do not measurably change compile times:
| Macro | Apple clang, model (`-O0` / `-O2`) | GCC 16, model (`-O0` / `-O2`) |
|------------------------------------------------------------------------|-----------------------------------:|------------------------------:|
| default | 776 ms / 846 ms | 1022 ms / 1120 ms |
| [`JSON_NO_IO`](../api/macros/json_no_io.md) | 764 ms / 836 ms | 1022 ms / 1117 ms |
| [`JSON_USE_GLOBAL_UDLS`](../api/macros/json_use_global_udls.md)`=0` | 763 ms / 852 ms | 1019 ms / 1106 ms |
`JSON_USE_GLOBAL_UDLS` only controls *where* the literals are declared; to avoid their cost, use
`JSON_NO_AUTOMATIC_UDLS` instead.
## See also
- [`JSON_NO_AUTOMATIC_UDLS`](../api/macros/json_no_automatic_udls.md) - do not include the user-defined string
literals automatically
- [Modules](../features/modules.md) - C++20 module support
- [Header only](index.md) - including the library
+4 -1
View File
@@ -15,4 +15,7 @@ Clang).
You can further use file You can further use file
[`single_include/nlohmann/json_fwd.hpp`](https://github.com/nlohmann/json/blob/develop/single_include/nlohmann/json_fwd.hpp) [`single_include/nlohmann/json_fwd.hpp`](https://github.com/nlohmann/json/blob/develop/single_include/nlohmann/json_fwd.hpp)
for forward declarations. for forward declarations (see [Compile times](compile_times.md)), and file
[`single_include/nlohmann/json_literals.hpp`](https://github.com/nlohmann/json/blob/develop/single_include/nlohmann/json_literals.hpp)
for the user-defined string literals if you define
[`JSON_NO_AUTOMATIC_UDLS`](../api/macros/json_no_automatic_udls.md).
+2
View File
@@ -106,6 +106,7 @@ nav:
- integration/cmake.md - integration/cmake.md
- integration/package_managers.md - integration/package_managers.md
- integration/pkg-config.md - integration/pkg-config.md
- integration/compile_times.md
- API Documentation: - API Documentation:
- basic_json: - basic_json:
- 'Overview': api/basic_json/index.md - 'Overview': api/basic_json/index.md
@@ -293,6 +294,7 @@ nav:
- 'JSON_HAS_STATIC_RTTI': api/macros/json_has_static_rtti.md - 'JSON_HAS_STATIC_RTTI': api/macros/json_has_static_rtti.md
- 'JSON_HAS_STD_FORMAT': api/macros/json_has_std_format.md - 'JSON_HAS_STD_FORMAT': api/macros/json_has_std_format.md
- 'JSON_HAS_THREE_WAY_COMPARISON': api/macros/json_has_three_way_comparison.md - 'JSON_HAS_THREE_WAY_COMPARISON': api/macros/json_has_three_way_comparison.md
- 'JSON_NO_AUTOMATIC_UDLS': api/macros/json_no_automatic_udls.md
- 'JSON_NOEXCEPTION': api/macros/json_noexception.md - 'JSON_NOEXCEPTION': api/macros/json_noexception.md
- 'JSON_NO_IO': api/macros/json_no_io.md - 'JSON_NO_IO': api/macros/json_no_io.md
- 'JSON_NO_THREAD_LOCAL': api/macros/json_no_thread_local.md - 'JSON_NO_THREAD_LOCAL': api/macros/json_no_thread_local.md
+168 -2
View File
@@ -1324,6 +1324,80 @@ class binary_reader
} }
} }
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*! /*!
@brief reads a definite-length CBOR byte array @brief reads a definite-length CBOR byte array
@@ -1568,7 +1642,7 @@ class binary_reader
if (top.is_object) if (top.is_object)
{ {
key.clear(); key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key))) if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
{ {
return false; return false;
} }
@@ -2069,6 +2143,98 @@ class binary_reader
} }
} }
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*! /*!
@brief reads a MessagePack byte array @brief reads a MessagePack byte array
@@ -2231,7 +2397,7 @@ class binary_reader
{ {
get(); get();
key.clear(); key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key))) if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
{ {
return false; return false;
} }
+76 -39
View File
@@ -206,7 +206,6 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter)) : ia(std::move(adapter))
, ignore_comments(ignore_comments_) , ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_) , discard_number_values(discard_number_values_)
{} {}
@@ -222,8 +221,7 @@ class lexer : public lexer_base<BasicJsonType>
// locales // locales
///////////////////// /////////////////////
/// return the locale-dependent decimal point /// return the decimal point of the current locale
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept static char get_decimal_point() noexcept
{ {
const auto* loc = localeconv(); const auto* loc = localeconv();
@@ -1092,9 +1090,10 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned, token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise token_type::parse_error otherwise
@note The scanner is independent of the current locale. Internally, the @note The scanner is independent of the current locale: token_buffer
locale's decimal point is used instead of `.` to work with the always holds `.`. Only the std::strtod fallback of convert_number()
locale-dependent converters. depends on the locale, and it looks up the decimal point right
before converting (see convert_float_locale_aware()).
*/ */
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated. token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{ {
@@ -1183,7 +1182,7 @@ scan_number_zero:
{ {
case '.': case '.':
{ {
add(decimal_point_char); add(current);
decimal_point_position = token_buffer.size() - 1; decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1; goto scan_number_decimal1;
} }
@@ -1220,7 +1219,7 @@ scan_number_any1:
case '.': case '.':
{ {
add(decimal_point_char); add(current);
decimal_point_position = token_buffer.size() - 1; decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1; goto scan_number_decimal1;
} }
@@ -1462,9 +1461,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only // Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the // while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note // answer - so this loop is skipped for all but a few tokens. The
// token_buffer holds the locale's decimal point, so the fraction is // fraction is located through decimal_point_position rather than by
// located through decimal_point_position rather than by searching '.'. // searching '.'.
if (lead_zero != 0) if (lead_zero != 0)
{ {
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -1482,8 +1481,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type @brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the scan_number() state machine or by the contiguous fast path) and holds '.'
locale decimal point in place of '.'. Integers are parsed first and fall as decimal point, independent of the locale. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce back to floating point on overflow. This is shared so both scanners produce
identical results. identical results.
@@ -1563,7 +1562,7 @@ scan_number_done:
// integer conversion above overflowed. Prefer std::from_chars // integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available; // (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the // otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod. // locale-aware strtof/strtod/strtold.
if (parse_float_from_chars(num_begin, num_end, value_float)) if (parse_float_from_chars(num_begin, num_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
@@ -1572,26 +1571,75 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on // extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json // high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end) if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float)) && parse_float_fast(num_begin, num_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
} }
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg) convert_float_locale_aware();
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
return token_type::value_float; return token_type::value_float;
} }
/*!
@brief convert the float in token_buffer with strtof/strtod/strtold
These functions expect the decimal point of the *current* locale, so it is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX
handler, or another thread) must not truncate the value (#5198). The
token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the
lookup and the call, and the conversion is repeated with the new decimal
point. If the decimal point did not change, a retry cannot succeed: the
locale's decimal point is not a single character (e.g., the two-byte
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
The value strtod parsed up to that point is kept, as before this change.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
*/
void convert_float_locale_aware()
{
const bool has_dot = decimal_point_position != std::string::npos;
char decimal_point = get_decimal_point();
for (;;)
{
const bool substitute = has_dot && decimal_point != '.';
if (substitute)
{
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
if (substitute)
{
// get_string() hands the token to the SAX interface with '.'
token_buffer[decimal_point_position] = '.';
}
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
const char current_decimal_point = get_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = current_decimal_point;
}
}
/*! /*!
@brief contiguous fast path for scanning a number @brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and (as scan_number() does) and
returns the token type. On anything it does not fully recognize as a returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -1707,16 +1755,11 @@ scan_number_done:
} }
#endif #endif
// materialize the token exactly as scan_number() would, substituting the // materialize the token exactly as scan_number() would. reset() already
// locale decimal point so convert_number()'s strtof fallback stays valid. // cleared token_buffer, so append() fills it (assign() is avoided
// reset() already cleared token_buffer, so append() fills it (assign() is // because custom string_t types need not provide it)
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len); token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos) decimal_point_position = dot_index;
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1); ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1); position.chars_read_total += (len - 1);
@@ -1983,11 +2026,7 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once) /// return current string value (implicitly resets the token; useful only once)
string_t& get_string() string_t& get_string()
{ {
// translate decimal points from locale back to '.' (#4084) // a number token holds '.' regardless of the locale (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
return token_buffer; return token_buffer;
} }
@@ -2283,9 +2322,7 @@ scan_number_done:
number_unsigned_t value_unsigned = 0; number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0; number_float_t value_float = 0;
/// the decimal point /// the position of the decimal point in token_buffer
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
std::size_t decimal_point_position = std::string::npos; std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the /// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
+8 -13
View File
@@ -118,14 +118,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload long double keep the std::strtof/std::strtold paths (see the templated overload
below). below).
@param[in] first pointer to the first character of the number @param[in] first pointer to the first character of the number
@param[in] last pointer past the last character @param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character @param[out] out the parsed value on success
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod @return true if the value was parsed exactly; false to fall back to strtod
*/ */
template<typename DecimalPointType> inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{ {
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 #if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in // Clinger's fast path is only exact when double operations are evaluated in
@@ -136,7 +134,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
// std::from_chars / std::strtod path. // std::from_chars / std::strtod path.
static_cast<void>(first); static_cast<void>(first);
static_cast<void>(last); static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out); static_cast<void>(out);
return false; return false;
#else #else
@@ -175,7 +172,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
++num_digits; ++num_digits;
fractional_digits += static_cast<int>(seen_dot); fractional_digits += static_cast<int>(seen_dot);
} }
else if (static_cast<DecimalPointType>(c) == decimal_point) else if (c == '.')
{ {
if (JSON_HEDLEY_UNLIKELY(seen_dot)) if (JSON_HEDLEY_UNLIKELY(seen_dot))
{ {
@@ -260,8 +257,8 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
} }
/// fast float path is only exact for `double`; decline for float/long double /// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType> template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
{ {
return false; return false;
} }
@@ -273,9 +270,7 @@ std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it __cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations). expects (side-stepping the P4168 divergence between implementations).
-4
View File
@@ -915,7 +915,3 @@ void templated_json_throw(ExceptionType exception)
#ifndef JSON_DISABLE_ENUM_SERIALIZATION #ifndef JSON_DISABLE_ENUM_SERIALIZATION
#define JSON_DISABLE_ENUM_SERIALIZATION 0 #define JSON_DISABLE_ENUM_SERIALIZATION 0
#endif #endif
#ifndef JSON_USE_GLOBAL_UDLS
#define JSON_USE_GLOBAL_UDLS 1
#endif
@@ -25,7 +25,6 @@
#undef JSON_INLINE_VARIABLE #undef JSON_INLINE_VARIABLE
#undef JSON_NO_UNIQUE_ADDRESS #undef JSON_NO_UNIQUE_ADDRESS
#undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_DISABLE_ENUM_SERIALIZATION
#undef JSON_USE_GLOBAL_UDLS
#ifndef JSON_TEST_KEEP_MACROS #ifndef JSON_TEST_KEEP_MACROS
#undef JSON_CATCH #undef JSON_CATCH
+1 -52
View File
@@ -48,8 +48,7 @@ enum class error_handler_t
{ {
strict, ///< throw a type_error exception in case of invalid UTF-8 strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences ignore ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
}; };
template<typename BasicJsonType> template<typename BasicJsonType>
@@ -1020,47 +1019,6 @@ class serializer
break; break;
} }
case error_handler_t::keep:
{
// drop whatever the incomplete sequence left in
// the buffer (only copied if !EnsureAscii) and copy
// the ill-formed bytes from the input instead
bytes = bytes_after_last_accept;
if (undumped_chars > 0)
{
// the pending bytes of the incomplete sequence
// are ill-formed; the current byte may be OK for
// itself, so we would like to read it again
for (std::size_t j = i - undumped_chars; j < i; ++j)
{
string_buffer[bytes++] = s[j];
}
--i;
}
else
{
// the current byte cannot start any sequence
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
} }
@@ -1106,15 +1064,6 @@ class serializer
break; break;
} }
case error_handler_t::keep:
{
// write all accepted bytes
put_buffer(string_buffer, bytes_after_last_accept);
// copy the bytes of the incomplete sequence unchanged
put_string(s, s.size() - undumped_chars, s.size());
break;
}
case error_handler_t::replace: case error_handler_t::replace:
{ {
// write all accepted bytes // write all accepted bytes
+8 -60
View File
@@ -6423,55 +6423,6 @@ std::string format_as(const NLOHMANN_BASIC_JSON_TPL& j)
return j.dump(); return j.dump();
} }
inline namespace literals
{
inline namespace json_literals
{
/// @brief user-defined string literal for JSON values
/// @sa https://json.nlohmann.me/api/basic_json/operator_literal_json/
JSON_HEDLEY_NON_NULL(1)
#if !defined(JSON_HEDLEY_GCC_VERSION) || JSON_HEDLEY_GCC_VERSION_CHECK(4,9,0)
inline nlohmann::json operator""_json(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json operator"" _json(const char* s, std::size_t n)
#endif
{
return nlohmann::json::parse(s, s + n);
}
#if defined(__cpp_char8_t)
JSON_HEDLEY_NON_NULL(1)
inline nlohmann::json operator""_json(const char8_t* s, std::size_t n)
{
return nlohmann::json::parse(reinterpret_cast<const char*>(s),
reinterpret_cast<const char*>(s) + n);
}
#endif
/// @brief user-defined string literal for JSON pointer
/// @sa https://json.nlohmann.me/api/basic_json/operator_literal_json_pointer/
JSON_HEDLEY_NON_NULL(1)
#if !defined(JSON_HEDLEY_GCC_VERSION) || JSON_HEDLEY_GCC_VERSION_CHECK(4,9,0)
inline nlohmann::json::json_pointer operator""_json_pointer(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json::json_pointer operator"" _json_pointer(const char* s, std::size_t n)
#endif
{
return nlohmann::json::json_pointer(std::string(s, n));
}
#if defined(__cpp_char8_t)
inline nlohmann::json::json_pointer operator""_json_pointer(const char8_t* s, std::size_t n)
{
return nlohmann::json::json_pointer(std::string(reinterpret_cast<const char*>(s), n));
}
#endif
} // namespace json_literals
} // namespace literals
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
/////////////////////// ///////////////////////
@@ -6606,17 +6557,6 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
} // namespace std } // namespace std
#if JSON_USE_GLOBAL_UDLS
#if !defined(JSON_HEDLEY_GCC_VERSION) || JSON_HEDLEY_GCC_VERSION_CHECK(4,9,0)
using nlohmann::literals::json_literals::operator""_json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator""_json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#else
// GCC 4.8 requires a space between "" and suffix
using nlohmann::literals::json_literals::operator"" _json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator"" _json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#endif
#endif
#include <nlohmann/detail/macro_unscope.hpp> #include <nlohmann/detail/macro_unscope.hpp>
// End of GCC diagnostic pragmas for C++ modules support // End of GCC diagnostic pragmas for C++ modules support
@@ -6624,4 +6564,12 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#pragma GCC diagnostic pop #pragma GCC diagnostic pop
#endif #endif
// The user-defined string literals are in a separate header, because their
// bodies instantiate the parser in every translation unit that includes them.
// Define JSON_NO_AUTOMATIC_UDLS to include <nlohmann/json_literals.hpp> only
// where needed.
#ifndef JSON_NO_AUTOMATIC_UDLS
#include <nlohmann/json_literals.hpp>
#endif
#endif // INCLUDE_NLOHMANN_JSON_HPP_ #endif // INCLUDE_NLOHMANN_JSON_HPP_
+83
View File
@@ -0,0 +1,83 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#ifndef INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#define INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#include <cstddef> // size_t
#include <string> // string
#include <nlohmann/json.hpp>
// This header is included at the end of <nlohmann/json.hpp> unless
// JSON_NO_AUTOMATIC_UDLS is defined, and can be included on its own after that.
// Either way, the library's internal macros are no longer defined here (and the
// amalgamation inlines macro_scope.hpp only once), so only standard and public
// macros may be used below.
NLOHMANN_JSON_NAMESPACE_BEGIN
inline namespace literals
{
inline namespace json_literals
{
/// @brief user-defined string literal for JSON values
/// @sa https://json.nlohmann.me/api/operator_literal_json/
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
inline nlohmann::json operator""_json(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json operator"" _json(const char* s, std::size_t n)
#endif
{
return nlohmann::json::parse(s, s + n);
}
#if defined(__cpp_char8_t)
inline nlohmann::json operator""_json(const char8_t* s, std::size_t n)
{
return nlohmann::json::parse(reinterpret_cast<const char*>(s),
reinterpret_cast<const char*>(s) + n);
}
#endif
/// @brief user-defined string literal for JSON pointer
/// @sa https://json.nlohmann.me/api/operator_literal_json_pointer/
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
inline nlohmann::json::json_pointer operator""_json_pointer(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json::json_pointer operator"" _json_pointer(const char* s, std::size_t n)
#endif
{
return nlohmann::json::json_pointer(std::string(s, n));
}
#if defined(__cpp_char8_t)
inline nlohmann::json::json_pointer operator""_json_pointer(const char8_t* s, std::size_t n)
{
return nlohmann::json::json_pointer(std::string(reinterpret_cast<const char*>(s), n));
}
#endif
} // namespace json_literals
} // namespace literals
NLOHMANN_JSON_NAMESPACE_END
#if !defined(JSON_USE_GLOBAL_UDLS) || JSON_USE_GLOBAL_UDLS
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
using nlohmann::literals::json_literals::operator""_json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator""_json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#else
// GCC 4.8 requires a space between "" and suffix
using nlohmann::literals::json_literals::operator"" _json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator"" _json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#endif
#endif
#endif // INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
+1
View File
@@ -15,6 +15,7 @@ nlohmann_json_multiple_headers = declare_dependency(
if not meson.is_subproject() if not meson.is_subproject()
install_headers('single_include/nlohmann/json.hpp', subdir: 'nlohmann') install_headers('single_include/nlohmann/json.hpp', subdir: 'nlohmann')
install_headers('single_include/nlohmann/json_fwd.hpp', subdir: 'nlohmann') install_headers('single_include/nlohmann/json_fwd.hpp', subdir: 'nlohmann')
install_headers('single_include/nlohmann/json_literals.hpp', subdir: 'nlohmann')
pkgc = import('pkgconfig') pkgc = import('pkgconfig')
pkgc.generate(name: 'nlohmann_json', pkgc.generate(name: 'nlohmann_json',
+346 -171
View File
@@ -3327,10 +3327,6 @@ void templated_json_throw(ExceptionType exception)
#define JSON_DISABLE_ENUM_SERIALIZATION 0 #define JSON_DISABLE_ENUM_SERIALIZATION 0
#endif #endif
#ifndef JSON_USE_GLOBAL_UDLS
#define JSON_USE_GLOBAL_UDLS 1
#endif
#if JSON_HAS_THREE_WAY_COMPARISON #if JSON_HAS_THREE_WAY_COMPARISON
#include <compare> // partial_ordering #include <compare> // partial_ordering
#endif #endif
@@ -8605,14 +8601,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload long double keep the std::strtof/std::strtold paths (see the templated overload
below). below).
@param[in] first pointer to the first character of the number @param[in] first pointer to the first character of the number
@param[in] last pointer past the last character @param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character @param[out] out the parsed value on success
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod @return true if the value was parsed exactly; false to fall back to strtod
*/ */
template<typename DecimalPointType> inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{ {
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 #if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in // Clinger's fast path is only exact when double operations are evaluated in
@@ -8623,7 +8617,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
// std::from_chars / std::strtod path. // std::from_chars / std::strtod path.
static_cast<void>(first); static_cast<void>(first);
static_cast<void>(last); static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out); static_cast<void>(out);
return false; return false;
#else #else
@@ -8662,7 +8655,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
++num_digits; ++num_digits;
fractional_digits += static_cast<int>(seen_dot); fractional_digits += static_cast<int>(seen_dot);
} }
else if (static_cast<DecimalPointType>(c) == decimal_point) else if (c == '.')
{ {
if (JSON_HEDLEY_UNLIKELY(seen_dot)) if (JSON_HEDLEY_UNLIKELY(seen_dot))
{ {
@@ -8747,8 +8740,8 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
} }
/// fast float path is only exact for `double`; decline for float/long double /// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType> template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
{ {
return false; return false;
} }
@@ -8760,9 +8753,7 @@ std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it __cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations). expects (side-stepping the P4168 divergence between implementations).
@@ -9303,7 +9294,6 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter)) : ia(std::move(adapter))
, ignore_comments(ignore_comments_) , ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_) , discard_number_values(discard_number_values_)
{} {}
@@ -9319,8 +9309,7 @@ class lexer : public lexer_base<BasicJsonType>
// locales // locales
///////////////////// /////////////////////
/// return the locale-dependent decimal point /// return the decimal point of the current locale
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept static char get_decimal_point() noexcept
{ {
const auto* loc = localeconv(); const auto* loc = localeconv();
@@ -10189,9 +10178,10 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned, token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise token_type::parse_error otherwise
@note The scanner is independent of the current locale. Internally, the @note The scanner is independent of the current locale: token_buffer
locale's decimal point is used instead of `.` to work with the always holds `.`. Only the std::strtod fallback of convert_number()
locale-dependent converters. depends on the locale, and it looks up the decimal point right
before converting (see convert_float_locale_aware()).
*/ */
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated. token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{ {
@@ -10280,7 +10270,7 @@ scan_number_zero:
{ {
case '.': case '.':
{ {
add(decimal_point_char); add(current);
decimal_point_position = token_buffer.size() - 1; decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1; goto scan_number_decimal1;
} }
@@ -10317,7 +10307,7 @@ scan_number_any1:
case '.': case '.':
{ {
add(decimal_point_char); add(current);
decimal_point_position = token_buffer.size() - 1; decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1; goto scan_number_decimal1;
} }
@@ -10559,9 +10549,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only // Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the // while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note // answer - so this loop is skipped for all but a few tokens. The
// token_buffer holds the locale's decimal point, so the fraction is // fraction is located through decimal_point_position rather than by
// located through decimal_point_position rather than by searching '.'. // searching '.'.
if (lead_zero != 0) if (lead_zero != 0)
{ {
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -10579,8 +10569,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type @brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the scan_number() state machine or by the contiguous fast path) and holds '.'
locale decimal point in place of '.'. Integers are parsed first and fall as decimal point, independent of the locale. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce back to floating point on overflow. This is shared so both scanners produce
identical results. identical results.
@@ -10660,7 +10650,7 @@ scan_number_done:
// integer conversion above overflowed. Prefer std::from_chars // integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available; // (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the // otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod. // locale-aware strtof/strtod/strtold.
if (parse_float_from_chars(num_begin, num_end, value_float)) if (parse_float_from_chars(num_begin, num_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
@@ -10669,26 +10659,75 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on // extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json // high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end) if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float)) && parse_float_fast(num_begin, num_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
} }
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg) convert_float_locale_aware();
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
return token_type::value_float; return token_type::value_float;
} }
/*!
@brief convert the float in token_buffer with strtof/strtod/strtold
These functions expect the decimal point of the *current* locale, so it is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX
handler, or another thread) must not truncate the value (#5198). The
token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the
lookup and the call, and the conversion is repeated with the new decimal
point. If the decimal point did not change, a retry cannot succeed: the
locale's decimal point is not a single character (e.g., the two-byte
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
The value strtod parsed up to that point is kept, as before this change.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
*/
void convert_float_locale_aware()
{
const bool has_dot = decimal_point_position != std::string::npos;
char decimal_point = get_decimal_point();
for (;;)
{
const bool substitute = has_dot && decimal_point != '.';
if (substitute)
{
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
if (substitute)
{
// get_string() hands the token to the SAX interface with '.'
token_buffer[decimal_point_position] = '.';
}
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
const char current_decimal_point = get_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = current_decimal_point;
}
}
/*! /*!
@brief contiguous fast path for scanning a number @brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and (as scan_number() does) and
returns the token type. On anything it does not fully recognize as a returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -10804,16 +10843,11 @@ scan_number_done:
} }
#endif #endif
// materialize the token exactly as scan_number() would, substituting the // materialize the token exactly as scan_number() would. reset() already
// locale decimal point so convert_number()'s strtof fallback stays valid. // cleared token_buffer, so append() fills it (assign() is avoided
// reset() already cleared token_buffer, so append() fills it (assign() is // because custom string_t types need not provide it)
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len); token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos) decimal_point_position = dot_index;
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1); ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1); position.chars_read_total += (len - 1);
@@ -11080,11 +11114,7 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once) /// return current string value (implicitly resets the token; useful only once)
string_t& get_string() string_t& get_string()
{ {
// translate decimal points from locale back to '.' (#4084) // a number token holds '.' regardless of the locale (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
return token_buffer; return token_buffer;
} }
@@ -11380,9 +11410,7 @@ scan_number_done:
number_unsigned_t value_unsigned = 0; number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0; number_float_t value_float = 0;
/// the decimal point /// the position of the decimal point in token_buffer
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
std::size_t decimal_point_position = std::string::npos; std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the /// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
@@ -14059,6 +14087,80 @@ class binary_reader
} }
} }
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*! /*!
@brief reads a definite-length CBOR byte array @brief reads a definite-length CBOR byte array
@@ -14303,7 +14405,7 @@ class binary_reader
if (top.is_object) if (top.is_object)
{ {
key.clear(); key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key))) if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
{ {
return false; return false;
} }
@@ -14804,6 +14906,98 @@ class binary_reader
} }
} }
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*! /*!
@brief reads a MessagePack byte array @brief reads a MessagePack byte array
@@ -14966,7 +15160,7 @@ class binary_reader
{ {
get(); get();
key.clear(); key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key))) if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
{ {
return false; return false;
} }
@@ -23931,8 +24125,7 @@ enum class error_handler_t
{ {
strict, ///< throw a type_error exception in case of invalid UTF-8 strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences ignore ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
}; };
template<typename BasicJsonType> template<typename BasicJsonType>
@@ -24903,47 +25096,6 @@ class serializer
break; break;
} }
case error_handler_t::keep:
{
// drop whatever the incomplete sequence left in
// the buffer (only copied if !EnsureAscii) and copy
// the ill-formed bytes from the input instead
bytes = bytes_after_last_accept;
if (undumped_chars > 0)
{
// the pending bytes of the incomplete sequence
// are ill-formed; the current byte may be OK for
// itself, so we would like to read it again
for (std::size_t j = i - undumped_chars; j < i; ++j)
{
string_buffer[bytes++] = s[j];
}
--i;
}
else
{
// the current byte cannot start any sequence
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
} }
@@ -24989,15 +25141,6 @@ class serializer
break; break;
} }
case error_handler_t::keep:
{
// write all accepted bytes
put_buffer(string_buffer, bytes_after_last_accept);
// copy the bytes of the incomplete sequence unchanged
put_string(s, s.size() - undumped_chars, s.size());
break;
}
case error_handler_t::replace: case error_handler_t::replace:
{ {
// write all accepted bytes // write all accepted bytes
@@ -32357,55 +32500,6 @@ std::string format_as(const NLOHMANN_BASIC_JSON_TPL& j)
return j.dump(); return j.dump();
} }
inline namespace literals
{
inline namespace json_literals
{
/// @brief user-defined string literal for JSON values
/// @sa https://json.nlohmann.me/api/basic_json/operator_literal_json/
JSON_HEDLEY_NON_NULL(1)
#if !defined(JSON_HEDLEY_GCC_VERSION) || JSON_HEDLEY_GCC_VERSION_CHECK(4,9,0)
inline nlohmann::json operator""_json(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json operator"" _json(const char* s, std::size_t n)
#endif
{
return nlohmann::json::parse(s, s + n);
}
#if defined(__cpp_char8_t)
JSON_HEDLEY_NON_NULL(1)
inline nlohmann::json operator""_json(const char8_t* s, std::size_t n)
{
return nlohmann::json::parse(reinterpret_cast<const char*>(s),
reinterpret_cast<const char*>(s) + n);
}
#endif
/// @brief user-defined string literal for JSON pointer
/// @sa https://json.nlohmann.me/api/basic_json/operator_literal_json_pointer/
JSON_HEDLEY_NON_NULL(1)
#if !defined(JSON_HEDLEY_GCC_VERSION) || JSON_HEDLEY_GCC_VERSION_CHECK(4,9,0)
inline nlohmann::json::json_pointer operator""_json_pointer(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json::json_pointer operator"" _json_pointer(const char* s, std::size_t n)
#endif
{
return nlohmann::json::json_pointer(std::string(s, n));
}
#if defined(__cpp_char8_t)
inline nlohmann::json::json_pointer operator""_json_pointer(const char8_t* s, std::size_t n)
{
return nlohmann::json::json_pointer(std::string(reinterpret_cast<const char*>(s), n));
}
#endif
} // namespace json_literals
} // namespace literals
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
/////////////////////// ///////////////////////
@@ -32540,17 +32634,6 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
} // namespace std } // namespace std
#if JSON_USE_GLOBAL_UDLS
#if !defined(JSON_HEDLEY_GCC_VERSION) || JSON_HEDLEY_GCC_VERSION_CHECK(4,9,0)
using nlohmann::literals::json_literals::operator""_json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator""_json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#else
// GCC 4.8 requires a space between "" and suffix
using nlohmann::literals::json_literals::operator"" _json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator"" _json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#endif
#endif
// #include <nlohmann/detail/macro_unscope.hpp> // #include <nlohmann/detail/macro_unscope.hpp>
// __ _____ _____ _____ // __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ // __| | __| | | | JSON for Modern C++
@@ -32579,7 +32662,6 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#undef JSON_INLINE_VARIABLE #undef JSON_INLINE_VARIABLE
#undef JSON_NO_UNIQUE_ADDRESS #undef JSON_NO_UNIQUE_ADDRESS
#undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_DISABLE_ENUM_SERIALIZATION
#undef JSON_USE_GLOBAL_UDLS
#ifndef JSON_TEST_KEEP_MACROS #ifndef JSON_TEST_KEEP_MACROS
#undef JSON_CATCH #undef JSON_CATCH
@@ -32772,4 +32854,97 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#pragma GCC diagnostic pop #pragma GCC diagnostic pop
#endif #endif
// The user-defined string literals are in a separate header, because their
// bodies instantiate the parser in every translation unit that includes them.
// Define JSON_NO_AUTOMATIC_UDLS to include <nlohmann/json_literals.hpp> only
// where needed.
#ifndef JSON_NO_AUTOMATIC_UDLS
// #include <nlohmann/json_literals.hpp>
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#ifndef INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#define INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#include <cstddef> // size_t
#include <string> // string
// #include <nlohmann/json.hpp>
// This header is included at the end of <nlohmann/json.hpp> unless
// JSON_NO_AUTOMATIC_UDLS is defined, and can be included on its own after that.
// Either way, the library's internal macros are no longer defined here (and the
// amalgamation inlines macro_scope.hpp only once), so only standard and public
// macros may be used below.
NLOHMANN_JSON_NAMESPACE_BEGIN
inline namespace literals
{
inline namespace json_literals
{
/// @brief user-defined string literal for JSON values
/// @sa https://json.nlohmann.me/api/operator_literal_json/
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
inline nlohmann::json operator""_json(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json operator"" _json(const char* s, std::size_t n)
#endif
{
return nlohmann::json::parse(s, s + n);
}
#if defined(__cpp_char8_t)
inline nlohmann::json operator""_json(const char8_t* s, std::size_t n)
{
return nlohmann::json::parse(reinterpret_cast<const char*>(s),
reinterpret_cast<const char*>(s) + n);
}
#endif
/// @brief user-defined string literal for JSON pointer
/// @sa https://json.nlohmann.me/api/operator_literal_json_pointer/
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
inline nlohmann::json::json_pointer operator""_json_pointer(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json::json_pointer operator"" _json_pointer(const char* s, std::size_t n)
#endif
{
return nlohmann::json::json_pointer(std::string(s, n));
}
#if defined(__cpp_char8_t)
inline nlohmann::json::json_pointer operator""_json_pointer(const char8_t* s, std::size_t n)
{
return nlohmann::json::json_pointer(std::string(reinterpret_cast<const char*>(s), n));
}
#endif
} // namespace json_literals
} // namespace literals
NLOHMANN_JSON_NAMESPACE_END
#if !defined(JSON_USE_GLOBAL_UDLS) || JSON_USE_GLOBAL_UDLS
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
using nlohmann::literals::json_literals::operator""_json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator""_json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#else
// GCC 4.8 requires a space between "" and suffix
using nlohmann::literals::json_literals::operator"" _json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator"" _json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#endif
#endif
#endif // INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#endif
#endif // INCLUDE_NLOHMANN_JSON_HPP_ #endif // INCLUDE_NLOHMANN_JSON_HPP_
+83
View File
@@ -0,0 +1,83 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#ifndef INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#define INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
#include <cstddef> // size_t
#include <string> // string
#include <nlohmann/json.hpp>
// This header is included at the end of <nlohmann/json.hpp> unless
// JSON_NO_AUTOMATIC_UDLS is defined, and can be included on its own after that.
// Either way, the library's internal macros are no longer defined here (and the
// amalgamation inlines macro_scope.hpp only once), so only standard and public
// macros may be used below.
NLOHMANN_JSON_NAMESPACE_BEGIN
inline namespace literals
{
inline namespace json_literals
{
/// @brief user-defined string literal for JSON values
/// @sa https://json.nlohmann.me/api/operator_literal_json/
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
inline nlohmann::json operator""_json(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json operator"" _json(const char* s, std::size_t n)
#endif
{
return nlohmann::json::parse(s, s + n);
}
#if defined(__cpp_char8_t)
inline nlohmann::json operator""_json(const char8_t* s, std::size_t n)
{
return nlohmann::json::parse(reinterpret_cast<const char*>(s),
reinterpret_cast<const char*>(s) + n);
}
#endif
/// @brief user-defined string literal for JSON pointer
/// @sa https://json.nlohmann.me/api/operator_literal_json_pointer/
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
inline nlohmann::json::json_pointer operator""_json_pointer(const char* s, std::size_t n)
#else
// GCC 4.8 requires a space between "" and suffix
inline nlohmann::json::json_pointer operator"" _json_pointer(const char* s, std::size_t n)
#endif
{
return nlohmann::json::json_pointer(std::string(s, n));
}
#if defined(__cpp_char8_t)
inline nlohmann::json::json_pointer operator""_json_pointer(const char8_t* s, std::size_t n)
{
return nlohmann::json::json_pointer(std::string(reinterpret_cast<const char*>(s), n));
}
#endif
} // namespace json_literals
} // namespace literals
NLOHMANN_JSON_NAMESPACE_END
#if !defined(JSON_USE_GLOBAL_UDLS) || JSON_USE_GLOBAL_UDLS
#if !defined(__GNUC__) || defined(__clang__) || __GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9)
using nlohmann::literals::json_literals::operator""_json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator""_json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#else
// GCC 4.8 requires a space between "" and suffix
using nlohmann::literals::json_literals::operator"" _json; // NOLINT(misc-unused-using-decls,google-global-names-in-headers)
using nlohmann::literals::json_literals::operator"" _json_pointer; //NOLINT(misc-unused-using-decls,google-global-names-in-headers)
#endif
#endif
#endif // INCLUDE_NLOHMANN_JSON_LITERALS_HPP_
+1
View File
@@ -18,6 +18,7 @@ module;
// See: https://github.com/nlohmann/json/issues/5103 // See: https://github.com/nlohmann/json/issues/5103
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
#include <nlohmann/json_literals.hpp>
export module nlohmann.json; export module nlohmann.json;
+43 -2
View File
@@ -1830,10 +1830,51 @@ TEST_CASE("CBOR")
SECTION("invalid string in map") SECTION("invalid string in map")
{ {
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01}), true, false).is_discarded()); CHECK(json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01}), true, false).is_discarded());
} }
SECTION("non-string key (see #2766 and #3381)")
{
// only text strings map to JSON object keys; any other key is
// rejected with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0xA1, 0x01, 0x01}, "an unsigned integer; last byte: 0x01"},
{{0xA1, 0x20, 0x01}, "a negative integer; last byte: 0x20"},
{{0xA1, 0x41, 0x61, 0x01}, "a byte string; last byte: 0x41"},
{{0xA1, 0x80, 0x01}, "an array; last byte: 0x80"},
{{0xA1, 0xA0, 0x01}, "a map; last byte: 0xA0"},
{{0xA1, 0xC0, 0x61, 0x61, 0x01}, "a tag; last byte: 0xC0"},
{{0xA1, 0xF4, 0x01}, "a boolean; last byte: 0xF4"},
{{0xA1, 0xF5, 0x01}, "a boolean; last byte: 0xF5"},
{{0xA1, 0xF6, 0x01}, "null; last byte: 0xF6"},
{{0xA1, 0xF7, 0x01}, "undefined; last byte: 0xF7"},
{{0xA1, 0xF9, 0x3C, 0x00, 0x01}, "a floating-point number; last byte: 0xF9"},
{{0xA1, 0xFA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFA"},
{{0xA1, 0xFB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFB"},
{{0xA1, 0xE0, 0x01}, "a simple value; last byte: 0xE0"},
{{0xA1, 0xF8, 0x20, 0x01}, "a simple value; last byte: 0xF8"},
// indefinite-length map
{{0xBF, 0x01, 0x01, 0xFF}, "an unsigned integer; last byte: 0x01"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_cbor(c.first, true, false).is_discarded());
}
// a key of major type 3 with a reserved length is still reported as
// a malformed string, and a missing key as the end of input
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)") SECTION("invalid UTF-8 in string (see #5529)")
{ {
// a two-character text string (major type 3) whose bytes are not // a two-character text string (major type 3) whose bytes are not
@@ -2284,7 +2325,7 @@ TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
SECTION("a break marker outside an indefinite-length string is not a string") SECTION("a break marker outside an indefinite-length string is not a string")
{ {
// 0xFF only closes a string that was opened; on its own it is not one // 0xFF only closes a string that was opened; on its own it is not one
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
} }
} }
+1 -1
View File
@@ -666,7 +666,7 @@ TEST_CASE("parse_float_fast declines what it cannot convert exactly")
// always safe: the caller then falls back to a slower, exact conversion. // always safe: the caller then falls back to a slower, exact conversion.
const auto fast = [](const std::string & s, double & out) const auto fast = [](const std::string & s, double & out)
{ {
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), '.', out); return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out);
}; };
double out = 0; double out = 0;
+210
View File
@@ -12,7 +12,12 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <array>
#include <clocale> #include <clocale>
#include <map>
#include <string>
#include <utility>
#include <vector>
struct ParserImpl final: public nlohmann::json_sax<json> struct ParserImpl final: public nlohmann::json_sax<json>
{ {
@@ -175,3 +180,208 @@ TEST_CASE("locale-dependent test (LC_NUMERIC=de_DE)")
MESSAGE("locale de_DE is not usable"); MESSAGE("locale de_DE is not usable");
} }
} }
namespace
{
// records the numbers of a flat array and switches LC_NUMERIC to the given
// locale once the array opens - after the lexer was constructed, but before
// any number in the array is lexed
struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
{
explicit LocaleSwitchingSax(const char* switch_to)
: locale_after_open(switch_to)
{}
bool null() override
{
return true;
}
bool boolean(bool /*val*/) override
{
return true;
}
bool number_integer(json::number_integer_t /*val*/) override
{
return true;
}
bool number_unsigned(json::number_unsigned_t /*val*/) override
{
return true;
}
bool number_float(json::number_float_t val, const json::string_t& s) override
{
values.push_back(val);
strings.push_back(s);
return true;
}
bool string(json::string_t& /*val*/) override
{
return true;
}
bool binary(json::binary_t& /*val*/) override
{
return true;
}
bool start_object(std::size_t /*val*/) override
{
return true;
}
bool key(json::string_t& /*val*/) override
{
return true;
}
bool end_object() override
{
return true;
}
bool start_array(std::size_t /*val*/) override
{
switched = std::setlocale(LC_NUMERIC, locale_after_open.c_str()) != nullptr;
return true;
}
bool end_array() override
{
return true;
}
bool parse_error(std::size_t /*val*/, const std::string& /*val*/, const nlohmann::detail::exception& /*val*/) override
{
return false;
}
std::string locale_after_open;
bool switched = false;
std::vector<json::number_float_t> values {}; // NOLINT(readability-redundant-member-init)
std::vector<json::string_t> strings {}; // NOLINT(readability-redundant-member-init)
};
} // namespace
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
{
// The numbers are chosen so that the conversion also takes the strtod
// fallback, which honors the locale that is current at conversion time:
// too many significant digits for Clinger's fast path, an underflow that
// std::from_chars rejects, and a plain value.
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
std::string text = "[";
for (const auto& n : numbers)
{
text += (text.size() == 1 ? "" : ",") + n;
}
text += "]";
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
// reference values, parsed without a locale switch
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
const json expected = json::parse(text);
const long_double_json expected_ld = long_double_json::parse(text);
const std::array<std::pair<const char*, const char*>, 2> transitions =
{
{
{"C", "de_DE"},
{"de_DE", "C"}
}
};
for (const auto& transition : transitions)
{
CAPTURE(transition.first);
CAPTURE(transition.second);
if (std::setlocale(LC_NUMERIC, transition.first) == nullptr)
{
MESSAGE("locale is not usable");
continue;
}
// SAX parsing
{
LocaleSwitchingSax sax(transition.second);
CHECK(json::sax_parse(text, &sax));
if (sax.switched)
{
CHECK(sax.values == expected.get<std::vector<json::number_float_t>>());
CHECK(sax.strings == numbers);
}
}
// DOM parsing with a callback
{
bool switched = false;
const auto cb = [&](int /*depth*/, json::parse_event_t event, json& /*parsed*/) noexcept
{
if (event == json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const json j = json::parse(text, cb);
if (switched)
{
CHECK(j == expected);
}
}
// a long double goes through std::strtold unless std::from_chars supports it
{
bool switched = false;
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
{
if (event == long_double_json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const long_double_json j = long_double_json::parse(text, cb);
if (switched)
{
CHECK(j == expected_ld);
}
}
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
TEST_CASE("locale with a multi-byte decimal point")
{
// Some locales use a decimal point that is not a single character, e.g.
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
// substituted in place for '.', so the strtod fallback stops early. The
// conversion must still terminate rather than retry forever.
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
bool tested = false;
for (const char* name : names)
{
if (std::setlocale(LC_NUMERIC, name) == nullptr)
{
continue;
}
const std::string decimal_point = std::localeconv()->decimal_point;
if (decimal_point.size() < 2)
{
continue;
}
CAPTURE(name);
tested = true;
// too many significant digits for Clinger's fast path, and an underflow
// that std::from_chars rejects: both reach the strtod fallback
json j;
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
CHECK(j.is_array());
CHECK(json::accept("3.14159265358979323846"));
// a value the locale-independent paths convert is not affected
CHECK(json::parse("12.5") == 12.5);
}
if (!tested)
{
MESSAGE("no locale with a multi-byte decimal point is usable");
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
+60 -1
View File
@@ -1551,10 +1551,69 @@ TEST_CASE("MessagePack")
SECTION("invalid string in map") SECTION("invalid string in map")
{ {
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found an integer; last byte: 0xFF", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01}), true, false).is_discarded()); CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01}), true, false).is_discarded());
} }
SECTION("non-string key (see #3381)")
{
// only strings map to JSON object keys; any other key is rejected
// with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0x81, 0xC0, 0x01}, "nil; last byte: 0xC0"},
{{0x81, 0xC2, 0x01}, "a boolean; last byte: 0xC2"},
{{0x81, 0xC3, 0x01}, "a boolean; last byte: 0xC3"},
{{0x81, 0xCA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a float; last byte: 0xCA"},
{{0x81, 0xCB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a float; last byte: 0xCB"},
{{0x81, 0xC4, 0x00, 0x01}, "a bin; last byte: 0xC4"},
{{0x81, 0xC5, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC5"},
{{0x81, 0xC6, 0x00, 0x00, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC6"},
{{0x81, 0xC7, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC7"},
{{0x81, 0xC8, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC8"},
{{0x81, 0xC9, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC9"},
{{0x81, 0xD4, 0x01, 0x00, 0x01}, "an ext; last byte: 0xD4"},
{{0x81, 0xD5, 0x01, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD5"},
{{0x81, 0xD6, 0x01, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD6"},
{{0x81, 0xD7, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD7"},
{{0x81, 0xD8, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD8"},
{{0x81, 0xCC, 0x01, 0x01}, "an integer; last byte: 0xCC"},
{{0x81, 0xCD, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCD"},
{{0x81, 0xCE, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCE"},
{{0x81, 0xCF, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCF"},
{{0x81, 0xD0, 0x01, 0x01}, "an integer; last byte: 0xD0"},
{{0x81, 0xD1, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD1"},
{{0x81, 0xD2, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD2"},
{{0x81, 0xD3, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD3"},
{{0x81, 0x00, 0x01}, "an integer; last byte: 0x00"},
{{0x81, 0x7F, 0x01}, "an integer; last byte: 0x7F"},
{{0x81, 0xE0, 0x01}, "an integer; last byte: 0xE0"},
{{0x81, 0x80, 0x01}, "a map; last byte: 0x80"},
{{0x81, 0x8F, 0x01}, "a map; last byte: 0x8F"},
{{0x81, 0xDE, 0x00, 0x00, 0x01}, "a map; last byte: 0xDE"},
{{0x81, 0xDF, 0x00, 0x00, 0x00, 0x00, 0x01}, "a map; last byte: 0xDF"},
{{0x81, 0x90, 0x01}, "an array; last byte: 0x90"},
{{0x81, 0x9F, 0x01}, "an array; last byte: 0x9F"},
{{0x81, 0xDC, 0x00, 0x00, 0x01}, "an array; last byte: 0xDC"},
{{0x81, 0xDD, 0x00, 0x00, 0x00, 0x00, 0x01}, "an array; last byte: 0xDD"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_msgpack(c.first, true, false).is_discarded());
}
json _;
// the unused byte 0xC1 is still reported as a malformed string
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xC1, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xC1", json::parse_error&);
// a missing key is still reported as the end of input
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)") SECTION("invalid UTF-8 in string (see #5529)")
{ {
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8 // a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
+97
View File
@@ -0,0 +1,97 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// This translation unit checks JSON_NO_AUTOMATIC_UDLS, which keeps
// <nlohmann/json.hpp> from including <nlohmann/json_literals.hpp> and thereby
// leaves out the user-defined string literals operator""_json and
// operator""_json_pointer (see #5294), and that including
// <nlohmann/json_literals.hpp> afterwards brings them back.
#define JSON_NO_AUTOMATIC_UDLS 1
#include "doctest_compatibility.h"
#include <cstddef>
#include <utility>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
// An argument type whose associated namespace is the library namespace, so
// argument-dependent lookup of a literal operator called by its function name
// also searches the inline namespaces nlohmann::literals::json_literals.
NLOHMANN_JSON_NAMESPACE_BEGIN
struct no_automatic_udls_probe
{
operator const char* () const // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
{
return "";
}
};
NLOHMANN_JSON_NAMESPACE_END
namespace
{
// The calls below use a dependent argument, so a literal operator that is not
// declared at all is a substitution failure rather than a hard error: lookup is
// deferred to the point of instantiation, where it considers the declarations
// visible from here (the global using-declarations of JSON_USE_GLOBAL_UDLS) plus
// argument-dependent lookup (the literals in the library namespace).
template<typename T>
using json_udl_t = decltype(operator""_json(std::declval<T>(), std::size_t()));
template<typename T>
using json_pointer_udl_t = decltype(operator""_json_pointer(std::declval<T>(), std::size_t()));
template<typename T>
using has_json_udl = nlohmann::detail::is_detected<json_udl_t, T>;
template<typename T>
using has_json_pointer_udl = nlohmann::detail::is_detected<json_pointer_udl_t, T>;
} // namespace
TEST_CASE("JSON_NO_AUTOMATIC_UDLS")
{
SECTION("literals are not declared")
{
// global namespace (JSON_USE_GLOBAL_UDLS defaults to 1)
CHECK_FALSE(has_json_udl<const char*>::value);
CHECK_FALSE(has_json_pointer_udl<const char*>::value);
// nlohmann::literals::json_literals
CHECK_FALSE(has_json_udl<nlohmann::no_automatic_udls_probe>::value);
CHECK_FALSE(has_json_pointer_udl<nlohmann::no_automatic_udls_probe>::value);
}
SECTION("the rest of the library keeps working")
{
const json j = json::parse(R"({"foo": {"bar": 42}})");
CHECK(j.dump() == R"({"foo":{"bar":42}})");
const json::json_pointer ptr("/foo/bar");
CHECK(j.at(ptr) == 42);
CHECK(j.contains(ptr));
}
}
// the literals can still be added where they are needed
#include <nlohmann/json_literals.hpp>
TEST_CASE("JSON_NO_AUTOMATIC_UDLS with <nlohmann/json_literals.hpp>")
{
SECTION("global namespace")
{
CHECK("[1,2]"_json == json({1, 2}));
CHECK("/a/0"_json_pointer == json::json_pointer("/a/0"));
}
SECTION("nlohmann::literals::json_literals")
{
using namespace nlohmann::literals::json_literals; // NOLINT(google-build-using-namespace)
CHECK(R"({"a":[42]})"_json.at("/a/0"_json_pointer) == 42);
}
}
+3 -3
View File
@@ -1018,7 +1018,7 @@ TEST_CASE("regression tests 1")
}; };
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x98", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an array; last byte: 0x98", json::parse_error&);
// related test case: nonempty UTF-8 string (indefinite length) // related test case: nonempty UTF-8 string (indefinite length)
std::vector<uint8_t> const vec1 {0x7f, 0x61, 0x61}; std::vector<uint8_t> const vec1 {0x7f, 0x61, 0x61};
@@ -1065,7 +1065,7 @@ TEST_CASE("regression tests 1")
}; };
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
// related test case: double-precision // related test case: double-precision
std::vector<uint8_t> const vec2 std::vector<uint8_t> const vec2
@@ -1077,7 +1077,7 @@ TEST_CASE("regression tests 1")
0x96, 0x96, 0xb4, 0xb4, 0xfa, 0x94, 0x94, 0x61, 0x96, 0x96, 0xb4, 0xb4, 0xfa, 0x94, 0x94, 0x61,
0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0xfb 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0xfb
}; };
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
} }
SECTION("issue #452 - Heap-buffer-overflow (OSS-Fuzz issue 585)") SECTION("issue #452 - Heap-buffer-overflow (OSS-Fuzz issue 585)")
-9
View File
@@ -766,15 +766,6 @@ TEST_CASE("regression tests 2")
CHECK(j == k); CHECK(j == k);
} }
SECTION("issue #4552 - UTF-8 invalid characters are not always ignored when dumping with error_handler_t::ignore")
{
json node;
node["test"] = "test\334\005";
CHECK(node.dump(-1, ' ', false, json::error_handler_t::ignore) == "{\"test\":\"test\\u0005\"}");
CHECK(node.dump(-1, ' ', false, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
CHECK(node.dump(-1, ' ', true, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
}
} }
TEST_CASE("regression test - parser callback must not lose a duplicate key's prior value") TEST_CASE("regression test - parser callback must not lose a duplicate key's prior value")
-37
View File
@@ -92,8 +92,6 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"äü\""); CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"äü\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"ä\xEF\xBF\xBDü\""); CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"ä\xEF\xBF\xBDü\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"\\u00e4\\ufffd\\u00fc\""); CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"\\u00e4\\ufffd\\u00fc\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"ä\xA9ü\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\\u00e4\xA9\\u00fc\"");
} }
SECTION("invalid character (regression guard for shared UTF-8 decoder, see #5529)") SECTION("invalid character (regression guard for shared UTF-8 decoder, see #5529)")
@@ -116,8 +114,6 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123\""); CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\""); CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd\""); CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xC2\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xC2\"");
} }
SECTION("unexpected character") SECTION("unexpected character")
@@ -130,39 +126,6 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123456\""); CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123456\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\x34\x35\x36\""); CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\x34\x35\x36\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd456\""); CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd456\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
}
SECTION("keep: valid characters are still escaped")
{
// an invalid byte followed by characters that must be escaped
const json j = "\xC2\"\\\n\xFF\x05";
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
}
SECTION("keep: truncated multibyte sequences")
{
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
}
SECTION("keep: long string with many invalid bytes")
{
// exceeds the internal string buffer several times
std::string input;
std::string expected = "\"";
for (int i = 0; i < 2000; ++i)
{
input += "\xFF\xE2\x82\n\xC3\xA4";
expected += "\xFF\xE2\x82\\n\xC3\xA4";
}
expected += "\"";
const json j = input;
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == expected);
} }
SECTION("U+FFFD Substitution of Maximal Subparts") SECTION("U+FFFD Substitution of Maximal Subparts")
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <algorithm>
#include <fstream> #include <fstream>
#include <sstream> #include <sstream>
#include <iostream> #include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2; static std::string s_replaced2;
static std::string s_replaced_ascii; static std::string s_replaced_ascii;
static std::string s_replaced2_ascii; static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case // dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore); s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace); s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected) if (success_expected)
{ {
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string // all dumps should agree on the string
CHECK(s_strict == s_ignored); CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced); CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
} }
else else
{ {
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character // check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos); CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
} }
// check that prefix and suffix are preserved // check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz"); CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc"); CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz"); CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
} }
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4); void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <algorithm>
#include <fstream> #include <fstream>
#include <sstream> #include <sstream>
#include <iostream> #include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2; static std::string s_replaced2;
static std::string s_replaced_ascii; static std::string s_replaced_ascii;
static std::string s_replaced2_ascii; static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case // dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore); s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace); s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected) if (success_expected)
{ {
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string // all dumps should agree on the string
CHECK(s_strict == s_ignored); CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced); CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
} }
else else
{ {
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character // check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos); CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
} }
// check that prefix and suffix are preserved // check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz"); CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc"); CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz"); CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
} }
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4); void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <algorithm>
#include <fstream> #include <fstream>
#include <sstream> #include <sstream>
#include <iostream> #include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2; static std::string s_replaced2;
static std::string s_replaced_ascii; static std::string s_replaced_ascii;
static std::string s_replaced2_ascii; static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case // dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore); s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace); s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected) if (success_expected)
{ {
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string // all dumps should agree on the string
CHECK(s_strict == s_ignored); CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced); CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
} }
else else
{ {
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character // check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos); CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
} }
// check that prefix and suffix are preserved // check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz"); CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc"); CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz"); CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
} }
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4); void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <algorithm>
#include <fstream> #include <fstream>
#include <sstream> #include <sstream>
#include <iostream> #include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2; static std::string s_replaced2;
static std::string s_replaced_ascii; static std::string s_replaced_ascii;
static std::string s_replaced2_ascii; static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case // dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore); s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore); s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace); s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace); s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected) if (success_expected)
{ {
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string // all dumps should agree on the string
CHECK(s_strict == s_ignored); CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced); CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
} }
else else
{ {
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character // check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos); CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
} }
// check that prefix and suffix are preserved // check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz"); CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc"); CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz"); CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
} }
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4); void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);