Compare commits

..
Author SHA1 Message Date
Niels Lohmann 15b0cc0561 Merge remote-tracking branch 'origin/develop' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:18:51 +02:00
Niels Lohmann 83302ff69d Merge branch 'develop' into json-view/02b-float-parser
Conflicts:
- number_parse.hpp: kept this branch's float parser, which replaces the
  Eisel-Lemire code that develop's side changed (#5750 made its digit
  counter unsigned; this parser has no such counter, and it compiles
  cleanly with GCC's -Wstrict-overflow=5).
- number_handling.md, template_parameters.md: kept this branch's
  description of the conversion and added develop's "Before version
  3.13.0" sentence.

Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:35:40 +02:00
Niels Lohmann 9d44e3f359 Merge branch 'develop' into json-view/02b-float-parser
Conflicted only in tests/src/unit-class_lexer.cpp, where develop's #5737
lint fix (CAPTURE(x); -> CAPTURE(x)) collided with this PR's rewrite of
the Eisel-Lemire float tests; kept the PR's new tests and applied the
lint-fixed CAPTURE style. single_include regenerated via make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:53:02 +02:00
Niels Lohmann 9d88ead578 Clarify that the strtold fallback substitutes the locale's decimal point
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:40:37 +02:00
Niels Lohmann 9c71689715 Convert long doubles under a multi-byte decimal point completely
The strtold fallback, which is left only for long double formats that
are not binary64 (x87, binary128), substituted the first byte of the
locale's decimal point for '.'. Under a locale whose decimal point is
longer than one byte, such as fa_IR.UTF-8 or ar_EG.UTF-8 (U+066B),
strtold stopped there and the value was truncated at the decimal point.
A longer decimal point is now put into a copy of the token.

The test "locale with a multi-byte decimal point" now compares the long
double values with those of the "C" locale; with x87 long doubles it
failed before.

Fixes #5660.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:11 +02:00
Niels Lohmann 44ec53c77b Convert float and double with the library's own correctly rounded parser
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple
arm64) are now converted by the library itself, correctly rounded and
independent of the locale and of the C and C++ libraries:

- The token is split into sign, significand w (at most 19 digits), and
  decimal exponent q, using the positions of the decimal point and the
  exponent that the scanners already recorded, so no character is
  classified again.
- Clinger's fast path where w and 10^|q| are exact.
- Eisel-Lemire otherwise, now templated for binary32 and binary64.
- For tokens with more than 19 digits whose w and w + 1 round differently,
  an exact big-integer comparison with the midpoint between the two
  candidates (the digit comparison of fast_float, simplified).

This replaces the separate token walks of Clinger's fast path and of
Eisel-Lemire, the significant-digit gate that avoided the former, and, for
float and double, std::from_chars and the locale-aware strtod. std::from_chars
and strtold remain only for other long double formats (x87, binary128,
double-double) and for types that are not IEEE-754. Values are bit-identical
to before wherever the previous conversion was correctly rounded; tokens
converted in a locale with a multi-byte decimal point are now also exact.
Overflow still gives out_of_range.406, underflow a signed zero.

convert_float() is the entry point for other parsers of JSON text: it
converts like the lexer, without allocation for binary32/binary64.

Tests: exact-bit tests for double and float (ties, subnormal and overflow
boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for
binary32, the round trips of 200,000 doubles and 100,000 floats without
declines, 508 generated hard cases with the expected bits of both formats
(float_hard_cases.hpp) through the converter and both scanners, and
JSON-level overflow/underflow checks for double and float. The locale tests
now check the values in a locale with a multi-byte decimal point.

Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float
credit now names the digit comparison.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:11 +02:00
25 changed files with 4714 additions and 2711 deletions
+6 -1
View File
@@ -72,4 +72,9 @@ build_script:
- cmake --build . --config "%configuration%" --parallel 2 - cmake --build . --config "%configuration%" --parallel 2
test_script: test_script:
- ctest -C "%configuration%" --parallel 2 --output-on-failure - if "%configuration%"=="Release" ctest -C "%configuration%" --parallel 2 --output-on-failure
# On Debug builds, skip test-unicode_all
# as it is extremely slow to run and cause
# occasional timeouts on AppVeyor.
# More info: https://github.com/nlohmann/json/pull/1570
- if "%configuration%"=="Debug" ctest --exclude-regex "test-unicode" -C "%configuration%" --parallel 2 --output-on-failure
+2 -2
View File
@@ -178,7 +178,7 @@ jobs:
- name: Build - name: Build
run: cmake --build build --parallel 10 run: cmake --build build --parallel 10
- name: Test - name: Test
run: cd build ; ctest -j 10 -C Debug --output-on-failure run: cd build ; ctest -j 10 -C Debug --exclude-regex "test-unicode" --output-on-failure
clang-cl-12: clang-cl-12:
runs-on: windows-2022 runs-on: windows-2022
@@ -195,7 +195,7 @@ jobs:
- name: Build - name: Build
run: cmake --build build --config Debug --parallel 10 run: cmake --build build --config Debug --parallel 10
- name: Test - name: Test
run: cd build ; ctest -j 10 -C Debug --output-on-failure run: cd build ; ctest -j 10 -C Debug --exclude-regex "test-unicode" --output-on-failure
ci_module_cpp20: ci_module_cpp20:
runs-on: windows-2022 runs-on: windows-2022
+1 -1
View File
@@ -1393,7 +1393,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I
- The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/) - The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/)
- The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). - The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
- The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0). - The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0).
- The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors - The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
<img align="right" src="https://git.fsfe.org/reuse/reuse-ci/raw/branch/master/reuse-horizontal.png" alt="REUSE Software"> <img align="right" src="https://git.fsfe.org/reuse/reuse-ci/raw/branch/master/reuse-horizontal.png" alt="REUSE Software">
+6 -7
View File
@@ -442,14 +442,13 @@ add_custom_target(ci_test_single_header
# Valgrind. # Valgrind.
############################################################################### ###############################################################################
# The Unicode test (~17M assertions) is too slow under Valgrind.
add_custom_target(ci_test_valgrind add_custom_target(ci_test_valgrind
COMMAND CXX=${GCC_TOOL} ${CMAKE_COMMAND} COMMAND CXX=${GCC_TOOL} ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja -DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_Valgrind=ON -DJSON_BuildTests=ON -DJSON_Valgrind=ON
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_valgrind -S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_valgrind
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_valgrind COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_valgrind
COMMAND cd ${PROJECT_BINARY_DIR}/build_valgrind && ${CMAKE_CTEST_COMMAND} -L valgrind --exclude-regex "test-unicode" --parallel ${N} --output-on-failure COMMAND cd ${PROJECT_BINARY_DIR}/build_valgrind && ${CMAKE_CTEST_COMMAND} -L valgrind --parallel ${N} --output-on-failure
COMMENT "Compile and test with Valgrind" COMMENT "Compile and test with Valgrind"
) )
@@ -774,7 +773,7 @@ foreach(COMPILER g++-4.8 g++-4.9 g++-5 g++-6 g++-7 g++-8 g++-9 g++-10 g++-11 cla
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_compiler_${COMPILER} -S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_compiler_${COMPILER}
${ADDITIONAL_FLAGS} ${ADDITIONAL_FLAGS}
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_compiler_${COMPILER} COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_compiler_${COMPILER}
COMMAND cd ${PROJECT_BINARY_DIR}/build_compiler_${COMPILER} && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure COMMAND cd ${PROJECT_BINARY_DIR}/build_compiler_${COMPILER} && ${CMAKE_CTEST_COMMAND} --parallel ${N} --exclude-regex "test-unicode" --output-on-failure
COMMENT "Compile and test with ${COMPILER}" COMMENT "Compile and test with ${COMPILER}"
) )
endif() endif()
@@ -788,7 +787,7 @@ add_custom_target(ci_test_compiler_default
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_compiler_default -S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_compiler_default
${ADDITIONAL_FLAGS} ${ADDITIONAL_FLAGS}
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_compiler_default --parallel ${N} COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_compiler_default --parallel ${N}
COMMAND cd ${PROJECT_BINARY_DIR}/build_compiler_default && ${CMAKE_CTEST_COMMAND} --parallel ${N} -LE git_required --output-on-failure COMMAND cd ${PROJECT_BINARY_DIR}/build_compiler_default && ${CMAKE_CTEST_COMMAND} --parallel ${N} --exclude-regex "test-unicode" -LE git_required --output-on-failure
COMMENT "Compile and test with default C++ compiler" COMMENT "Compile and test with default C++ compiler"
) )
@@ -826,7 +825,7 @@ add_custom_target(ci_icpc
-DJSON_BuildTests=ON -DJSON_FastTests=ON -DJSON_BuildTests=ON -DJSON_FastTests=ON
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_icpc -S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_icpc
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_icpc COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_icpc
COMMAND cd ${PROJECT_BINARY_DIR}/build_icpc && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure COMMAND cd ${PROJECT_BINARY_DIR}/build_icpc && ${CMAKE_CTEST_COMMAND} --parallel ${N} --exclude-regex "test-unicode" --output-on-failure
COMMENT "Compile and test with ICPC" COMMENT "Compile and test with ICPC"
) )
@@ -837,7 +836,7 @@ add_custom_target(ci_icpx
-DJSON_BuildTests=ON -DJSON_FastTests=ON -DJSON_BuildTests=ON -DJSON_FastTests=ON
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_icpx -S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_icpx
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_icpx COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_icpx
COMMAND cd ${PROJECT_BINARY_DIR}/build_icpx && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure COMMAND cd ${PROJECT_BINARY_DIR}/build_icpx && ${CMAKE_CTEST_COMMAND} --parallel ${N} --exclude-regex "test-unicode" --output-on-failure
COMMENT "Compile and test with ICPX (Intel oneAPI DPC++/C++)" COMMENT "Compile and test with ICPX (Intel oneAPI DPC++/C++)"
) )
@@ -873,7 +872,7 @@ add_custom_target(ci_nvhpc
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_nvhpc COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_nvhpc
# the pipes are escaped so the surrounding shell passes them to ctest verbatim # the pipes are escaped so the surrounding shell passes them to ctest verbatim
# instead of treating them as shell pipe operators # instead of treating them as shell pipe operators
COMMAND cd ${PROJECT_BINARY_DIR}/build_nvhpc && ${CMAKE_CTEST_COMMAND} --parallel ${N} --exclude-regex "test-comparison_cpp20\\|test-comparison_legacy_cpp20\\|test-constructor1_cpp11\\|test-deserialization_cpp20" --output-on-failure COMMAND cd ${PROJECT_BINARY_DIR}/build_nvhpc && ${CMAKE_CTEST_COMMAND} --parallel ${N} --exclude-regex "test-unicode\\|test-comparison_cpp20\\|test-comparison_legacy_cpp20\\|test-constructor1_cpp11\\|test-deserialization_cpp20" --output-on-failure
COMMENT "Compile and test with NVIDIA HPC SDK (nvc++)" COMMENT "Compile and test with NVIDIA HPC SDK (nvc++)"
) )
@@ -23,9 +23,10 @@ type to use.
## Template parameters ## Template parameters
`NumberFloatType` `NumberFloatType`
: the type to store floating-point numbers. Parsing and serialization are implemented in terms of : the type to store floating-point numbers. The parser converts `#!cpp float`, `#!cpp double`, and a
`#!cpp std::strtof`/`#!cpp std::strtod`/`#!cpp std::strtold` and `#!cpp std::snprintf`, so the type must be `#!cpp long double` that is IEEE 754 binary64 itself and other `#!cpp long double` formats with
`#!cpp float`, `#!cpp double`, or `#!cpp long double`. The `#!cpp std::from_chars` or `#!cpp std::strtold`, and serialization falls back to `#!cpp std::snprintf`, so the
type must be `#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
[binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`, [binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`,
because they have no encoding for `#!cpp long double`. See because they have no encoding for `#!cpp long double`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype). [Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype).
@@ -82,12 +82,13 @@ flowchart TD
- Numbers with a decimal digit or scientific notation are always stored as `#!c double`. - Numbers with a decimal digit or scientific notation are always stored as `#!c double`.
- The number types can be changed, see [Template number types](#template-number-types). - The number types can be changed, see [Template number types](#template-number-types).
- Integers are converted by the library's own digit parser. Floating-point numbers are converted with - The library converts integers and floating-point numbers itself, independent of the locale. Floating-point
[`std::from_chars`](https://en.cppreference.com/w/cpp/utility/from_chars) if the library is compiled with C++17 numbers are correctly rounded (to nearest, ties to even). Only a `#!c long double` that is not IEEE 754 binary64
and the standard library supports it, then with an exact fast path for `#!c double` values with few significant (e.g., the 80-bit x87 format) is converted with `#!cpp std::from_chars` where available, or else with
digits, and otherwise with the locale-aware [`std::strtold`](https://en.cppreference.com/w/cpp/string/byte/strtof). For that call, the library temporarily
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) (`std::strtof`/`std::strtold` for the replaces the `.` with the decimal point of the current locale (which may be longer than one byte, e.g., in
other floating-point types). Before version 3.13.0, the conversion was realized by `fa_IR.UTF-8`), so the result does not depend on the locale either. Changing the locale in another thread during
parsing is undefined behavior of the C library, though. Before version 3.13.0, the conversion was realized by
[`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul), [`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul),
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and `std::strtod`, respectively. [`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and `std::strtod`, respectively.
@@ -100,10 +101,10 @@ flowchart TD
### Number limits ### Number limits
- Any 64-bit signed or unsigned integer can be stored without loss of precision. - Any 64-bit signed or unsigned integer can be stored without loss of precision.
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion via - Numbers exceeding the limits of `#!c double` (i.e., numbers whose rounded value is not satisfying
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) are not satisfying
[`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception [`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing. [`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing. Numbers too
small for `#!c double` (such as `#!c 1E-400`) become zero, with the sign of the number.
- Floating-point numbers are rounded to the next number representable as `double`. For instance - Floating-point numbers are rounded to the next number representable as `double`. For instance
`#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18). `#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18).
This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`. This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`.
@@ -26,9 +26,9 @@ Requirements are split into two groups:
diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside
the library. Four violations are not caught at compile time at all: the library. Four violations are not caught at compile time at all:
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and can silently misparse - A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers
floating-point numbers, because the lexer may hand the buffer to `#!cpp std::strtod`, which reads up to the stored as a `#!cpp long double` that is not IEEE 754 binary64 (e.g., the 80-bit x87 format), because the lexer
terminating null character. hands the buffer to `#!cpp std::strtold`.
- A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation, - A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation,
and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance. and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance.
- The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion - The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion
@@ -537,9 +537,10 @@ therefore silently changes parse results rather than raising an error. See
`NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`: `NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`:
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::from_chars` or, as a fallback, with - The [parser](../parsing/index.md) converts number literals to `#!cpp float`, `#!cpp double`, and a
`#!cpp std::strtof`, `#!cpp std::strtod`, or `#!cpp std::strtold`; the library provides overloads for exactly these `#!cpp long double` that is IEEE 754 binary64 itself; other `#!cpp long double` formats are converted with
three types. `#!cpp std::from_chars` where available, or with `#!cpp std::strtold`. The library provides overloads for exactly
these three types.
- [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion - [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`). (`#!cpp float` is promoted to `#!cpp double`).
+1 -1
View File
@@ -20,4 +20,4 @@ The class contains a slightly modified version of the Grisu2 algorithm from Flor
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
+13 -10
View File
@@ -1044,9 +1044,11 @@ class lexer : public lexer_base<BasicJsonType>
token_type::parse_error otherwise token_type::parse_error otherwise
@note The scanner is independent of the current locale: token_buffer @note The scanner is independent of the current locale: token_buffer
always holds `.`. Only the std::strtod fallback of convert_number() always holds `.`. The conversion of float and double does not use
depends on the locale, and it looks up the decimal point right the locale either. Only the std::strtold fallback of
before converting (see detail::convert_float_locale_aware()). convert_number() for long double formats other than binary64
depends on it, and it looks up the decimal point right before
converting (see detail::convert_float_locale_aware()).
*/ */
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated. token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{ {
@@ -1059,7 +1061,7 @@ class lexer : public lexer_base<BasicJsonType>
// offset just past the last mantissa byte in token_buffer (i.e. the // offset just past the last mantissa byte in token_buffer (i.e. the
// index of 'e'/'E', or the whole token when there is no exponent). // index of 'e'/'E', or the whole token when there is no exponent).
// convert_number() uses it to count significant digits; npos means // convert_number() uses it to split the token; npos means
// "not seen an exponent yet" and is resolved at scan_number_done // "not seen an exponent yet" and is resolved at scan_number_done
std::size_t mantissa_end = std::string::npos; std::size_t mantissa_end = std::string::npos;
@@ -1389,8 +1391,8 @@ scan_number_done:
@param[in] mantissa_end offset just past the last mantissa byte in @param[in] mantissa_end offset just past the last mantissa byte in
token_buffer (the index of 'e'/'E', or token_buffer (the index of 'e'/'E', or
token_buffer.size() when there is no exponent); token_buffer.size() when there is no exponent);
used to skip Clinger's fast path when it cannot with decimal_point_position, it locates the parts
possibly succeed - see detail::mantissa_fits_clinger() of a float token without scanning it again
*/ */
token_type convert_number(token_type number_type, std::size_t mantissa_end) token_type convert_number(token_type number_type, std::size_t mantissa_end)
{ {
@@ -1444,10 +1446,11 @@ scan_number_done:
} }
// this code is reached if we parse a floating-point number or if an // this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars // integer conversion above overflowed. float and double (and long
// (Eisel-Lemire, locale-independent, correctly rounded) when available; // double where it is binary64) are converted by the library itself,
// otherwise the exact Clinger fast path (double only); otherwise the // correctly rounded and independent of the locale; other long double
// locale-aware strtof/strtod/strtold. // formats use std::from_chars when available, otherwise the
// locale-aware strtold.
if (convert_float_fast(num_begin, num_end, decimal_point_position, mantissa_end, value_float)) if (convert_float_fast(num_begin, num_end, decimal_point_position, mantissa_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+3
View File
@@ -134,6 +134,9 @@ json_test_set_test_options(test-disabled_exceptions
#$<$<CXX_COMPILER_ID:MSVC>:/EH> #$<$<CXX_COMPILER_ID:MSVC>:/EH>
) )
# raise timeout of expensive Unicode test
json_test_set_test_options(test-unicode4 TEST_PROPERTIES TIMEOUT 3000)
# only the #972 regression test needs thirdparty/fifo_map on its include path # only the #972 regression test needs thirdparty/fifo_map on its include path
json_test_set_test_options(test-regression1 LINK_LIBRARIES fifo_map_include) json_test_set_test_options(test-regression1 LINK_LIBRARIES fifo_map_include)
+599
View File
@@ -0,0 +1,599 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstdint> // uint32_t, uint64_t
// Number tokens that are hard to round correctly, with the IEEE-754 binary64
// and binary32 bits of their correctly rounded values (ties to even; infinity
// for an overflow, a signed zero for an underflow).
//
// For doubles and floats around 0, the smallest normal number, 1, 2^24, 2^53,
// 0.1, and the largest finite number, and for random ones, the exact midpoint
// m to the next number gives: m, m with one unit more and less in the last
// digit, m with "01" and "0...01" appended, m with trailing zeros, and m cut
// after 17 to 30 digits (rounded down and up, so that the rounding is decided
// after the 19th digit), in fixed and exponent notation, 30% of them negative.
// Tokens longer than 80 characters are left out, except for four of 700 digits
// and more. Zeros, underflow, overflow, huge exponents, and integers beyond 64
// bits complete the set. Of the 508 tokens, 134 (as double) and 150 (as
// float) need the exact comparison with the midpoint (detail::digit_comparison()).
//
// The expected bits were computed with exact rational arithmetic in Python
// (fractions.Fraction) and cross-checked with Python's float(); strtod_l and
// strtof_l of Apple's libc and of glibc agree. Generated by
// compact_hard_cases.py 5 (with hard_cases.py), see the pull request that
// added this file.
namespace float_hard_cases
{
struct hard_case
{
const char* token;
std::uint64_t bits64;
std::uint32_t bits32;
};
inline const std::array<hard_case, 508>& cases()
{
static const std::array<hard_case, 508> table =
{
{
{"-2.4703282292062327e-324", 0x8000000000000000u, 0x80000000u},
{"24703282292062328e-340", 0x0000000000000001u, 0x00000000u},
{"247032822920623272e-341", 0x0000000000000000u, 0x00000000u},
{"-0.2470328229206232721e-323", 0x8000000000000001u, 0x80000000u},
{"-0.24703282292062327208e-323", 0x8000000000000000u, 0x80000000u},
{"-2.4703282292062327209e-324", 0x8000000000000001u, 0x80000000u},
{"2.47032822920623272088e-324", 0x0000000000000000u, 0x00000000u},
{"247032822920623272089e-344", 0x0000000000000001u, 0x00000000u},
{"-247032822920623272088284396434e-353", 0x8000000000000000u, 0x80000000u},
{"0.247032822920623272088284396435e-323", 0x0000000000000001u, 0x00000000u},
{"-74109846876186981e-340", 0x8000000000000001u, 0x80000000u},
{"0.74109846876186982e-323", 0x0000000000000002u, 0x00000000u},
{"-0.7410984687618698162e-323", 0x8000000000000001u, 0x80000000u},
{"-7.410984687618698163e-324", 0x8000000000000002u, 0x80000000u},
{"7.4109846876186981626e-324", 0x0000000000000001u, 0x00000000u},
{"-74109846876186981627e-343", 0x8000000000000002u, 0x80000000u},
{"-741098468761869816264e-344", 0x8000000000000001u, 0x80000000u},
{"0.741098468761869816265e-323", 0x0000000000000002u, 0x00000000u},
{"0.741098468761869816264853189302e-323", 0x0000000000000001u, 0x00000000u},
{"-7.41098468761869816264853189303e-324", 0x8000000000000002u, 0x80000000u},
{"0.22250738585072006e-307", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"2.2250738585072007e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"2.225073858507200641e-308", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"-2225073858507200642e-326", 0x800FFFFFFFFFFFFFu, 0x80000000u},
{"22250738585072006419e-327", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"0.2225073858507200642e-307", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"0.222507385850720064199e-307", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"2.225073858507200642e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"-2.22507385850720064199176395546e-308", 0x800FFFFFFFFFFFFEu, 0x80000000u},
{"222507385850720064199176395547e-337", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"-2.2250738585072011e-308", 0x800FFFFFFFFFFFFFu, 0x80000000u},
{"-22250738585072012e-324", 0x8010000000000000u, 0x80000000u},
{"-2225073858507201136e-326", 0x800FFFFFFFFFFFFFu, 0x80000000u},
{"0.2225073858507201137e-307", 0x0010000000000000u, 0x00000000u},
{"0.2225073858507201136e-307", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"-2.2250738585072011361e-308", 0x8010000000000000u, 0x80000000u},
{"2.22507385850720113605e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"222507385850720113606e-328", 0x0010000000000000u, 0x00000000u},
{"22250738585072011360574097967e-336", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"0.222507385850720113605740979671e-307", 0x0010000000000000u, 0x00000000u},
{"22250738585072016e-324", 0x0010000000000000u, 0x00000000u},
{"0.22250738585072017e-307", 0x0010000000000001u, 0x00000000u},
{"0.222507385850720163e-307", 0x0010000000000000u, 0x00000000u},
{"2.225073858507201631e-308", 0x0010000000000001u, 0x00000000u},
{"-2.2250738585072016301e-308", 0x8010000000000000u, 0x80000000u},
{"22250738585072016302e-327", 0x0010000000000001u, 0x00000000u},
{"-222507385850720163012e-328", 0x8010000000000000u, 0x80000000u},
{"0.222507385850720163013e-307", 0x0010000000000001u, 0x00000000u},
{"0.222507385850720163012305563795e-307", 0x0010000000000000u, 0x00000000u},
{"-2.22507385850720163012305563796e-308", 0x8010000000000001u, 0x80000000u},
{"0.17976931348623156E+309", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"1.7976931348623157e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"1.797693134862315608e308", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"-1797693134862315609e290", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u},
{"-17976931348623156083e289", 0xFFEFFFFFFFFFFFFEu, 0xFF800000u},
{"-0.17976931348623156084E+309", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u},
{"0.179769313486231560835E+309", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"-1.79769313486231560836e308", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u},
{"1.79769313486231560835325876058e308", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"179769313486231560835325876059e279", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"1.7976931348623158e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"17976931348623159e292", 0x7FF0000000000000u, 0x7F800000u},
{"1797693134862315807e290", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"0.1797693134862315808E+309", 0x7FF0000000000000u, 0x7F800000u},
{"0.17976931348623158079E+309", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"-1.797693134862315808e308", 0xFFF0000000000000u, 0xFF800000u},
{"1.79769313486231580793e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"179769313486231580794e288", 0x7FF0000000000000u, 0x7F800000u},
{"179769313486231580793728971405e279", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"-0.179769313486231580793728971406E+309", 0xFFF0000000000000u, 0xFF800000u},
{"100000000000000011102230246251565404236316680908203125e-53", 0x3FF0000000000000u, 0x3F800000u},
{"-1.00000000000000011102230246251565404236316680908203126", 0xBFF0000000000001u, 0xBF800000u},
{"1.00000000000000011102230246251565404236316680908203124e0", 0x3FF0000000000000u, 0x3F800000u},
{"10000000000000001110223024625156540423631668090820312501e-55", 0x3FF0000000000001u, 0x3F800000u},
{"1.00000000000000011102230246251565404236316680908203125000000000000000000001", 0x3FF0000000000001u, 0x3F800000u},
{"10000000000000001e-16", 0x3FF0000000000000u, 0x3F800000u},
{"1.0000000000000002", 0x3FF0000000000001u, 0x3F800000u},
{"1.000000000000000111", 0x3FF0000000000000u, 0x3F800000u},
{"1.000000000000000112e0", 0x3FF0000000000001u, 0x3F800000u},
{"1.000000000000000111e0", 0x3FF0000000000000u, 0x3F800000u},
{"-10000000000000001111e-19", 0xBFF0000000000001u, 0xBF800000u},
{"-100000000000000011102e-20", 0xBFF0000000000000u, 0xBF800000u},
{"-1.00000000000000011103", 0xBFF0000000000001u, 0xBF800000u},
{"1.00000000000000011102230246251", 0x3FF0000000000000u, 0x3F800000u},
{"1.00000000000000011102230246252e0", 0x3FF0000000000001u, 0x3F800000u},
{"-0.999999999999999944488848768742172978818416595458984375", 0xBFF0000000000000u, 0xBF800000u},
{"-9.99999999999999944488848768742172978818416595458984376e-1", 0xBFF0000000000000u, 0xBF800000u},
{"999999999999999944488848768742172978818416595458984374e-54", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"0.99999999999999994448884876874217297881841659545898437501", 0x3FF0000000000000u, 0x3F800000u},
{"9.99999999999999944488848768742172978818416595458984375000000000000000000001e-1", 0x3FF0000000000000u, 0x3F800000u},
{"-0.99999999999999994", 0xBFEFFFFFFFFFFFFFu, 0xBF800000u},
{"9.9999999999999995e-1", 0x3FF0000000000000u, 0x3F800000u},
{"9.999999999999999444e-1", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"9999999999999999445e-19", 0x3FF0000000000000u, 0x3F800000u},
{"99999999999999994448e-20", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"-0.99999999999999994449", 0xBFF0000000000000u, 0xBF800000u},
{"0.999999999999999944488", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"-9.99999999999999944489e-1", 0xBFF0000000000000u, 0xBF800000u},
{"9.99999999999999944488848768742e-1", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"999999999999999944488848768743e-30", 0x3FF0000000000000u, 0x3F800000u},
{"-9.007199254740993e15", 0xC340000000000000u, 0xDA000000u},
{"9007199254740994e0", 0x4340000000000001u, 0x5A000000u},
{"9007199254740992", 0x4340000000000000u, 0x5A000000u},
{"9.00719925474099301e15", 0x4340000000000001u, 0x5A000000u},
{"9007199254740993000000000000000000001e-21", 0x4340000000000001u, 0x5A000000u},
{"-9007199254740993.000000000000000000000000000000", 0xC340000000000000u, 0xDA000000u},
{"90071992547409915e-1", 0x4340000000000000u, 0x5A000000u},
{"-9007199254740991.6", 0xC340000000000000u, 0xDA000000u},
{"9.0071992547409914e15", 0x433FFFFFFFFFFFFFu, 0x5A000000u},
{"-9007199254740991501e-3", 0xC340000000000000u, 0xDA000000u},
{"9007199254740991.5000000000000000000001", 0x4340000000000000u, 0x5A000000u},
{"9.0071992547409915000000000000000000000000000000e15", 0x4340000000000000u, 0x5A000000u},
{"0.100000000000000012490009027033011079765856266021728515625", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"1.00000000000000012490009027033011079765856266021728515626e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"100000000000000012490009027033011079765856266021728515624e-57", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"0.10000000000000001249000902703301107976585626602172851562501", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"0.10000000000000001", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"1.0000000000000002e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"-1.000000000000000124e-1", 0xBFB999999999999Au, 0xBDCCCCCDu},
{"1000000000000000125e-19", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"-10000000000000001249e-20", 0xBFB999999999999Au, 0xBDCCCCCDu},
{"0.1000000000000000125", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"0.10000000000000001249", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"1.00000000000000012491e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"1.00000000000000012490009027033e-1", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"100000000000000012490009027034e-30", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"2.45134755833537796875e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"-245134755833537796876e-6", 0xC2EBDE5C4164D83Au, 0xD75EF2E2u},
{"-245134755833537.796874", 0xC2EBDE5C4164D839u, 0xD75EF2E2u},
{"2.4513475583353779687501e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"245134755833537796875000000000000000000001e-27", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"245134755833537.796875000000000000000000000000000000", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"2.4513475583353779e14", 0x42EBDE5C4164D839u, 0x575EF2E2u},
{"2451347558335378e-1", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"2451347558335377968e-4", 0x42EBDE5C4164D839u, 0x575EF2E2u},
{"245134755833537.7969", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"245134755833537.79687", 0x42EBDE5C4164D839u, 0x575EF2E2u},
{"2.4513475583353779688e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"181510327827821147441864013671875e-23", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"-1815103278.27821147441864013671876", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"1.81510327827821147441864013671874e9", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"18151032782782114744186401367187501e-25", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"1815103278.27821147441864013671875000000000000000000001", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"-1.81510327827821147441864013671875000000000000000000000000000000e9", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"18151032782782114e-7", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"1815103278.2782115", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"1815103278.278211474", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"-1.815103278278211475e9", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"1.8151032782782114744e9", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"-18151032782782114745e-10", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"181510327827821147441e-11", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"1815103278.27821147442", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"1815103278.27821147441864013671", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"1.81510327827821147441864013672e9", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"3809325632181785344", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3.809325632181785345e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3809325632181785343e0", 0x43CA6EB8BD69FE29u, 0x5E5375C6u},
{"3809325632181785344.01", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3.809325632181785344000000000000000000001e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3809325632181785344000000000000000000000000000000e-30", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3809325632181785300", 0x43CA6EB8BD69FE29u, 0x5E5375C6u},
{"3.8093256321817854e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"4.046966549916366943359375e12", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916366943359376e-12", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916.366943359374", 0x428D7210076CE2EFu, 0x546B9080u},
{"4.04696654991636694335937501e12", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916366943359375000000000000000000001e-33", 0x428D7210076CE2F0u, 0x546B9080u},
{"-4046966549916.366943359375000000000000000000000000000000", 0xC28D7210076CE2F0u, 0xD46B9080u},
{"4.0469665499163669e12", 0x428D7210076CE2EFu, 0x546B9080u},
{"4046966549916367e-3", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916366943e-6", 0x428D7210076CE2EFu, 0x546B9080u},
{"4046966549916.366944", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916.3669433", 0x428D7210076CE2EFu, 0x546B9080u},
{"4.0469665499163669434e12", 0x428D7210076CE2F0u, 0x546B9080u},
{"-4.04696654991636694335e12", 0xC28D7210076CE2EFu, 0xD46B9080u},
{"404696654991636694336e-8", 0x428D7210076CE2F0u, 0x546B9080u},
{"28093802557000874e154", 0x63529C3B77330BDBu, 0x7F800000u},
{"0.28093802557000875E+171", 0x63529C3B77330BDCu, 0x7F800000u},
{"0.2809380255700087447E+171", 0x63529C3B77330BDBu, 0x7F800000u},
{"2.809380255700087448e170", 0x63529C3B77330BDCu, 0x7F800000u},
{"2.8093802557000874472e170", 0x63529C3B77330BDBu, 0x7F800000u},
{"28093802557000874473e151", 0x63529C3B77330BDCu, 0x7F800000u},
{"280938025570008744728e150", 0x63529C3B77330BDBu, 0x7F800000u},
{"-0.280938025570008744729E+171", 0xE3529C3B77330BDCu, 0xFF800000u},
{"-0.280938025570008744728403667979E+171", 0xE3529C3B77330BDBu, 0xFF800000u},
{"2.8093802557000874472840366798e170", 0x63529C3B77330BDCu, 0x7F800000u},
{"0.39523280297734525e-154", 0x1FE0F51BF17FD374u, 0x00000000u},
{"-3.9523280297734526e-155", 0x9FE0F51BF17FD375u, 0x80000000u},
{"3.952328029773452547e-155", 0x1FE0F51BF17FD374u, 0x00000000u},
{"3952328029773452548e-173", 0x1FE0F51BF17FD375u, 0x00000000u},
{"-39523280297734525478e-174", 0x9FE0F51BF17FD374u, 0x80000000u},
{"0.39523280297734525479e-154", 0x1FE0F51BF17FD375u, 0x00000000u},
{"-0.395232802977345254787e-154", 0x9FE0F51BF17FD374u, 0x80000000u},
{"-3.95232802977345254788e-155", 0x9FE0F51BF17FD375u, 0x80000000u},
{"-3.95232802977345254787245825501e-155", 0x9FE0F51BF17FD374u, 0x80000000u},
{"395232802977345254787245825502e-184", 0x1FE0F51BF17FD375u, 0x00000000u},
{"-1.0790205420931879e-276", 0x86A3209CA6233255u, 0x80000000u},
{"-1079020542093188e-291", 0x86A3209CA6233256u, 0x80000000u},
{"-1079020542093187947e-294", 0x86A3209CA6233255u, 0x80000000u},
{"0.1079020542093187948e-275", 0x06A3209CA6233256u, 0x00000000u},
{"0.1079020542093187947e-275", 0x06A3209CA6233255u, 0x00000000u},
{"1.0790205420931879471e-276", 0x06A3209CA6233256u, 0x00000000u},
{"1.07902054209318794701e-276", 0x06A3209CA6233255u, 0x00000000u},
{"-107902054209318794702e-296", 0x86A3209CA6233256u, 0x80000000u},
{"107902054209318794701153285302e-305", 0x06A3209CA6233255u, 0x00000000u},
{"-0.107902054209318794701153285303e-275", 0x86A3209CA6233256u, 0x80000000u},
{"58530471071351308e-228", 0x1413B446E6A16A3Bu, 0x00000000u},
{"-0.58530471071351309e-211", 0x9413B446E6A16A3Cu, 0x80000000u},
{"0.5853047107135130893e-211", 0x1413B446E6A16A3Bu, 0x00000000u},
{"-5.853047107135130894e-212", 0x9413B446E6A16A3Cu, 0x80000000u},
{"5.853047107135130893e-212", 0x1413B446E6A16A3Bu, 0x00000000u},
{"58530471071351308931e-231", 0x1413B446E6A16A3Cu, 0x00000000u},
{"-585304710713513089304e-232", 0x9413B446E6A16A3Bu, 0x80000000u},
{"0.585304710713513089305e-211", 0x1413B446E6A16A3Cu, 0x00000000u},
{"-0.585304710713513089304248824438e-211", 0x9413B446E6A16A3Bu, 0x80000000u},
{"5.85304710713513089304248824439e-212", 0x1413B446E6A16A3Cu, 0x00000000u},
{"0.19334214893983531e-78", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"1.9334214893983532e-79", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"1.933421489398353102e-79", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"1933421489398353103e-97", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"19334214893983531023e-98", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"0.19334214893983531024e-78", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"0.193342148939835310231e-78", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"1.93342148939835310232e-79", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"-1.93342148939835310231359014704e-79", 0xAF96ECBF1CFB10F6u, 0x80000000u},
{"193342148939835310231359014705e-108", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"2.9873358928024455e227", 0x6F2938807814E8A2u, 0x7F800000u},
{"29873358928024456e211", 0x6F2938807814E8A3u, 0x7F800000u},
{"298733589280244551e210", 0x6F2938807814E8A2u, 0x7F800000u},
{"0.2987335892802445511E+228", 0x6F2938807814E8A3u, 0x7F800000u},
{"0.29873358928024455109E+228", 0x6F2938807814E8A2u, 0x7F800000u},
{"2.987335892802445511e227", 0x6F2938807814E8A3u, 0x7F800000u},
{"-2.98733589280244551098e227", 0xEF2938807814E8A2u, 0xFF800000u},
{"-298733589280244551099e207", 0xEF2938807814E8A3u, 0xFF800000u},
{"298733589280244551098081559931e198", 0x6F2938807814E8A2u, 0x7F800000u},
{"0.298733589280244551098081559932E+228", 0x6F2938807814E8A3u, 0x7F800000u},
{"7.0064923216240853e-46", 0x3690000000000000u, 0x00000000u},
{"70064923216240854e-62", 0x3690000000000000u, 0x00000001u},
{"7006492321624085354e-64", 0x3690000000000000u, 0x00000000u},
{"-0.7006492321624085355e-45", 0xB690000000000000u, 0x80000001u},
{"0.70064923216240853546e-45", 0x3690000000000000u, 0x00000000u},
{"7.0064923216240853547e-46", 0x3690000000000000u, 0x00000001u},
{"7.00649232162408535461e-46", 0x3690000000000000u, 0x00000000u},
{"-700649232162408535462e-66", 0xB690000000000000u, 0x80000001u},
{"-700649232162408535461864791644e-75", 0xB690000000000000u, 0x80000000u},
{"0.700649232162408535461864791645e-45", 0x3690000000000000u, 0x00000001u},
{"21019476964872256e-61", 0x36A8000000000000u, 0x00000001u},
{"0.21019476964872257e-44", 0x36A8000000000000u, 0x00000002u},
{"-0.2101947696487225606e-44", 0xB6A8000000000000u, 0x80000001u},
{"-2.101947696487225607e-45", 0xB6A8000000000000u, 0x80000002u},
{"2.1019476964872256063e-45", 0x36A8000000000000u, 0x00000001u},
{"21019476964872256064e-64", 0x36A8000000000000u, 0x00000002u},
{"-210194769648722560638e-65", 0xB6A8000000000000u, 0x80000001u},
{"-0.210194769648722560639e-44", 0xB6A8000000000000u, 0x80000002u},
{"-0.210194769648722560638559437493e-44", 0xB6A8000000000000u, 0x80000001u},
{"2.10194769648722560638559437494e-45", 0x36A8000000000000u, 0x00000002u},
{"0.11754941406275178e-37", 0x380FFFFFA0000000u, 0x007FFFFEu},
{"1.1754941406275179e-38", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"-1.175494140627517859e-38", 0xB80FFFFFA0000000u, 0x807FFFFEu},
{"117549414062751786e-55", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"-11754941406275178592e-57", 0xB80FFFFFA0000000u, 0x807FFFFEu},
{"0.11754941406275178593e-37", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"0.117549414062751785924e-37", 0x380FFFFFA0000000u, 0x007FFFFEu},
{"1.17549414062751785925e-38", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"-1.17549414062751785924617589866e-38", 0xB80FFFFFA0000000u, 0x807FFFFEu},
{"117549414062751785924617589867e-67", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"1.1754942807573642e-38", 0x380FFFFFDFFFFFFFu, 0x007FFFFFu},
{"11754942807573643e-54", 0x380FFFFFE0000000u, 0x00800000u},
{"1175494280757364291e-56", 0x380FFFFFE0000000u, 0x007FFFFFu},
{"0.1175494280757364292e-37", 0x380FFFFFE0000000u, 0x00800000u},
{"0.11754942807573642917e-37", 0x380FFFFFE0000000u, 0x007FFFFFu},
{"-1.1754942807573642918e-38", 0xB80FFFFFE0000000u, 0x80800000u},
{"-1.17549428075736429172e-38", 0xB80FFFFFE0000000u, 0x807FFFFFu},
{"-117549428075736429173e-58", 0xB80FFFFFE0000000u, 0x80800000u},
{"117549428075736429172788299103e-67", 0x380FFFFFE0000000u, 0x007FFFFFu},
{"0.117549428075736429172788299104e-37", 0x380FFFFFE0000000u, 0x00800000u},
{"11754944208872107e-54", 0x3810000010000000u, 0x00800000u},
{"-0.11754944208872108e-37", 0xB810000010000000u, 0x80800001u},
{"0.1175494420887210724e-37", 0x3810000010000000u, 0x00800000u},
{"1.175494420887210725e-38", 0x3810000010000000u, 0x00800001u},
{"1.1754944208872107242e-38", 0x3810000010000000u, 0x00800000u},
{"-11754944208872107243e-57", 0xB810000010000000u, 0x80800001u},
{"-11754944208872107242e-57", 0xB810000010000000u, 0x80800000u},
{"0.117549442088721072421e-37", 0x3810000010000000u, 0x00800001u},
{"0.11754944208872107242095900834e-37", 0x3810000010000000u, 0x00800000u},
{"-1.17549442088721072420959008341e-38", 0xB810000010000000u, 0x80800001u},
{"340282336497324057985868971510891282432", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"-3.40282336497324057985868971510891282433e38", 0xC7EFFFFFD0000000u, 0xFF7FFFFFu},
{"340282336497324057985868971510891282431e0", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324057985868971510891282432.01", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.40282336497324057985868971510891282432000000000000000000001e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"340282336497324057985868971510891282432000000000000000000000000000000e-30", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324050000000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"3.4028233649732406e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.402823364973240579e38", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324058e21", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"34028233649732405798e19", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324057990000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"340282336497324057985000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"3.40282336497324057986e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.4028233649732405798586897151e38", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324057985868971511e9", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.40282356779733661637539395458142568448e38", 0x47EFFFFFF0000000u, 0x7F800000u},
{"-340282356779733661637539395458142568449e0", 0xC7EFFFFFF0000000u, 0xFF800000u},
{"340282356779733661637539395458142568447", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"3.4028235677973366163753939545814256844801e38", 0x47EFFFFFF0000000u, 0x7F800000u},
{"340282356779733661637539395458142568448000000000000000000001e-21", 0x47EFFFFFF0000000u, 0x7F800000u},
{"340282356779733661637539395458142568448.000000000000000000000000000000", 0x47EFFFFFF0000000u, 0x7F800000u},
{"3.4028235677973366e38", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"-34028235677973367e22", 0xC7EFFFFFF0000000u, 0xFF800000u},
{"3402823567797336616e20", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"340282356779733661700000000000000000000", 0x47EFFFFFF0000000u, 0x7F800000u},
{"340282356779733661630000000000000000000", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"3.4028235677973366164e38", 0x47EFFFFFF0000000u, 0x7F800000u},
{"3.40282356779733661637e38", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"340282356779733661638e18", 0x47EFFFFFF0000000u, 0x7F800000u},
{"-340282356779733661637539395458e9", 0xC7EFFFFFF0000000u, 0xFF7FFFFFu},
{"340282356779733661637539395459000000000", 0x47EFFFFFF0000000u, 0x7F800000u},
{"-1000000059604644775390625e-24", 0xBFF0000010000000u, 0xBF800000u},
{"-1.000000059604644775390626", 0xBFF0000010000000u, 0xBF800001u},
{"-1.000000059604644775390624e0", 0xBFF0000010000000u, 0xBF800000u},
{"100000005960464477539062501e-26", 0x3FF0000010000000u, 0x3F800001u},
{"1.000000059604644775390625000000000000000000001", 0x3FF0000010000000u, 0x3F800001u},
{"1.000000059604644775390625000000000000000000000000000000e0", 0x3FF0000010000000u, 0x3F800000u},
{"-10000000596046447e-16", 0xBFF0000010000000u, 0xBF800000u},
{"1.0000000596046448", 0x3FF0000010000000u, 0x3F800001u},
{"1.000000059604644775", 0x3FF0000010000000u, 0x3F800000u},
{"-1.000000059604644776e0", 0xBFF0000010000000u, 0xBF800001u},
{"-1.0000000596046447753e0", 0xBFF0000010000000u, 0xBF800000u},
{"10000000596046447754e-19", 0x3FF0000010000000u, 0x3F800001u},
{"100000005960464477539e-20", 0x3FF0000010000000u, 0x3F800000u},
{"-1.0000000596046447754", 0xBFF0000010000000u, 0xBF800001u},
{"0.9999999701976776123046875", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"9.999999701976776123046876e-1", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"9999999701976776123046874e-25", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu},
{"0.999999970197677612304687501", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"-9.999999701976776123046875000000000000000000001e-1", 0xBFEFFFFFF0000000u, 0xBF800000u},
{"-9999999701976776123046875000000000000000000000000000000e-55", 0xBFEFFFFFF0000000u, 0xBF800000u},
{"0.99999997019767761", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu},
{"-9.9999997019767762e-1", 0xBFEFFFFFF0000000u, 0xBF800000u},
{"-9.999999701976776123e-1", 0xBFEFFFFFF0000000u, 0xBF7FFFFFu},
{"9999999701976776124e-19", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"9999999701976776123e-19", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu},
{"0.99999997019767761231", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"-0.999999970197677612304", 0xBFEFFFFFF0000000u, 0xBF7FFFFFu},
{"9.99999970197677612305e-1", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"-1.6777217e7", 0xC170000010000000u, 0xCB800000u},
{"16777218e0", 0x4170000020000000u, 0x4B800001u},
{"16777216", 0x4170000000000000u, 0x4B800000u},
{"-1.677721701e7", 0xC17000001028F5C3u, 0xCB800001u},
{"-16777217000000000000000000001e-21", 0xC170000010000000u, 0xCB800001u},
{"16777217.000000000000000000000000000000", 0x4170000010000000u, 0x4B800000u},
{"167772155e-1", 0x416FFFFFF0000000u, 0x4B800000u},
{"16777215.6", 0x416FFFFFF3333333u, 0x4B800000u},
{"1.67772154e7", 0x416FFFFFECCCCCCDu, 0x4B7FFFFFu},
{"16777215501e-3", 0x416FFFFFF0083127u, 0x4B800000u},
{"-16777215.5000000000000000000001", 0xC16FFFFFF0000000u, 0xCB800000u},
{"-1.67772155000000000000000000000000000000e7", 0xC16FFFFFF0000000u, 0xCB800000u},
{"0.1000000052154064178466796875", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"1.000000052154064178466796876e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"-1000000052154064178466796874e-28", 0xBFB99999B0000000u, 0xBDCCCCCDu},
{"0.100000005215406417846679687501", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"1.000000052154064178466796875000000000000000000001e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"-1000000052154064178466796875000000000000000000000000000000e-58", 0xBFB99999B0000000u, 0xBDCCCCCEu},
{"-0.10000000521540641", 0xBFB99999AFFFFFFFu, 0xBDCCCCCDu},
{"-1.0000000521540642e-1", 0xBFB99999B0000000u, 0xBDCCCCCEu},
{"1.000000052154064178e-1", 0x3FB99999B0000000u, 0x3DCCCCCDu},
{"-1000000052154064179e-19", 0xBFB99999B0000000u, 0xBDCCCCCEu},
{"10000000521540641784e-20", 0x3FB99999B0000000u, 0x3DCCCCCDu},
{"0.10000000521540641785", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"0.100000005215406417846", 0x3FB99999B0000000u, 0x3DCCCCCDu},
{"1.00000005215406417847e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"5.429001220703125e3", 0x40B5350050000000u, 0x45A9A802u},
{"-5429001220703126e-12", 0xC0B5350050000001u, 0xC5A9A803u},
{"-5429.001220703124", 0xC0B535004FFFFFFFu, 0xC5A9A802u},
{"5.42900122070312501e3", 0x40B5350050000000u, 0x45A9A803u},
{"5429001220703125000000000000000000001e-33", 0x40B5350050000000u, 0x45A9A803u},
{"-5429.001220703125000000000000000000000000000000", 0xC0B5350050000000u, 0xC5A9A802u},
{"503719056e0", 0x41BE062490000000u, 0x4DF03124u},
{"503719057", 0x41BE062491000000u, 0x4DF03125u},
{"5.03719055e8", 0x41BE06248F000000u, 0x4DF03124u},
{"50371905601e-2", 0x41BE062490028F5Cu, 0x4DF03125u},
{"503719056.000000000000000000001", 0x41BE062490000000u, 0x4DF03125u},
{"5.03719056000000000000000000000000000000e8", 0x41BE062490000000u, 0x4DF03124u},
{"-92331620", 0xC196037990000000u, 0xCCB01BCCu},
{"9.233163e7", 0x41960379B8000000u, 0x4CB01BCEu},
{"9233161e1", 0x4196037968000000u, 0x4CB01BCBu},
{"92331620.1", 0x4196037990666666u, 0x4CB01BCDu},
{"9.233162000000000000000000001e7", 0x4196037990000000u, 0x4CB01BCDu},
{"9233162000000000000000000000000000000e-29", 0x4196037990000000u, 0x4CB01BCCu},
{"3.002458625e6", 0x4146E82D50000000u, 0x4A37416Au},
{"3002458626e-3", 0x4146E82D5020C49Cu, 0x4A37416Bu},
{"3002458.624", 0x4146E82D4FDF3B64u, 0x4A37416Au},
{"-3.00245862501e6", 0xC146E82D500053E3u, 0xCA37416Bu},
{"3002458625000000000000000000001e-24", 0x4146E82D50000000u, 0x4A37416Bu},
{"3002458.625000000000000000000000000000000", 0x4146E82D50000000u, 0x4A37416Au},
{"-1095485584696182596504479582065262592e1", 0xC7A07BA830000000u, 0xFD03DD42u},
{"10954855846961825965044795820652625930", 0x47A07BA830000000u, 0x7D03DD42u},
{"1.095485584696182596504479582065262591e37", 0x47A07BA830000000u, 0x7D03DD41u},
{"109548558469618259650447958206526259201e-1", 0x47A07BA830000000u, 0x7D03DD42u},
{"-10954855846961825965044795820652625920.00000000000000000001", 0xC7A07BA830000000u, 0xFD03DD42u},
{"1.095485584696182596504479582065262592000000000000000000000000000000e37", 0x47A07BA830000000u, 0x7D03DD42u},
{"10954855846961825e21", 0x47A07BA830000000u, 0x7D03DD41u},
{"10954855846961826000000000000000000000", 0x47A07BA830000000u, 0x7D03DD42u},
{"10954855846961825960000000000000000000", 0x47A07BA830000000u, 0x7D03DD41u},
{"1.095485584696182597e37", 0x47A07BA830000000u, 0x7D03DD42u},
{"1.0954855846961825965e37", 0x47A07BA830000000u, 0x7D03DD41u},
{"-10954855846961825966e18", 0xC7A07BA830000000u, 0xFD03DD42u},
{"10954855846961825965e18", 0x47A07BA830000000u, 0x7D03DD41u},
{"10954855846961825965100000000000000000", 0x47A07BA830000000u, 0x7D03DD42u},
{"10954855846961825965044795820600000000", 0x47A07BA830000000u, 0x7D03DD41u},
{"-1.09548558469618259650447958207e37", 0xC7A07BA830000000u, 0xFD03DD42u},
{"1.6449216019182103706535606608388384863861375606575165875256061553955078126e-21", 0x3B9F125A50000000u, 0x1CF892D3u},
{"-16449216019182103706535606608388384863861375606575165875256061553955078124e-94", 0xBB9F125A50000000u, 0x9CF892D2u},
{"-0.0000000000000000000016449216019182103", 0xBB9F125A50000000u, 0x9CF892D2u},
{"-1.6449216019182104e-21", 0xBB9F125A50000000u, 0x9CF892D3u},
{"-1.64492160191821037e-21", 0xBB9F125A50000000u, 0x9CF892D2u},
{"-1644921601918210371e-39", 0xBB9F125A50000000u, 0x9CF892D3u},
{"-16449216019182103706e-40", 0xBB9F125A50000000u, 0x9CF892D2u},
{"0.0000000000000000000016449216019182103707", 0x3B9F125A50000000u, 0x1CF892D3u},
{"0.00000000000000000000164492160191821037065", 0x3B9F125A50000000u, 0x1CF892D2u},
{"1.64492160191821037066e-21", 0x3B9F125A50000000u, 0x1CF892D3u},
{"1.64492160191821037065356066083e-21", 0x3B9F125A50000000u, 0x1CF892D2u},
{"164492160191821037065356066084e-50", 0x3B9F125A50000000u, 0x1CF892D3u},
{"6.565061509609222412109375e-1", 0x3FE5021930000000u, 0x3F2810CAu},
{"6565061509609222412109376e-25", 0x3FE5021930000000u, 0x3F2810CAu},
{"0.6565061509609222412109374", 0x3FE5021930000000u, 0x3F2810C9u},
{"6.56506150960922241210937501e-1", 0x3FE5021930000000u, 0x3F2810CAu},
{"6565061509609222412109375000000000000000000001e-46", 0x3FE5021930000000u, 0x3F2810CAu},
{"-0.6565061509609222412109375000000000000000000000000000000", 0xBFE5021930000000u, 0xBF2810CAu},
{"6.5650615096092224e-1", 0x3FE5021930000000u, 0x3F2810C9u},
{"-65650615096092225e-17", 0xBFE5021930000000u, 0xBF2810CAu},
{"6565061509609222412e-19", 0x3FE5021930000000u, 0x3F2810C9u},
{"0.6565061509609222413", 0x3FE5021930000000u, 0x3F2810CAu},
{"0.65650615096092224121", 0x3FE5021930000000u, 0x3F2810C9u},
{"6.5650615096092224122e-1", 0x3FE5021930000000u, 0x3F2810CAu},
{"-6.5650615096092224121e-1", 0xBFE5021930000000u, 0xBF2810C9u},
{"656506150960922241211e-21", 0x3FE5021930000000u, 0x3F2810CAu},
{"18014627239033005156980393746124491372029297053813934326171875e-77", 0x3CA9F63970000000u, 0x254FB1CCu},
{"0.00000000000000018014627239033005156980393746124491372029297053813934326171876", 0x3CA9F63970000000u, 0x254FB1CCu},
{"-1.8014627239033005156980393746124491372029297053813934326171874e-16", 0xBCA9F63970000000u, 0xA54FB1CBu},
{"1801462723903300515698039374612449137202929705381393432617187501e-79", 0x3CA9F63970000000u, 0x254FB1CCu},
{"18014627239033005e-32", 0x3CA9F63970000000u, 0x254FB1CBu},
{"0.00000000000000018014627239033006", 0x3CA9F63970000000u, 0x254FB1CCu},
{"0.0000000000000001801462723903300515", 0x3CA9F63970000000u, 0x254FB1CBu},
{"-1.801462723903300516e-16", 0xBCA9F63970000000u, 0xA54FB1CCu},
{"-1.8014627239033005156e-16", 0xBCA9F63970000000u, 0xA54FB1CBu},
{"18014627239033005157e-35", 0x3CA9F63970000000u, 0x254FB1CCu},
{"180146272390330051569e-36", 0x3CA9F63970000000u, 0x254FB1CBu},
{"-0.00000000000000018014627239033005157", 0xBCA9F63970000000u, 0xA54FB1CCu},
{"0.000000000000000180146272390330051569803937461", 0x3CA9F63970000000u, 0x254FB1CBu},
{"1.80146272390330051569803937462e-16", 0x3CA9F63970000000u, 0x254FB1CCu},
{"0.05534819327294826507568359375", 0x3FAC569930000000u, 0x3D62B4CAu},
{"-5.534819327294826507568359376e-2", 0xBFAC569930000000u, 0xBD62B4CAu},
{"5534819327294826507568359374e-29", 0x3FAC569930000000u, 0x3D62B4C9u},
{"-0.0553481932729482650756835937501", 0xBFAC569930000000u, 0xBD62B4CAu},
{"5.534819327294826507568359375000000000000000000001e-2", 0x3FAC569930000000u, 0x3D62B4CAu},
{"5534819327294826507568359375000000000000000000000000000000e-59", 0x3FAC569930000000u, 0x3D62B4CAu},
{"0.055348193272948265", 0x3FAC569930000000u, 0x3D62B4C9u},
{"5.5348193272948266e-2", 0x3FAC569930000000u, 0x3D62B4CAu},
{"5.534819327294826507e-2", 0x3FAC569930000000u, 0x3D62B4C9u},
{"5534819327294826508e-20", 0x3FAC569930000000u, 0x3D62B4CAu},
{"55348193272948265075e-21", 0x3FAC569930000000u, 0x3D62B4C9u},
{"-0.055348193272948265076", 0xBFAC569930000000u, 0xBD62B4CAu},
{"0.0553481932729482650756", 0x3FAC569930000000u, 0x3D62B4C9u},
{"-5.53481932729482650757e-2", 0xBFAC569930000000u, 0xBD62B4CAu},
{"5.179692133247783258005389047985340416e36", 0x478F2C9450000000u, 0x7C7964A2u},
{"5179692133247783258005389047985340417e0", 0x478F2C9450000000u, 0x7C7964A3u},
{"5179692133247783258005389047985340415", 0x478F2C9450000000u, 0x7C7964A2u},
{"-5.17969213324778325800538904798534041601e36", 0xC78F2C9450000000u, 0xFC7964A3u},
{"-5179692133247783258005389047985340416000000000000000000001e-21", 0xC78F2C9450000000u, 0xFC7964A3u},
{"-5179692133247783258005389047985340416.000000000000000000000000000000", 0xC78F2C9450000000u, 0xFC7964A2u},
{"5.1796921332477832e36", 0x478F2C9450000000u, 0x7C7964A2u},
{"51796921332477833e20", 0x478F2C9450000000u, 0x7C7964A3u},
{"-5179692133247783258e18", 0xC78F2C9450000000u, 0xFC7964A2u},
{"5179692133247783259000000000000000000", 0x478F2C9450000000u, 0x7C7964A3u},
{"5179692133247783258000000000000000000", 0x478F2C9450000000u, 0x7C7964A2u},
{"5.1796921332477832581e36", 0x478F2C9450000000u, 0x7C7964A3u},
{"-5.179692133247783258e36", 0xC78F2C9450000000u, 0xFC7964A2u},
{"-517969213324778325801e16", 0xC78F2C9450000000u, 0xFC7964A3u},
{"517969213324778325800538904798e7", 0x478F2C9450000000u, 0x7C7964A2u},
{"5179692133247783258005389047990000000", 0x478F2C9450000000u, 0x7C7964A3u},
{
"0.22250738585072011360574097967091319759348195463516456480234261097248222220210769455165295239081350"
"8791414915891303962110687008643869459464552765720740782062174337998814106326732925355228688137214901"
"2981122451451889849057222307285255133155755015914397476397983411801999323962548289017107081850690630"
"6666559949382757725720157630626906633326475653000092458883164330377797918696120494973903778297049050"
"5108060994073026293712895895000358379996720725430436028407889577179615094551674824347103070260914462"
"1572289880258182545180325707018860872113128079512233426288368622321503775666622503982534335974568884"
"4239002654981983854879482922068947216898310996983658468140228542433306603398508864458040010349339704"
"2756718644338377048603786162277173854562306587467901408672332763671875e-307", 0x0010000000000000u, 0x00000000u
},
{
"2.22507385850720113605740979670913197593481954635164564802342610972482222202107694551652952390813508"
"7914149158913039621106870086438694594645527657207407820621743379988141063267329253552286881372149012"
"9811224514518898490572223072852551331557550159143974763979834118019993239625482890171070818506906306"
"6665599493827577257201576306269066333264756530000924588831643303777979186961204949739037782970490505"
"1080609940730262937128958950003583799967207254304360284078895771796150945516748243471030702609144621"
"5722898802581825451803257070188608721131280795122334262883686223215037756666225039825343359745688844"
"2390026549819838548794829220689472168983109969836584681402285424333066033985088644580400103493397042"
"756718644338377048603786162277173854562306587467901408672332763671875000000000000000000001e-308", 0x0010000000000000u, 0x00000000u
},
{
"0.11754942807573642917278829910357665133228589927589904276829631184250030649651730385585324256680905"
"8189392089843750000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"00000000000000000000000000000000000000000000000000000000000000000e-37", 0x380FFFFFE0000000u, 0x00800000u
},
{
"1175494280757364291727882991035766513322858992758990427682963118425003064965173038558532425668090581"
"8939208984375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"00000000000001e-751", 0x380FFFFFE0000000u, 0x00800000u
},
{"0", 0x0000000000000000u, 0x00000000u},
{"-0", 0x8000000000000000u, 0x80000000u},
{"0.0", 0x0000000000000000u, 0x00000000u},
{"-0.0", 0x8000000000000000u, 0x80000000u},
{"0e999999999999999999999", 0x0000000000000000u, 0x00000000u},
{"-0.000e-99999", 0x8000000000000000u, 0x80000000u},
{"1e-400", 0x0000000000000000u, 0x00000000u},
{"-1e-400", 0x8000000000000000u, 0x80000000u},
{"1e400", 0x7FF0000000000000u, 0x7F800000u},
{"-1e400", 0xFFF0000000000000u, 0xFF800000u},
{"1e-50", 0x358DEE7A4AD4B81Fu, 0x00000000u},
{"-1e-50", 0xB58DEE7A4AD4B81Fu, 0x80000000u},
{"1e39", 0x48078287F49C4A1Du, 0x7F800000u},
{"-1e39", 0xC8078287F49C4A1Du, 0xFF800000u},
{"1e99999999999999999999999999", 0x7FF0000000000000u, 0x7F800000u},
{"1e-99999999999999999999999999", 0x0000000000000000u, 0x00000000u},
{"1e0000000000000000000000000000000000000000308", 0x7FE1CCF385EBC8A0u, 0x7F800000u},
{"123456789012345678901234567890e-30", 0x3FBF9ADD3746F65Fu, 0x3DFCD6EAu},
{"18446744073709551615", 0x43F0000000000000u, 0x5F800000u},
{"18446744073709551616", 0x43F0000000000000u, 0x5F800000u},
{"-9223372036854775808", 0xC3E0000000000000u, 0xDF000000u},
{"-9223372036854775809", 0xC3E0000000000000u, 0xDF000000u},
}
};
return table;
}
} // namespace float_hard_cases
-28
View File
@@ -8,7 +8,6 @@
#pragma once #pragma once
#include <array> // array
#include <cstdint> // uint8_t #include <cstdint> // uint8_t
#include <cstddef> // size_t #include <cstddef> // size_t
#include <fstream> // ifstream, ios #include <fstream> // ifstream, ios
@@ -44,33 +43,6 @@ T next_integer_sample(T i, T last, T stride)
return n < last ? n : last; return n < last ? n : last;
} }
// UTF-8 continuation bytes in [lo, hi] that stand in for all of them in the
// ill-formed UTF-8 tests. Both the lexer's range checks and the serializer's
// decoder (detail::decode) only distinguish the classes 0x80..0x8F, 0x90..0x9F,
// and 0xA0..0xBF, so the first and last byte of each class within [lo, hi]
// exercise every behavior while a test sweeps another byte position through
// all 256 values (#5418). Define JSON_TEST_UTF8_EXHAUSTIVE to get every byte.
inline std::vector<int> utf8_continuation_bytes(int lo, int hi)
{
std::vector<int> result;
#ifdef JSON_TEST_UTF8_EXHAUSTIVE
for (int byte = lo; byte <= hi; ++byte)
{
result.push_back(byte);
}
#else
static const std::array<int, 6> class_ends = {{0x80, 0x8F, 0x90, 0x9F, 0xA0, 0xBF}};
for (const int byte : class_ends)
{
if (lo <= byte && byte <= hi)
{
result.push_back(byte);
}
}
#endif
return result;
}
inline std::vector<std::uint8_t> read_binary_file(const std::string& filename) inline std::vector<std::uint8_t> read_binary_file(const std::string& filename)
{ {
std::ifstream file(filename, std::ios::binary); std::ifstream file(filename, std::ios::binary);
+43 -45
View File
@@ -14,7 +14,7 @@ using nlohmann::json;
#include <fstream> #include <fstream>
#include "make_test_data_available.hpp" #include "make_test_data_available.hpp"
TEST_CASE("Binary Formats") TEST_CASE("Binary Formats" * doctest::skip())
{ {
SECTION("canada.json") SECTION("canada.json")
{ {
@@ -142,6 +142,48 @@ TEST_CASE("Binary Formats")
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(84.963)); CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(84.963));
} }
SECTION("jeopardy.json")
{
const auto* filename = TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json";
json j = json::parse(std::ifstream(filename));
const auto json_size = j.dump().size();
const auto bjdata_1_size = json::to_bjdata(j).size();
const auto bjdata_2_size = json::to_bjdata(j, true).size();
const auto bjdata_3_size = json::to_bjdata(j, true, true).size();
const auto bon8_size = json::to_bon8(j).size();
const auto bson_size = json::to_bson({{"", j}}).size(); // wrap array in object for BSON
const auto cbor_size = json::to_cbor(j).size();
const auto msgpack_size = json::to_msgpack(j).size();
const auto ubjson_1_size = json::to_ubjson(j).size();
const auto ubjson_2_size = json::to_ubjson(j, true).size();
const auto ubjson_3_size = json::to_ubjson(j, true, true).size();
CHECK(json_size == 52508728);
CHECK(bjdata_1_size == 50710965);
CHECK(bjdata_2_size == 51144830);
CHECK(bjdata_3_size == 51144830);
CHECK(bon8_size == 45942080);
CHECK(bson_size == 56008520);
CHECK(cbor_size == 46187320);
CHECK(msgpack_size == 46158575);
CHECK(ubjson_1_size == 50710965);
CHECK(ubjson_2_size == 51144830);
CHECK(ubjson_3_size == 49861422);
CHECK((100.0 * double(json_size) / double(json_size)) == Approx(100.0));
CHECK((100.0 * double(bjdata_1_size) / double(json_size)) == Approx(96.576));
CHECK((100.0 * double(bjdata_2_size) / double(json_size)) == Approx(97.402));
CHECK((100.0 * double(bjdata_3_size) / double(json_size)) == Approx(97.402));
CHECK((100.0 * double(bon8_size) / double(json_size)) == Approx(87.494));
CHECK((100.0 * double(bson_size) / double(json_size)) == Approx(106.665));
CHECK((100.0 * double(cbor_size) / double(json_size)) == Approx(87.961));
CHECK((100.0 * double(msgpack_size) / double(json_size)) == Approx(87.906));
CHECK((100.0 * double(ubjson_1_size) / double(json_size)) == Approx(96.576));
CHECK((100.0 * double(ubjson_2_size) / double(json_size)) == Approx(97.402));
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(94.958));
}
SECTION("sample.json") SECTION("sample.json")
{ {
const auto* filename = TEST_DATA_DIRECTORY "/json_testsuite/sample.json"; const auto* filename = TEST_DATA_DIRECTORY "/json_testsuite/sample.json";
@@ -182,47 +224,3 @@ TEST_CASE("Binary Formats")
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450)); CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450));
} }
} }
// jeopardy.json is 52 MB and produces ~500 MB of serialization output, so it
// is kept apart from the cheap corpus files above (#5418)
TEST_CASE("Binary Formats (jeopardy.json)" * doctest::skip())
{
const auto* filename = TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json";
json j = json::parse(std::ifstream(filename));
const auto json_size = j.dump().size();
const auto bjdata_1_size = json::to_bjdata(j).size();
const auto bjdata_2_size = json::to_bjdata(j, true).size();
const auto bjdata_3_size = json::to_bjdata(j, true, true).size();
const auto bon8_size = json::to_bon8(j).size();
const auto bson_size = json::to_bson({{"", j}}).size(); // wrap array in object for BSON
const auto cbor_size = json::to_cbor(j).size();
const auto msgpack_size = json::to_msgpack(j).size();
const auto ubjson_1_size = json::to_ubjson(j).size();
const auto ubjson_2_size = json::to_ubjson(j, true).size();
const auto ubjson_3_size = json::to_ubjson(j, true, true).size();
CHECK(json_size == 52508728);
CHECK(bjdata_1_size == 50710965);
CHECK(bjdata_2_size == 51144830);
CHECK(bjdata_3_size == 51144830);
CHECK(bon8_size == 45942080);
CHECK(bson_size == 56008520);
CHECK(cbor_size == 46187320);
CHECK(msgpack_size == 46158575);
CHECK(ubjson_1_size == 50710965);
CHECK(ubjson_2_size == 51144830);
CHECK(ubjson_3_size == 49861422);
CHECK((100.0 * double(json_size) / double(json_size)) == Approx(100.0));
CHECK((100.0 * double(bjdata_1_size) / double(json_size)) == Approx(96.576));
CHECK((100.0 * double(bjdata_2_size) / double(json_size)) == Approx(97.402));
CHECK((100.0 * double(bjdata_3_size) / double(json_size)) == Approx(97.402));
CHECK((100.0 * double(bon8_size) / double(json_size)) == Approx(87.494));
CHECK((100.0 * double(bson_size) / double(json_size)) == Approx(106.665));
CHECK((100.0 * double(cbor_size) / double(json_size)) == Approx(87.961));
CHECK((100.0 * double(msgpack_size) / double(json_size)) == Approx(87.906));
CHECK((100.0 * double(ubjson_1_size) / double(json_size)) == Approx(96.576));
CHECK((100.0 * double(ubjson_2_size) / double(json_size)) == Approx(97.402));
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(94.958));
}
+278 -111
View File
@@ -13,15 +13,18 @@
using nlohmann::json; using nlohmann::json;
#include <array> // array #include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstdint> // uint32_t, uint64_t #include <cstdint> // uint32_t, uint64_t
#include <cstdio> // snprintf
#include <cstdlib> // strtod #include <cstdlib> // strtod
#include <cstring> // memcpy #include <cstring> // memcpy
#include <map> // map
#include <sstream> // stringstream #include <sstream> // stringstream
#include <string> // string #include <string> // string
#include <utility> // pair #include <utility> // pair
#include <vector> // vector #include <vector> // vector
#include "float_hard_cases.hpp"
namespace namespace
{ {
// shortcut to scan a string literal // shortcut to scan a string literal
@@ -257,7 +260,7 @@ TEST_CASE("lexer number fast path")
"123456789012345678901234567890", // huge -> float "123456789012345678901234567890", // huge -> float
"0.30000000000000004", "2.2250738585072014e-308", "1e308", "0.30000000000000004", "2.2250738585072014e-308", "1e308",
// high-precision / wide-exponent values that exercise the // high-precision / wide-exponent values that exercise the
// std::from_chars (Eisel-Lemire) path beyond the Clinger subset // Eisel-Lemire path beyond the Clinger subset
"1.7976931348623157e308", "1.2345678901234567e-250", "1.7976931348623157e308", "1.2345678901234567e-250",
"9007199254740993", "5e-324", "1e-320" "9007199254740993", "5e-324", "1e-320"
}; };
@@ -279,20 +282,18 @@ TEST_CASE("lexer number fast path")
} }
} }
SECTION("significant-digit gate for the Clinger fast path") SECTION("significant digits around Clinger's fast path")
{ {
// Clinger's fast path needs a significand below 2^53, so it cannot // Clinger's fast path needs a significand of at most 2^53, which
// succeed once the mantissa has 17 or more significant digits (the // tokens with 17 or more significant digits exceed. The conversion
// significand would be at least 10^16). The lexer skips the attempt // splits the token at the positions the scanners recorded, so leading
// there. That is only allowed to save work: every value must still come // zeros must not count as digits - "0.1234567890123456" has 16
// out bit-exactly, and both scanners must agree. In particular the gate // significant digits, not 17 - and both scanners must agree.
// must not fire for tokens whose leading zeros merely look like extra
// digits - "0.1234567890123456" has 16 significant digits, not 17.
const std::vector<std::string> numbers = const std::vector<std::string> numbers =
{ {
"1234567890123456", // 16 significant digits "1234567890123456", // 16 significant digits
"12345678901234567", // 17 -> attempt skipped "12345678901234567", // 17
"123456789012345678", // 18 -> attempt skipped "123456789012345678", // 18
"0.1234567890123456", // 16: the leading "0" is not significant "0.1234567890123456", // 16: the leading "0" is not significant
"0.12345678901234567", // 17 "0.12345678901234567", // 17
"0.00000000000000001", // 1, in a long token "0.00000000000000001", // 1, in a long token
@@ -663,46 +664,145 @@ TEST_CASE("lexer string fast path")
} }
} }
TEST_CASE("parse_float_fast declines what it cannot convert exactly") namespace
{ {
// The lexer only hands well-formed numbers to parse_float_fast, so the // the index of the decimal point (or npos) and of the end of the mantissa of a
// malformed ones below can only be passed to it directly. Declining is // number token, which the lexer records while scanning it
// always safe: the caller then falls back to a slower, exact conversion. std::pair<std::size_t, std::size_t> float_token_layout(const std::string& s)
const auto fast = [](const std::string & s, double & out) {
std::size_t dot = std::string::npos;
std::size_t mantissa_end = s.size();
for (std::size_t i = 0; i < s.size(); ++i)
{ {
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out); if (s[i] == '.')
{
dot = i;
}
else if (s[i] == 'e' || s[i] == 'E')
{
mantissa_end = i;
break;
}
}
return {dot, mantissa_end};
}
template<typename FloatType>
FloatType parse_native(const std::string& s)
{
const auto layout = float_token_layout(s);
return nlohmann::detail::parse_float_native<FloatType>(s.data(), s.data() + s.size(), layout.first, layout.second);
}
std::uint64_t bits_of(double d)
{
std::uint64_t b = 0;
std::memcpy(&b, &d, sizeof(b));
return b;
}
std::uint32_t bits_of(float f)
{
std::uint32_t b = 0;
std::memcpy(&b, &f, sizeof(b));
return b;
}
std::uint64_t native_bits64(const std::string& s)
{
return bits_of(parse_native<double>(s));
}
std::uint32_t native_bits32(const std::string& s)
{
return bits_of(parse_native<float>(s));
}
} // namespace
TEST_CASE("parse_float_native rounds correctly")
{
SECTION("double")
{
CHECK(native_bits64("1.5") == 0x3FF8000000000000u);
CHECK(native_bits64("0.1") == 0x3FB999999999999Au);
CHECK(native_bits64("-0.0") == 0x8000000000000000u);
CHECK(native_bits64("0e999999999999999999999") == 0u);
// 2^53 + 1 is exactly between two doubles: ties to even, unless more digits follow
CHECK(native_bits64("9007199254740993") == 0x4340000000000000u);
CHECK(native_bits64("9007199254740993.0000000000000000001") == 0x4340000000000001u);
CHECK(native_bits64("9007199254740992.9999999999999999999") == 0x4340000000000000u);
// 1 + 2^-53 exactly (a tie), and one unit in the 55th digit around it
CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203125") == 0x3FF0000000000000u);
CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203126") == 0x3FF0000000000001u);
CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203124") == 0x3FF0000000000000u);
// subnormal and overflow boundaries
CHECK(native_bits64("2.4703282292062327e-324") == 0u);
CHECK(native_bits64("2.4703282292062328e-324") == 1u);
CHECK(native_bits64("2.2250738585072011e-308") == 0x000FFFFFFFFFFFFFu);
CHECK(native_bits64("2.2250738585072012e-308") == 0x0010000000000000u);
CHECK(native_bits64("1.7976931348623157e308") == 0x7FEFFFFFFFFFFFFFu);
CHECK(native_bits64("1.7976931348623159e308") == 0x7FF0000000000000u);
CHECK(native_bits64("-1e400") == 0xFFF0000000000000u);
CHECK(native_bits64("-1e-400") == 0x8000000000000000u);
// exponents and zeros far beyond the range cancel out
CHECK(native_bits64("0." + std::string(1000, '0') + "1e1001") == 0x3FF0000000000000u);
CHECK(native_bits64("1" + std::string(1000, '0') + "e-1000") == 0x3FF0000000000000u);
CHECK(native_bits64("1e-99999999999999999999999") == 0u);
CHECK(native_bits64("1E+99999999999999999999999") == 0x7FF0000000000000u);
// more digits than any midpoint has (769): only whether a nonzero digit follows matters
const std::string tie = "1.00000000000000011102230246251565404236316680908203125";
CHECK(native_bits64(tie + std::string(800, '0')) == 0x3FF0000000000000u);
CHECK(native_bits64(tie + std::string(800, '0') + "1") == 0x3FF0000000000001u);
}
SECTION("float")
{
CHECK(native_bits32("1.5") == 0x3FC00000u);
CHECK(native_bits32("0.1") == 0x3DCCCCCDu);
CHECK(native_bits32("-0.0") == 0x80000000u);
// 2^24 + 1 is exactly between two floats
CHECK(native_bits32("16777217") == 0x4B800000u);
CHECK(native_bits32("16777217.000000000000000000001") == 0x4B800001u);
CHECK(native_bits32("16777218.999999999999999999999") == 0x4B800001u);
CHECK(native_bits32("16777219") == 0x4B800002u);
// subnormal and overflow boundaries
CHECK(native_bits32("3.4028235677973366e38") == 0x7F7FFFFFu);
CHECK(native_bits32("3.4028235677973367e38") == 0x7F800000u);
CHECK(native_bits32("7.006492321624085e-46") == 0u);
CHECK(native_bits32("7.006492321624086e-46") == 1u);
CHECK(native_bits32("1.1754942e-38") == 0x007FFFFFu);
CHECK(native_bits32("-1.17549435e-38") == 0x80800000u);
CHECK(native_bits32("1e39") == 0x7F800000u);
CHECK(native_bits32("-1e-50") == 0x80000000u);
// not rounded through double: its double would round to another float
CHECK(native_bits32("1.00000005960464477539062500000000001") == 0x3F800001u);
CHECK(native_bits32("9007199254740993") == 0x5A000000u);
}
SECTION("the conversion shared with other parsers")
{
// convert_float() gives the lexer's results, for every type
const std::vector<std::string> tokens =
{
"0", "-0.0", "1.5", "0.1", "1e-400", "-2.5E+3", "123456789012345678901234567890",
"9007199254740993.0000000000000000001", "4.9406564584124654e-324"
}; };
double out = 0; using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 for (const auto& t : tokens)
// without true double precision, the fast path declines everything {
CHECK_FALSE(fast("1.5", out)); CAPTURE(t);
#else const auto layout = float_token_layout(t);
CHECK(fast("1.5", out)); const char* const first = t.data();
CHECK(out == 1.5); const char* const last = first + t.size();
CHECK(fast("+2.5e1", out)); const auto d = nlohmann::detail::convert_float<double>(first, last, layout.first, layout.second);
CHECK(out == 25.0); const auto f = nlohmann::detail::convert_float<float>(first, last, layout.first, layout.second);
CHECK(fast("-25E-1", out)); const auto ld = nlohmann::detail::convert_float<long double>(first, last, layout.first, layout.second);
CHECK(out == -2.5); CHECK(bits_of(d) == bits_of(json::parse(t).get<double>()));
CHECK(fast("1e", out)); CHECK(bits_of(f) == bits_of(float_json::parse(t).get<float>()));
CHECK(out == 1.0); CHECK(ld == long_double_json::parse(t).get<long double>());
#endif }
}
// not a number
CHECK_FALSE(fast("", out));
CHECK_FALSE(fast("-", out));
CHECK_FALSE(fast(".", out));
CHECK_FALSE(fast("1.2.3", out));
CHECK_FALSE(fast("1x", out));
CHECK_FALSE(fast("1e+", out));
CHECK_FALSE(fast("1e1x", out));
// numbers that are not represented exactly on the fast path
CHECK_FALSE(fast("12345678901234567890", out));
CHECK_FALSE(fast("1e10000", out));
CHECK_FALSE(fast("9007199254740993", out));
CHECK_FALSE(fast("1e23", out));
CHECK_FALSE(fast("1e-23", out));
} }
namespace namespace
@@ -806,40 +906,6 @@ std::size_t big_bit_length(const big_uint& a)
} }
return n; return n;
} }
std::uint64_t bits_of(double d)
{
std::uint64_t b = 0;
std::memcpy(&b, &d, sizeof(b));
return b;
}
bool eisel_lemire(const std::string& s, double& out)
{
return nlohmann::detail::parse_float_eisel_lemire(s.data(), s.data() + s.size(), out);
}
// significant digits of a token, without trailing zeros
std::size_t significant_digits(const std::string& s)
{
std::string digits;
for (const char c : s)
{
if (c == 'e' || c == 'E')
{
break;
}
if (c >= '0' && c <= '9' && !(digits.empty() && c == '0'))
{
digits += c;
}
}
while (!digits.empty() && digits.back() == '0')
{
digits.pop_back();
}
return digits.size();
}
} // namespace } // namespace
TEST_CASE("Eisel-Lemire float conversion") TEST_CASE("Eisel-Lemire float conversion")
@@ -1237,26 +1303,33 @@ TEST_CASE("Eisel-Lemire float conversion")
for (const auto& c : known) for (const auto& c : known)
{ {
CAPTURE(c.first) CAPTURE(c.first)
double out = 0; CHECK(native_bits64(c.first) == c.second);
if (eisel_lemire(c.first, out)) }
}
SECTION("binary32")
{ {
CHECK(bits_of(out) == c.second); using binary32 = nlohmann::detail::ieee_binary_format<24>;
} CHECK(nlohmann::detail::eisel_lemire<binary32>(0, 1) == 0x3F800000u);
else CHECK(nlohmann::detail::eisel_lemire<binary32>(-1, 1) == 0x3DCCCCCDu);
{ CHECK(nlohmann::detail::eisel_lemire<binary32>(-1, 15) == 0x3FC00000u);
// only tokens with more than 19 significant digits are left to CHECK(nlohmann::detail::eisel_lemire<binary32>(0, 16777217) == 0x4B800000u); // tie, to even
// strtod: those whose value lies too close to a tie CHECK(nlohmann::detail::eisel_lemire<binary32>(0, 16777219) == 0x4B800002u); // tie, to even
CHECK(significant_digits(c.first) > 19); CHECK(nlohmann::detail::eisel_lemire<binary32>(-45, 1) == 0x00000001u);
} CHECK(nlohmann::detail::eisel_lemire<binary32>(-46, 7) == 0x00000000u);
} CHECK(nlohmann::detail::eisel_lemire<binary32>(-46, 8) == 0x00000001u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-65, 9999999999999999999u) == 0x00000000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(20, 3402823466385288598u) == 0x7F7FFFFFu);
CHECK(nlohmann::detail::eisel_lemire<binary32>(20, 3402823669209384635u) == 0x7F800000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(39, 1) == 0x7F800000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-5, 0) == 0x00000000u);
} }
SECTION("round trip") SECTION("round trip")
{ {
// every double written by to_chars and read back, also with trailing // every double written by to_chars and read back, and its 17-digit
// digits that make the token longer than 19 digits // form with trailing digits that make the token longer than 19 digits
std::uint64_t state = 5295; std::uint64_t state = 5295;
std::size_t declined = 0;
for (int i = 0; i < 200000; ++i) for (int i = 0; i < 200000; ++i)
{ {
state ^= state << 13u; state ^= state << 13u;
@@ -1278,30 +1351,51 @@ TEST_CASE("Eisel-Lemire float conversion")
const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), d); const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), d);
const std::string token(buffer.data(), static_cast<std::size_t>(end - buffer.data())); const std::string token(buffer.data(), static_cast<std::size_t>(end - buffer.data()));
CAPTURE(token) CAPTURE(token)
double out = 0; CHECK(native_bits64(token) == b);
REQUIRE(eisel_lemire(token, out));
CHECK(bits_of(out) == b);
// insert digits before the exponent: the value moves by far less // insert digits before the exponent of the 17-digit form: that
// than the distance to the rounding boundary, so it must not change // form lies strictly inside the rounding interval of the double
std::string longer = token; // (the shortest one may lie on its boundary), and the digits move
// it by far less than the distance to the boundary, so the value
// must not change
std::array<char, 64> digits17{};
static_cast<void>(std::snprintf(digits17.data(), digits17.size(), "%.17g", d)); // NOLINT(cppcoreguidelines-pro-type-vararg,hicpp-vararg)
std::string longer = digits17.data();
const std::size_t e = longer.find('e'); const std::size_t e = longer.find('e');
const std::size_t dot = longer.find('.'); const std::size_t dot = longer.find('.');
const std::string extra = dot == std::string::npos ? ".000000000000000000001" : "000000000000000000001"; const std::string extra = dot == std::string::npos ? ".000000000000000000001" : "000000000000000000001";
longer.insert(e == std::string::npos ? longer.size() : e, extra); longer.insert(e == std::string::npos ? longer.size() : e, extra);
CAPTURE(longer) CAPTURE(longer)
if (eisel_lemire(longer, out)) CHECK(native_bits64(longer) == b);
}
}
SECTION("round trip, binary32")
{ {
CHECK(bits_of(out) == b); std::uint32_t state = 5295;
} for (int i = 0; i < 100000; ++i)
else
{ {
// w and w + 1 round differently: only when the value is very state ^= state << 13u;
// close to a rounding boundary state ^= state >> 17u;
++declined; state ^= state << 5u;
std::uint32_t b = state;
if ((b & 0x7F800000u) == 0x7F800000u)
{
continue; // infinity or NaN
} }
if (i % 4 == 0)
{
b &= 0x807FFFFFu; // subnormals
}
float f = 0;
std::memcpy(&f, &b, sizeof(f));
std::array<char, 64> buffer{};
const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), f);
const std::string token(buffer.data(), static_cast<std::size_t>(end - buffer.data()));
CAPTURE(token);
CHECK(native_bits32(token) == b);
} }
CHECK(declined < 1000); // 107 of the 200,000
} }
SECTION("used by the lexer") SECTION("used by the lexer")
@@ -1315,3 +1409,76 @@ TEST_CASE("Eisel-Lemire float conversion")
"[json.exception.out_of_range.406] number overflow parsing '1.7976931348623159e308'", json::out_of_range&); "[json.exception.out_of_range.406] number overflow parsing '1.7976931348623159e308'", json::out_of_range&);
} }
} }
namespace
{
using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
// the bits of the float that parse() gives for a token, via both scanners;
// the value must be the same for both
template<typename Json, typename Bits>
void check_parse(const std::string& token, Bits expected, Bits infinity)
{
std::stringstream stream(token);
if ((expected & ~(Bits{1} << (8 * sizeof(Bits) - 1))) == infinity)
{
Json _;
CHECK_THROWS_WITH_AS(_ = Json::parse(token), ("[json.exception.out_of_range.406] number overflow parsing '" + token + "'").c_str(), typename Json::out_of_range&);
CHECK_THROWS_WITH_AS(_ = Json::parse(stream), ("[json.exception.out_of_range.406] number overflow parsing '" + token + "'").c_str(), typename Json::out_of_range&);
return;
}
const Json contiguous = Json::parse(token);
const Json streamed = Json::parse(stream);
if (contiguous.is_number_float()) // not an integer that fits
{
CHECK(bits_of(contiguous.template get<typename Json::number_float_t>()) == expected);
CHECK(bits_of(streamed.template get<typename Json::number_float_t>()) == expected);
}
else
{
CHECK(streamed.is_number_integer());
}
}
} // namespace
TEST_CASE("float conversion of hard cases")
{
// see float_hard_cases.hpp
for (const auto& c : float_hard_cases::cases())
{
const std::string token = c.token;
CAPTURE(token);
CHECK(native_bits64(token) == c.bits64);
CHECK(native_bits32(token) == c.bits32);
check_parse<json>(token, c.bits64, std::uint64_t{0x7FF0000000000000u});
check_parse<float_json>(token, c.bits32, std::uint32_t{0x7F800000u});
}
}
TEST_CASE("float overflow and underflow in the parser")
{
SECTION("double")
{
check_parse<json>("1.7976931348623157e308", std::uint64_t{0x7FEFFFFFFFFFFFFFu}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("1.7976931348623159e308", std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("-1e309", std::uint64_t{0xFFF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("1" + std::string(400, '0'), std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("1e99999999999999999999", std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
// an underflow gives a zero with the sign of the token
check_parse<json>("1e-400", std::uint64_t{0}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("-1e-400", std::uint64_t{0x8000000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("-2.4703282292062327e-324", std::uint64_t{0x8000000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("0." + std::string(400, '0') + "1", std::uint64_t{0}, std::uint64_t{0x7FF0000000000000u});
}
SECTION("float")
{
check_parse<float_json>("3.4028234e38", std::uint32_t{0x7F7FFFFFu}, std::uint32_t{0x7F800000u});
check_parse<float_json>("3.4028236e38", std::uint32_t{0x7F800000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-1e39", std::uint32_t{0xFF800000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("1e-46", std::uint32_t{0}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-1e-46", std::uint32_t{0x80000000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-7.006492321624085e-46", std::uint32_t{0x80000000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-7.006492321624086e-46", std::uint32_t{0x80000001u}, std::uint32_t{0x7F800000u});
}
}
+25 -8
View File
@@ -260,10 +260,11 @@ struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
TEST_CASE("locale changes between lexer construction and number conversion (#5198)") TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
{ {
// The numbers are chosen so that the conversion also takes the strtod // float and double are converted without the locale. A long double that
// fallback, which honors the locale that is current at conversion time: // is not binary64 can take the strtold fallback, which honors the locale
// too many significant digits for Clinger's fast path, an underflow that // that is current at conversion time. The numbers are chosen so that it
// std::from_chars rejects, and a plain value. // does: too many significant digits for Clinger's fast path, an underflow
// that std::from_chars rejects, and a plain value.
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"}; const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
std::string text = "["; std::string text = "[";
for (const auto& n : numbers) for (const auto& n : numbers)
@@ -327,7 +328,8 @@ TEST_CASE("locale changes between lexer construction and number conversion (#519
} }
} }
// a long double goes through std::strtold unless std::from_chars supports it // a long double goes through std::strtold unless it is binary64 or
// std::from_chars supports it
{ {
bool switched = false; bool switched = false;
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
@@ -353,8 +355,15 @@ TEST_CASE("locale with a multi-byte decimal point")
{ {
// Some locales use a decimal point that is not a single character, e.g. // Some locales use a decimal point that is not a single character, e.g.
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be // U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
// substituted in place for '.', so the strtod fallback stops early. The // substituted in place for '.', so the strtold fallback (only for long
// conversion must still terminate rather than retry forever. // double formats other than binary64) converts a copy of the token with
// the whole decimal point instead (#5660). The values must be those of the
// "C" locale.
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
const char* const long_double_numbers = "[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]";
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
const long_double_json expected_long_double = long_double_json::parse(long_double_numbers);
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}}; const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
bool tested = false; bool tested = false;
for (const char* name : names) for (const char* name : names)
@@ -372,12 +381,20 @@ TEST_CASE("locale with a multi-byte decimal point")
tested = true; tested = true;
// too many significant digits for Clinger's fast path, and an underflow // too many significant digits for Clinger's fast path, and an underflow
// that std::from_chars rejects: both reach the strtod fallback // that std::from_chars rejects: double does not depend on the locale
json j; json j;
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]")); CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
CHECK(j.is_array()); CHECK(j.is_array());
CHECK(j[0] == 3.14159265358979323846);
CHECK(j[1] == 0.0);
CHECK(j[2] == -0.000123456789012345678);
CHECK(json::accept("3.14159265358979323846")); CHECK(json::accept("3.14159265358979323846"));
// a long double that reaches the strtold fallback is not truncated
long_double_json ld;
CHECK_NOTHROW(ld = long_double_json::parse(long_double_numbers));
CHECK(ld == expected_long_double);
// a value the locale-independent paths convert is not affected // a value the locale-independent paths convert is not affected
CHECK(json::parse("12.5") == 12.5); CHECK(json::parse("12.5") == 12.5);
} }
-99
View File
@@ -1,99 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// This file contains the C++17-only part of unit-msgpack.cpp (std::byte
// input). It is kept in a separate translation unit so the (much larger)
// unit-msgpack.cpp does not need to be compiled and run a second time just
// for this one test case (#5418).
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
using nlohmann::json;
#ifdef JSON_HAS_CPP_17
#include <cstddef>
#include <vector>
// Test suite for verifying MessagePack handling with std::byte input
TEST_CASE("MessagePack with std::byte")
{
SECTION("std::byte compatibility")
{
SECTION("vector roundtrip")
{
json original =
{
{"name", "test"},
{"value", 42},
{"array", {1, 2, 3}}
};
std::vector<uint8_t> temp = json::to_msgpack(original);
// Convert the uint8_t vector to std::byte vector
std::vector<std::byte> msgpack_data(temp.size());
for (size_t i = 0; i < temp.size(); ++i)
{
msgpack_data[i] = std::byte(temp[i]);
}
// Deserialize from std::byte vector back to JSON
json from_bytes;
CHECK_NOTHROW(from_bytes = json::from_msgpack(msgpack_data));
CHECK(from_bytes == original);
}
SECTION("empty vector")
{
const std::vector<std::byte> empty_data;
CHECK_THROWS_WITH_AS([&]()
{
[[maybe_unused]] auto result = json::from_msgpack(empty_data);
return true;
}
(),
"[json.exception.parse_error.110] parse error at byte 1: syntax error while parsing MessagePack value: unexpected end of input",
json::parse_error&);
}
SECTION("comparison with workaround")
{
json original =
{
{"string", "hello"},
{"integer", 42},
{"float", 3.14},
{"boolean", true},
{"null", nullptr},
{"array", {1, 2, 3}},
{"object", {{"key", "value"}}}
};
std::vector<uint8_t> temp = json::to_msgpack(original);
std::vector<std::byte> msgpack_data(temp.size());
for (size_t i = 0; i < temp.size(); ++i)
{
msgpack_data[i] = std::byte(temp[i]);
}
// Attempt direct deserialization using std::byte input
const json direct_result = json::from_msgpack(msgpack_data);
// Test the workaround approach: reinterpret as unsigned char* and use iterator range
const auto* const char_start = reinterpret_cast<unsigned char const*>(msgpack_data.data());
const auto* const char_end = char_start + msgpack_data.size();
json workaround_result = json::from_msgpack(char_start, char_end);
// Verify that the final deserialized JSON matches the original JSON
CHECK(direct_result == workaround_result);
CHECK(direct_result == original);
}
}
}
#endif
+79
View File
@@ -2122,6 +2122,85 @@ TEST_CASE("MessagePack roundtrips" * doctest::skip())
} }
} }
#ifdef JSON_HAS_CPP_17
// Test suite for verifying MessagePack handling with std::byte input
TEST_CASE("MessagePack with std::byte")
{
SECTION("std::byte compatibility")
{
SECTION("vector roundtrip")
{
json original =
{
{"name", "test"},
{"value", 42},
{"array", {1, 2, 3}}
};
std::vector<uint8_t> temp = json::to_msgpack(original);
// Convert the uint8_t vector to std::byte vector
std::vector<std::byte> msgpack_data(temp.size());
for (size_t i = 0; i < temp.size(); ++i)
{
msgpack_data[i] = std::byte(temp[i]);
}
// Deserialize from std::byte vector back to JSON
json from_bytes;
CHECK_NOTHROW(from_bytes = json::from_msgpack(msgpack_data));
CHECK(from_bytes == original);
}
SECTION("empty vector")
{
const std::vector<std::byte> empty_data;
CHECK_THROWS_WITH_AS([&]()
{
[[maybe_unused]] auto result = json::from_msgpack(empty_data);
return true;
}
(),
"[json.exception.parse_error.110] parse error at byte 1: syntax error while parsing MessagePack value: unexpected end of input",
json::parse_error&);
}
SECTION("comparison with workaround")
{
json original =
{
{"string", "hello"},
{"integer", 42},
{"float", 3.14},
{"boolean", true},
{"null", nullptr},
{"array", {1, 2, 3}},
{"object", {{"key", "value"}}}
};
std::vector<uint8_t> temp = json::to_msgpack(original);
std::vector<std::byte> msgpack_data(temp.size());
for (size_t i = 0; i < temp.size(); ++i)
{
msgpack_data[i] = std::byte(temp[i]);
}
// Attempt direct deserialization using std::byte input
const json direct_result = json::from_msgpack(msgpack_data);
// Test the workaround approach: reinterpret as unsigned char* and use iterator range
const auto* const char_start = reinterpret_cast<unsigned char const*>(msgpack_data.data());
const auto* const char_end = char_start + msgpack_data.size();
json workaround_result = json::from_msgpack(char_start, char_end);
// Verify that the final deserialized JSON matches the original JSON
CHECK(direct_result == workaround_result);
CHECK(direct_result == original);
}
}
}
#endif
// the fake sizes below do not fit into a 32-bit std::size_t // the fake sizes below do not fit into a 32-bit std::size_t
// with clang and libstdc++ 10, the std::filesystem::path conversion that // with clang and libstdc++ 10, the std::filesystem::path conversion that
// C++17 builds consider for every string type is ambiguous for a class // C++17 builds consider for every string type is ambiguous for a class
File diff suppressed because it is too large Load Diff
+623
View File
@@ -0,0 +1,623 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale>
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <fstream>
#include <sstream>
#include <iomanip>
#include "make_test_data_available.hpp"
#include "test_utils.hpp"
TEST_CASE("Unicode (1/5)" * doctest::skip())
{
SECTION("\\uxxxx sequences")
{
// create an escaped string from a code point
const auto codepoint_to_unicode = [](std::size_t cp)
{
// code points are represented as a six-character sequence: a
// reverse solidus, followed by the lowercase letter u, followed
// by four hexadecimal digits that encode the character's code
// point
std::stringstream ss;
ss << "\\u" << std::setw(4) << std::setfill('0') << std::hex << cp;
return ss.str();
};
SECTION("correct sequences")
{
// generate all UTF-8 code points; in total, 1112064 code points are
// generated: 0x1FFFFF code points - 2048 invalid values between
// 0xD800 and 0xDFFF.
for (std::size_t cp = 0; cp <= 0x10FFFFu; ++cp)
{
// string to store the code point as in \uxxxx format
std::string json_text = "\"";
// decide whether to use one or two \uxxxx sequences
if (cp < 0x10000u)
{
// The Unicode standard permanently reserves these code point
// values for UTF-16 encoding of the high and low surrogates, and
// they will never be assigned a character, so there should be no
// reason to encode them. The official Unicode standard says that
// no UTF forms, including UTF-16, can encode these code points.
if (cp >= 0xD800u && cp <= 0xDFFFu)
{
// if we would not skip these code points, we would get a
// "missing low surrogate" exception
continue;
}
// code points in the Basic Multilingual Plane can be
// represented with one \uxxxx sequence
json_text += codepoint_to_unicode(cp);
}
else
{
// To escape an extended character that is not in the Basic
// Multilingual Plane, the character is represented as a
// 12-character sequence, encoding the UTF-16 surrogate pair
const auto codepoint1 = 0xd800u + (((cp - 0x10000u) >> 10) & 0x3ffu);
const auto codepoint2 = 0xdc00u + ((cp - 0x10000u) & 0x3ffu);
json_text += codepoint_to_unicode(codepoint1) + codepoint_to_unicode(codepoint2);
}
json_text += "\"";
CAPTURE(json_text)
json _;
CHECK_NOTHROW(_ = json::parse(json_text));
}
}
SECTION("incorrect sequences")
{
SECTION("incorrect surrogate values")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uDC00\\uDC00\""), "[json.exception.parse_error.101] parse error at line 1, column 7: syntax error while parsing value - invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF; last read: '\"\\uDC00'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD7FF\\uDC00\""), "[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF; last read: '\"\\uD7FF\\uDC00'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD800]\""), "[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD800]'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD800\\v\""), "[json.exception.parse_error.101] parse error at line 1, column 9: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD800\\v'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD800\\u123\""), "[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\uD800\\u123\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD800\\uDBFF\""), "[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD800\\uDBFF'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD800\\uE000\""), "[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD800\\uE000'", json::parse_error&);
}
}
#if 0 // NOLINT(readability-avoid-unconditional-preprocessor-if)
SECTION("incorrect sequences")
{
SECTION("high surrogate without low surrogate")
{
// D800..DBFF are high surrogates and must be followed by low
// surrogates DC00..DFFF; here, nothing follows
for (std::size_t cp = 0xD800u; cp <= 0xDBFFu; ++cp)
{
std::string json_text = "\"" + codepoint_to_unicode(cp) + "\"";
CAPTURE(json_text)
CHECK_THROWS_AS(json::parse(json_text), json::parse_error&);
}
}
SECTION("high surrogate with wrong low surrogate")
{
// D800..DBFF are high surrogates and must be followed by low
// surrogates DC00..DFFF; here a different sequence follows
for (std::size_t cp1 = 0xD800u; cp1 <= 0xDBFFu; ++cp1)
{
for (std::size_t cp2 = 0x0000u; cp2 <= 0xFFFFu; ++cp2)
{
if (0xDC00u <= cp2 && cp2 <= 0xDFFFu)
{
continue;
}
std::string json_text = "\"" + codepoint_to_unicode(cp1) + codepoint_to_unicode(cp2) + "\"";
CAPTURE(json_text)
CHECK_THROWS_AS(json::parse(json_text), json::parse_error&);
}
}
}
SECTION("low surrogate without high surrogate")
{
// low surrogates DC00..DFFF must follow high surrogates; here,
// they occur alone
for (std::size_t cp = 0xDC00u; cp <= 0xDFFFu; ++cp)
{
std::string json_text = "\"" + codepoint_to_unicode(cp) + "\"";
CAPTURE(json_text)
CHECK_THROWS_AS(json::parse(json_text), json::parse_error&);
}
}
}
#endif
}
SECTION("read all unicode characters")
{
// read a file with all Unicode characters stored as single-character
// strings in a JSON array
std::ifstream f(TEST_DATA_DIRECTORY "/json_nlohmann_tests/all_unicode.json");
json j;
CHECK_NOTHROW(f >> j);
// the array has 1112064 + 1 elements (a terminating "null" value)
// Note: 1112064 = 0x1FFFFF code points - 2048 invalid values between
// 0xD800 and 0xDFFF.
CHECK(j.size() == 1112065);
SECTION("check JSON Pointers")
{
for (const auto& s : j)
{
// skip non-string JSON values
if (!s.is_string())
{
continue;
}
auto ptr = s.get<std::string>();
// tilde must be followed by 0 or 1
if (ptr == "~")
{
ptr += "0";
}
// JSON Pointers must begin with "/"
ptr.insert(0, "/");
CHECK_NOTHROW(json::json_pointer("/" + ptr));
// check escape/unescape roundtrip
auto escaped = nlohmann::detail::escape(ptr);
nlohmann::detail::unescape(escaped);
CHECK(escaped == ptr);
}
}
}
SECTION("ignore byte-order-mark")
{
SECTION("in a stream")
{
// read a file with a UTF-8 BOM
std::ifstream f(TEST_DATA_DIRECTORY "/json_nlohmann_tests/bom.json");
json j;
CHECK_NOTHROW(f >> j);
}
SECTION("with an iterator")
{
std::string i = "\xef\xbb\xbf{\n \"foo\": true\n}";
json _;
CHECK_NOTHROW(_ = json::parse(i.begin(), i.end()));
}
}
SECTION("error for incomplete/wrong BOM")
{
json _;
CHECK_THROWS_AS(_ = json::parse("\xef\xbb"), json::parse_error&);
CHECK_THROWS_AS(_ = json::parse("\xef\xbb\xbb"), json::parse_error&);
}
}
namespace
{
void roundtrip(bool success_expected, const std::string& s);
void roundtrip(bool success_expected, const std::string& s)
{
CAPTURE(s)
json _;
// create JSON string value
const json j = s;
// create JSON text
const std::string ps = std::string("\"") + s + "\"";
if (success_expected)
{
// serialization succeeds
// dump() is nodiscard; this only checks that dumping does not throw
CHECK_NOTHROW(utils::ignore_return_value(j.dump()));
// exclude parse test for U+0000
if (s[0] != '\0')
{
// parsing JSON text succeeds
CHECK_NOTHROW(_ = json::parse(ps));
}
// roundtrip succeeds
CHECK_NOTHROW(_ = json::parse(j.dump()));
// after roundtrip, the same string is stored
const json jr = json::parse(j.dump());
CHECK(jr.get<std::string>() == s);
}
else
{
// serialization fails
// dump() is nodiscard; the exception is thrown by dump() itself before it would return
CHECK_THROWS_AS(utils::ignore_return_value(j.dump()), json::type_error&);
// parsing JSON text fails
CHECK_THROWS_AS(_ = json::parse(ps), json::parse_error&);
}
}
} // namespace
TEST_CASE("Markus Kuhn's UTF-8 decoder capability and stress test")
{
// Markus Kuhn <http://www.cl.cam.ac.uk/~mgk25/> - 2015-08-28 - CC BY 4.0
// http://www.cl.cam.ac.uk/~mgk25/ucs/examples/UTF-8-test.txt
SECTION("1 Some correct UTF-8 text")
{
roundtrip(true, "κόσμε");
}
SECTION("2 Boundary condition test cases")
{
SECTION("2.1 First possible sequence of a certain length")
{
// 2.1.1 1 byte (U-00000000)
roundtrip(true, std::string("\0", 1));
// 2.1.2 2 bytes (U-00000080)
roundtrip(true, "\xc2\x80");
// 2.1.3 3 bytes (U-00000800)
roundtrip(true, "\xe0\xa0\x80");
// 2.1.4 4 bytes (U-00010000)
roundtrip(true, "\xf0\x90\x80\x80");
// 2.1.5 5 bytes (U-00200000)
roundtrip(false, "\xF8\x88\x80\x80\x80");
// 2.1.6 6 bytes (U-04000000)
roundtrip(false, "\xFC\x84\x80\x80\x80\x80");
}
SECTION("2.2 Last possible sequence of a certain length")
{
// 2.2.1 1 byte (U-0000007F)
roundtrip(true, "\x7f");
// 2.2.2 2 bytes (U-000007FF)
roundtrip(true, "\xdf\xbf");
// 2.2.3 3 bytes (U-0000FFFF)
roundtrip(true, "\xef\xbf\xbf");
// 2.2.4 4 bytes (U-001FFFFF)
roundtrip(false, "\xF7\xBF\xBF\xBF");
// 2.2.5 5 bytes (U-03FFFFFF)
roundtrip(false, "\xFB\xBF\xBF\xBF\xBF");
// 2.2.6 6 bytes (U-7FFFFFFF)
roundtrip(false, "\xFD\xBF\xBF\xBF\xBF\xBF");
}
SECTION("2.3 Other boundary conditions")
{
// 2.3.1 U-0000D7FF = ed 9f bf
roundtrip(true, "\xed\x9f\xbf");
// 2.3.2 U-0000E000 = ee 80 80
roundtrip(true, "\xee\x80\x80");
// 2.3.3 U-0000FFFD = ef bf bd
roundtrip(true, "\xef\xbf\xbd");
// 2.3.4 U-0010FFFF = f4 8f bf bf
roundtrip(true, "\xf4\x8f\xbf\xbf");
// 2.3.5 U-00110000 = f4 90 80 80
roundtrip(false, "\xf4\x90\x80\x80");
}
}
SECTION("3 Malformed sequences")
{
SECTION("3.1 Unexpected continuation bytes")
{
// Each unexpected continuation byte should be separately signalled as a
// malformed sequence of its own.
// 3.1.1 First continuation byte 0x80
roundtrip(false, "\x80");
// 3.1.2 Last continuation byte 0xbf
roundtrip(false, "\xbf");
// 3.1.3 2 continuation bytes
roundtrip(false, "\x80\xbf");
// 3.1.4 3 continuation bytes
roundtrip(false, "\x80\xbf\x80");
// 3.1.5 4 continuation bytes
roundtrip(false, "\x80\xbf\x80\xbf");
// 3.1.6 5 continuation bytes
roundtrip(false, "\x80\xbf\x80\xbf\x80");
// 3.1.7 6 continuation bytes
roundtrip(false, "\x80\xbf\x80\xbf\x80\xbf");
// 3.1.8 7 continuation bytes
roundtrip(false, "\x80\xbf\x80\xbf\x80\xbf\x80");
// 3.1.9 Sequence of all 64 possible continuation bytes (0x80-0xbf)
roundtrip(false, "\x80\x81\x82\x83\x84\x85\x86\x87\x88\x89\x8a\x8b\x8c\x8d\x8e\x8f\x90\x91\x92\x93\x94\x95\x96\x97\x98\x99\x9a\x9b\x9c\x9d\x9e\x9f\xa0\xa1\xa2\xa3\xa4\xa5\xa6\xa7\xa8\xa9\xaa\xab\xac\xad\xae\xaf\xb0\xb1\xb2\xb3\xb4\xb5\xb6\xb7\xb8\xb9\xba\xbb\xbc\xbd\xbe\xbf");
}
SECTION("3.2 Lonely start characters")
{
// 3.2.1 All 32 first bytes of 2-byte sequences (0xc0-0xdf)
roundtrip(false, "\xc0 \xc1 \xc2 \xc3 \xc4 \xc5 \xc6 \xc7 \xc8 \xc9 \xca \xcb \xcc \xcd \xce \xcf \xd0 \xd1 \xd2 \xd3 \xd4 \xd5 \xd6 \xd7 \xd8 \xd9 \xda \xdb \xdc \xdd \xde \xdf");
// 3.2.2 All 16 first bytes of 3-byte sequences (0xe0-0xef)
roundtrip(false, "\xe0 \xe1 \xe2 \xe3 \xe4 \xe5 \xe6 \xe7 \xe8 \xe9 \xea \xeb \xec \xed \xee \xef");
// 3.2.3 All 8 first bytes of 4-byte sequences (0xf0-0xf7)
roundtrip(false, "\xf0 \xf1 \xf2 \xf3 \xf4 \xf5 \xf6 \xf7");
// 3.2.4 All 4 first bytes of 5-byte sequences (0xf8-0xfb)
roundtrip(false, "\xf8 \xf9 \xfa \xfb");
// 3.2.5 All 2 first bytes of 6-byte sequences (0xfc-0xfd)
roundtrip(false, "\xfc \xfd");
}
SECTION("3.3 Sequences with last continuation byte missing")
{
// All bytes of an incomplete sequence should be signalled as a single
// malformed sequence, i.e., you should see only a single replacement
// character in each of the next 10 tests. (Characters as in section 2)
// 3.3.1 2-byte sequence with last byte missing (U+0000)
roundtrip(false, "\xc0");
// 3.3.2 3-byte sequence with last byte missing (U+0000)
roundtrip(false, "\xe0\x80");
// 3.3.3 4-byte sequence with last byte missing (U+0000)
roundtrip(false, "\xf0\x80\x80");
// 3.3.4 5-byte sequence with last byte missing (U+0000)
roundtrip(false, "\xf8\x80\x80\x80");
// 3.3.5 6-byte sequence with last byte missing (U+0000)
roundtrip(false, "\xfc\x80\x80\x80\x80");
// 3.3.6 2-byte sequence with last byte missing (U-000007FF)
roundtrip(false, "\xdf");
// 3.3.7 3-byte sequence with last byte missing (U-0000FFFF)
roundtrip(false, "\xef\xbf");
// 3.3.8 4-byte sequence with last byte missing (U-001FFFFF)
roundtrip(false, "\xf7\xbf\xbf");
// 3.3.9 5-byte sequence with last byte missing (U-03FFFFFF)
roundtrip(false, "\xfb\xbf\xbf\xbf");
// 3.3.10 6-byte sequence with last byte missing (U-7FFFFFFF)
roundtrip(false, "\xfd\xbf\xbf\xbf\xbf");
}
SECTION("3.4 Concatenation of incomplete sequences")
{
// All the 10 sequences of 3.3 concatenated, you should see 10 malformed
// sequences being signalled:
roundtrip(false, "\xc0\xe0\x80\xf0\x80\x80\xf8\x80\x80\x80\xfc\x80\x80\x80\x80\xdf\xef\xbf\xf7\xbf\xbf\xfb\xbf\xbf\xbf\xfd\xbf\xbf\xbf\xbf");
}
SECTION("3.5 Impossible bytes")
{
// The following two bytes cannot appear in a correct UTF-8 string
// 3.5.1 fe
roundtrip(false, "\xfe");
// 3.5.2 ff
roundtrip(false, "\xff");
// 3.5.3 fe fe ff ff
roundtrip(false, "\xfe\xfe\xff\xff");
}
}
SECTION("4 Overlong sequences")
{
// The following sequences are not malformed according to the letter of
// the Unicode 2.0 standard. However, they are longer then necessary and
// a correct UTF-8 encoder is not allowed to produce them. A "safe UTF-8
// decoder" should reject them just like malformed sequences for two
// reasons: (1) It helps to debug applications if overlong sequences are
// not treated as valid representations of characters, because this helps
// to spot problems more quickly. (2) Overlong sequences provide
// alternative representations of characters, that could maliciously be
// used to bypass filters that check only for ASCII characters. For
// instance, a 2-byte encoded line feed (LF) would not be caught by a
// line counter that counts only 0x0a bytes, but it would still be
// processed as a line feed by an unsafe UTF-8 decoder later in the
// pipeline. From a security point of view, ASCII compatibility of UTF-8
// sequences means also, that ASCII characters are *only* allowed to be
// represented by ASCII bytes in the range 0x00-0x7f. To ensure this
// aspect of ASCII compatibility, use only "safe UTF-8 decoders" that
// reject overlong UTF-8 sequences for which a shorter encoding exists.
SECTION("4.1 Examples of an overlong ASCII character")
{
// With a safe UTF-8 decoder, all the following five overlong
// representations of the ASCII character slash ("/") should be rejected
// like a malformed UTF-8 sequence, for instance by substituting it with
// a replacement character. If you see a slash below, you do not have a
// safe UTF-8 decoder!
// 4.1.1 U+002F = c0 af
roundtrip(false, "\xc0\xaf");
// 4.1.2 U+002F = e0 80 af
roundtrip(false, "\xe0\x80\xaf");
// 4.1.3 U+002F = f0 80 80 af
roundtrip(false, "\xf0\x80\x80\xaf");
// 4.1.4 U+002F = f8 80 80 80 af
roundtrip(false, "\xf8\x80\x80\x80\xaf");
// 4.1.5 U+002F = fc 80 80 80 80 af
roundtrip(false, "\xfc\x80\x80\x80\x80\xaf");
}
SECTION("4.2 Maximum overlong sequences")
{
// Below you see the highest Unicode value that is still resulting in an
// overlong sequence if represented with the given number of bytes. This
// is a boundary test for safe UTF-8 decoders. All five characters should
// be rejected like malformed UTF-8 sequences.
// 4.2.1 U-0000007F = c1 bf
roundtrip(false, "\xc1\xbf");
// 4.2.2 U-000007FF = e0 9f bf
roundtrip(false, "\xe0\x9f\xbf");
// 4.2.3 U-0000FFFF = f0 8f bf bf
roundtrip(false, "\xf0\x8f\xbf\xbf");
// 4.2.4 U-001FFFFF = f8 87 bf bf bf
roundtrip(false, "\xf8\x87\xbf\xbf\xbf");
// 4.2.5 U-03FFFFFF = fc 83 bf bf bf bf
roundtrip(false, "\xfc\x83\xbf\xbf\xbf\xbf");
}
SECTION("4.3 Overlong representation of the NUL character")
{
// The following five sequences should also be rejected like malformed
// UTF-8 sequences and should not be treated like the ASCII NUL
// character.
// 4.3.1 U+0000 = c0 80
roundtrip(false, "\xc0\x80");
// 4.3.2 U+0000 = e0 80 80
roundtrip(false, "\xe0\x80\x80");
// 4.3.3 U+0000 = f0 80 80 80
roundtrip(false, "\xf0\x80\x80\x80");
// 4.3.4 U+0000 = f8 80 80 80 80
roundtrip(false, "\xf8\x80\x80\x80\x80");
// 4.3.5 U+0000 = fc 80 80 80 80 80
roundtrip(false, "\xfc\x80\x80\x80\x80\x80");
}
}
SECTION("5 Illegal code positions")
{
// The following UTF-8 sequences should be rejected like malformed
// sequences, because they never represent valid ISO 10646 characters and
// a UTF-8 decoder that accepts them might introduce security problems
// comparable to overlong UTF-8 sequences.
SECTION("5.1 Single UTF-16 surrogates")
{
// 5.1.1 U+D800 = ed a0 80
roundtrip(false, "\xed\xa0\x80");
// 5.1.2 U+DB7F = ed ad bf
roundtrip(false, "\xed\xad\xbf");
// 5.1.3 U+DB80 = ed ae 80
roundtrip(false, "\xed\xae\x80");
// 5.1.4 U+DBFF = ed af bf
roundtrip(false, "\xed\xaf\xbf");
// 5.1.5 U+DC00 = ed b0 80
roundtrip(false, "\xed\xb0\x80");
// 5.1.6 U+DF80 = ed be 80
roundtrip(false, "\xed\xbe\x80");
// 5.1.7 U+DFFF = ed bf bf
roundtrip(false, "\xed\xbf\xbf");
}
SECTION("5.2 Paired UTF-16 surrogates")
{
// 5.2.1 U+D800 U+DC00 = ed a0 80 ed b0 80
roundtrip(false, "\xed\xa0\x80\xed\xb0\x80");
// 5.2.2 U+D800 U+DFFF = ed a0 80 ed bf bf
roundtrip(false, "\xed\xa0\x80\xed\xbf\xbf");
// 5.2.3 U+DB7F U+DC00 = ed ad bf ed b0 80
roundtrip(false, "\xed\xad\xbf\xed\xb0\x80");
// 5.2.4 U+DB7F U+DFFF = ed ad bf ed bf bf
roundtrip(false, "\xed\xad\xbf\xed\xbf\xbf");
// 5.2.5 U+DB80 U+DC00 = ed ae 80 ed b0 80
roundtrip(false, "\xed\xae\x80\xed\xb0\x80");
// 5.2.6 U+DB80 U+DFFF = ed ae 80 ed bf bf
roundtrip(false, "\xed\xae\x80\xed\xbf\xbf");
// 5.2.7 U+DBFF U+DC00 = ed af bf ed b0 80
roundtrip(false, "\xed\xaf\xbf\xed\xb0\x80");
// 5.2.8 U+DBFF U+DFFF = ed af bf ed bf bf
roundtrip(false, "\xed\xaf\xbf\xed\xbf\xbf");
}
SECTION("5.3 Noncharacter code positions")
{
// The following "noncharacters" are "reserved for internal use" by
// applications, and according to older versions of the Unicode Standard
// "should never be interchanged". Unicode Corrigendum #9 dropped the
// latter restriction. Nevertheless, their presence in incoming UTF-8 data
// can remain a potential security risk, depending on what use is made of
// these codes subsequently. Examples of such internal use:
//
// - Some file APIs with 16-bit characters may use the integer value -1
// = U+FFFF to signal an end-of-file (EOF) or error condition.
//
// - In some UTF-16 receivers, code point U+FFFE might trigger a
// byte-swap operation (to convert between UTF-16LE and UTF-16BE).
//
// With such internal use of noncharacters, it may be desirable and safer
// to block those code points in UTF-8 decoders, as they should never
// occur legitimately in incoming UTF-8 data, and could trigger unsafe
// behaviour in subsequent processing.
// Particularly problematic noncharacters in 16-bit applications:
// 5.3.1 U+FFFE = ef bf be
roundtrip(true, "\xef\xbf\xbe");
// 5.3.2 U+FFFF = ef bf bf
roundtrip(true, "\xef\xbf\xbf");
// 5.3.3 U+FDD0 .. U+FDEF
roundtrip(true, "\xEF\xB7\x90");
roundtrip(true, "\xEF\xB7\x91");
roundtrip(true, "\xEF\xB7\x92");
roundtrip(true, "\xEF\xB7\x93");
roundtrip(true, "\xEF\xB7\x94");
roundtrip(true, "\xEF\xB7\x95");
roundtrip(true, "\xEF\xB7\x96");
roundtrip(true, "\xEF\xB7\x97");
roundtrip(true, "\xEF\xB7\x98");
roundtrip(true, "\xEF\xB7\x99");
roundtrip(true, "\xEF\xB7\x9A");
roundtrip(true, "\xEF\xB7\x9B");
roundtrip(true, "\xEF\xB7\x9C");
roundtrip(true, "\xEF\xB7\x9D");
roundtrip(true, "\xEF\xB7\x9E");
roundtrip(true, "\xEF\xB7\x9F");
roundtrip(true, "\xEF\xB7\xA0");
roundtrip(true, "\xEF\xB7\xA1");
roundtrip(true, "\xEF\xB7\xA2");
roundtrip(true, "\xEF\xB7\xA3");
roundtrip(true, "\xEF\xB7\xA4");
roundtrip(true, "\xEF\xB7\xA5");
roundtrip(true, "\xEF\xB7\xA6");
roundtrip(true, "\xEF\xB7\xA7");
roundtrip(true, "\xEF\xB7\xA8");
roundtrip(true, "\xEF\xB7\xA9");
roundtrip(true, "\xEF\xB7\xAA");
roundtrip(true, "\xEF\xB7\xAB");
roundtrip(true, "\xEF\xB7\xAC");
roundtrip(true, "\xEF\xB7\xAD");
roundtrip(true, "\xEF\xB7\xAE");
roundtrip(true, "\xEF\xB7\xAF");
// 5.3.4 U+nFFFE U+nFFFF (for n = 1..10)
roundtrip(true, "\xF0\x9F\xBF\xBF");
roundtrip(true, "\xF0\xAF\xBF\xBF");
roundtrip(true, "\xF0\xBF\xBF\xBF");
roundtrip(true, "\xF1\x8F\xBF\xBF");
roundtrip(true, "\xF1\x9F\xBF\xBF");
roundtrip(true, "\xF1\xAF\xBF\xBF");
roundtrip(true, "\xF1\xBF\xBF\xBF");
roundtrip(true, "\xF2\x8F\xBF\xBF");
roundtrip(true, "\xF2\x9F\xBF\xBF");
roundtrip(true, "\xF2\xAF\xBF\xBF");
}
}
}
+612
View File
@@ -0,0 +1,612 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale>
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <fstream>
#include <sstream>
#include <iostream>
#include <iomanip>
#include "make_test_data_available.hpp"
#include "test_utils.hpp"
// this test suite uses static variables with non-trivial destructors
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wexit-time-destructors")
namespace
{
extern size_t calls;
size_t calls = 0;
void check_utf8dump(bool success_expected, int byte1, int byte2, int byte3, int byte4);
void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
static std::string json_string;
json_string.clear();
CAPTURE(byte1)
CAPTURE(byte2)
CAPTURE(byte3)
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
json_string += std::string(1, static_cast<char>(byte4));
}
CAPTURE(json_string)
// store the string in a JSON value
static json j;
static json j2;
j = json_string;
j2 = "abc" + json_string + "xyz";
static std::string s_ignored;
static std::string s_ignored2;
static std::string s_ignored_ascii;
static std::string s_ignored2_ascii;
static std::string s_replaced;
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
s_ignored2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::ignore);
s_replaced = j.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
if (success_expected)
{
static std::string s_strict;
// strict mode must not throw if success is expected
s_strict = j.dump();
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
}
else
{
// strict mode must throw if success is not expected
// dump() is nodiscard; the exception is thrown by dump() itself before it would return
CHECK_THROWS_AS(utils::ignore_return_value(j.dump()), json::type_error&);
// ignore and replace must create different dumps
CHECK(s_ignored != s_replaced);
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
}
// check that prefix and suffix are preserved
CHECK(s_ignored2.substr(1, 3) == "abc");
CHECK(s_ignored2.substr(s_ignored2.size() - 4, 3) == "xyz");
CHECK(s_ignored2_ascii.substr(1, 3) == "abc");
CHECK(s_ignored2_ascii.substr(s_ignored2_ascii.size() - 4, 3) == "xyz");
CHECK(s_replaced2.substr(1, 3) == "abc");
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
// create and check a JSON string with up to four UTF-8 bytes
void check_utf8string(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
if (++calls % 100000 == 0)
{
std::cout << calls << " of 455355 UTF-8 strings checked" << std::endl; // NOLINT(performance-avoid-endl)
}
static std::string json_string;
json_string = "\"";
CAPTURE(byte1)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
CAPTURE(byte2)
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
CAPTURE(byte3)
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte4));
}
json_string += "\"";
CAPTURE(json_string)
json _;
if (success_expected)
{
CHECK_NOTHROW(_ = json::parse(json_string));
}
else
{
CHECK_THROWS_AS(_ = json::parse(json_string), json::parse_error&);
}
}
} // namespace
TEST_CASE("Unicode (2/5)" * doctest::skip())
{
SECTION("RFC 3629")
{
/*
RFC 3629 describes in Sect. 4 the syntax of UTF-8 byte sequences as
follows:
A UTF-8 string is a sequence of octets representing a sequence of UCS
characters. An octet sequence is valid UTF-8 only if it matches the
following syntax, which is derived from the rules for encoding UTF-8
and is expressed in the ABNF of [RFC2234].
UTF8-octets = *( UTF8-char )
UTF8-char = UTF8-1 / UTF8-2 / UTF8-3 / UTF8-4
UTF8-1 = %x00-7F
UTF8-2 = %xC2-DF UTF8-tail
UTF8-3 = %xE0 %xA0-BF UTF8-tail / %xE1-EC 2( UTF8-tail ) /
%xED %x80-9F UTF8-tail / %xEE-EF 2( UTF8-tail )
UTF8-4 = %xF0 %x90-BF 2( UTF8-tail ) / %xF1-F3 3( UTF8-tail ) /
%xF4 %x80-8F 2( UTF8-tail )
UTF8-tail = %x80-BF
*/
SECTION("ill-formed first byte")
{
for (int byte1 = 0x80; byte1 <= 0xC1; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
for (int byte1 = 0xF5; byte1 <= 0xFF; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("UTF8-1 (x00-x7F)")
{
SECTION("well-formed")
{
for (int byte1 = 0x00; byte1 <= 0x7F; ++byte1)
{
// unescaped control characters are parse errors in JSON
if (0x00 <= byte1 && byte1 <= 0x1F)
{
check_utf8string(false, byte1);
continue;
}
// a single quote is a parse error in JSON
if (byte1 == 0x22)
{
check_utf8string(false, byte1);
continue;
}
// a single backslash is a parse error in JSON
if (byte1 == 0x5C)
{
check_utf8string(false, byte1);
continue;
}
// all other characters are OK
check_utf8string(true, byte1);
check_utf8dump(true, byte1);
}
}
}
SECTION("UTF8-2 (xC2-xDF UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xC2; byte1 <= 0xDF; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
check_utf8string(true, byte1, byte2);
check_utf8dump(true, byte1, byte2);
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xC2; byte1 <= 0xDF; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xC2; byte1 <= 0xDF; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x80 <= byte2 && byte2 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
}
SECTION("UTF8-3 (xE0 xA0-BF UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xE0; byte1 <= 0xE0; ++byte1)
{
for (int byte2 = 0xA0; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(true, byte1, byte2, byte3);
check_utf8dump(true, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xE0; byte1 <= 0xE0; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xE0; byte1 <= 0xE0; ++byte1)
{
for (int byte2 = 0xA0; byte2 <= 0xBF; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xE0; byte1 <= 0xE0; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0xA0 <= byte2 && byte2 <= 0xBF)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xE0; byte1 <= 0xE0; ++byte1)
{
for (int byte2 = 0xA0; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
}
SECTION("UTF8-3 (xE1-xEC UTF8-tail UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xE1; byte1 <= 0xEC; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(true, byte1, byte2, byte3);
check_utf8dump(true, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xE1; byte1 <= 0xEC; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xE1; byte1 <= 0xEC; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xE1; byte1 <= 0xEC; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x80 <= byte2 && byte2 <= 0xBF)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xE1; byte1 <= 0xEC; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
}
SECTION("UTF8-3 (xED x80-9F UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xED; byte1 <= 0xED; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x9F; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(true, byte1, byte2, byte3);
check_utf8dump(true, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xED; byte1 <= 0xED; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xED; byte1 <= 0xED; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x9F; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xED; byte1 <= 0xED; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x80 <= byte2 && byte2 <= 0x9F)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xED; byte1 <= 0xED; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x9F; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
}
SECTION("UTF8-3 (xEE-xEF UTF8-tail UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xEE; byte1 <= 0xEF; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(true, byte1, byte2, byte3);
check_utf8dump(true, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xEE; byte1 <= 0xEF; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xEE; byte1 <= 0xEF; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xEE; byte1 <= 0xEF; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x80 <= byte2 && byte2 <= 0xBF)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xEE; byte1 <= 0xEF; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
}
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+326
View File
@@ -0,0 +1,326 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale>
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <fstream>
#include <sstream>
#include <iostream>
#include <iomanip>
#include "make_test_data_available.hpp"
#include "test_utils.hpp"
// this test suite uses static variables with non-trivial destructors
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wexit-time-destructors")
namespace
{
extern size_t calls;
size_t calls = 0;
void check_utf8dump(bool success_expected, int byte1, int byte2, int byte3, int byte4);
void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
static std::string json_string;
json_string.clear();
CAPTURE(byte1)
CAPTURE(byte2)
CAPTURE(byte3)
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
json_string += std::string(1, static_cast<char>(byte4));
}
CAPTURE(json_string)
// store the string in a JSON value
static json j;
static json j2;
j = json_string;
j2 = "abc" + json_string + "xyz";
static std::string s_ignored;
static std::string s_ignored2;
static std::string s_ignored_ascii;
static std::string s_ignored2_ascii;
static std::string s_replaced;
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
s_ignored2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::ignore);
s_replaced = j.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
if (success_expected)
{
static std::string s_strict;
// strict mode must not throw if success is expected
s_strict = j.dump();
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
}
else
{
// strict mode must throw if success is not expected
// dump() is nodiscard; the exception is thrown by dump() itself before it would return
CHECK_THROWS_AS(utils::ignore_return_value(j.dump()), json::type_error&);
// ignore and replace must create different dumps
CHECK(s_ignored != s_replaced);
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
}
// check that prefix and suffix are preserved
CHECK(s_ignored2.substr(1, 3) == "abc");
CHECK(s_ignored2.substr(s_ignored2.size() - 4, 3) == "xyz");
CHECK(s_ignored2_ascii.substr(1, 3) == "abc");
CHECK(s_ignored2_ascii.substr(s_ignored2_ascii.size() - 4, 3) == "xyz");
CHECK(s_replaced2.substr(1, 3) == "abc");
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
// create and check a JSON string with up to four UTF-8 bytes
void check_utf8string(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
if (++calls % 100000 == 0)
{
std::cout << calls << " of 1641521 UTF-8 strings checked" << std::endl; // NOLINT(performance-avoid-endl)
}
static std::string json_string;
json_string = "\"";
CAPTURE(byte1)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
CAPTURE(byte2)
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
CAPTURE(byte3)
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte4));
}
json_string += "\"";
CAPTURE(json_string)
json _;
if (success_expected)
{
CHECK_NOTHROW(_ = json::parse(json_string));
}
else
{
CHECK_THROWS_AS(_ = json::parse(json_string), json::parse_error&);
}
}
} // namespace
TEST_CASE("Unicode (3/5)" * doctest::skip())
{
SECTION("RFC 3629")
{
/*
RFC 3629 describes in Sect. 4 the syntax of UTF-8 byte sequences as
follows:
A UTF-8 string is a sequence of octets representing a sequence of UCS
characters. An octet sequence is valid UTF-8 only if it matches the
following syntax, which is derived from the rules for encoding UTF-8
and is expressed in the ABNF of [RFC2234].
UTF8-octets = *( UTF8-char )
UTF8-char = UTF8-1 / UTF8-2 / UTF8-3 / UTF8-4
UTF8-1 = %x00-7F
UTF8-2 = %xC2-DF UTF8-tail
UTF8-3 = %xE0 %xA0-BF UTF8-tail / %xE1-EC 2( UTF8-tail ) /
%xED %x80-9F UTF8-tail / %xEE-EF 2( UTF8-tail )
UTF8-4 = %xF0 %x90-BF 2( UTF8-tail ) / %xF1-F3 3( UTF8-tail ) /
%xF4 %x80-8F 2( UTF8-tail )
UTF8-tail = %x80-BF
*/
SECTION("UTF8-4 (xF0 x90-BF UTF8-tail UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
for (int byte2 = 0x90; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(true, byte1, byte2, byte3, byte4);
check_utf8dump(true, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
for (int byte2 = 0x90; byte2 <= 0xBF; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: missing fourth byte")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
for (int byte2 = 0x90; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x90 <= byte2 && byte2 <= 0xBF)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
for (int byte2 = 0x90; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: wrong fourth byte")
{
for (int byte1 = 0xF0; byte1 <= 0xF0; ++byte1)
{
for (int byte2 = 0x90; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x00; byte4 <= 0xFF; ++byte4)
{
// skip correct fourth byte
if (0x80 <= byte4 && byte4 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
}
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+326
View File
@@ -0,0 +1,326 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale>
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <fstream>
#include <sstream>
#include <iostream>
#include <iomanip>
#include "make_test_data_available.hpp"
#include "test_utils.hpp"
// this test suite uses static variables with non-trivial destructors
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wexit-time-destructors")
namespace
{
extern size_t calls;
size_t calls = 0;
void check_utf8dump(bool success_expected, int byte1, int byte2, int byte3, int byte4);
void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
static std::string json_string;
json_string.clear();
CAPTURE(byte1)
CAPTURE(byte2)
CAPTURE(byte3)
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
json_string += std::string(1, static_cast<char>(byte4));
}
CAPTURE(json_string)
// store the string in a JSON value
static json j;
static json j2;
j = json_string;
j2 = "abc" + json_string + "xyz";
static std::string s_ignored;
static std::string s_ignored2;
static std::string s_ignored_ascii;
static std::string s_ignored2_ascii;
static std::string s_replaced;
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
s_ignored2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::ignore);
s_replaced = j.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
if (success_expected)
{
static std::string s_strict;
// strict mode must not throw if success is expected
s_strict = j.dump();
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
}
else
{
// strict mode must throw if success is not expected
// dump() is nodiscard; the exception is thrown by dump() itself before it would return
CHECK_THROWS_AS(utils::ignore_return_value(j.dump()), json::type_error&);
// ignore and replace must create different dumps
CHECK(s_ignored != s_replaced);
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
}
// check that prefix and suffix are preserved
CHECK(s_ignored2.substr(1, 3) == "abc");
CHECK(s_ignored2.substr(s_ignored2.size() - 4, 3) == "xyz");
CHECK(s_ignored2_ascii.substr(1, 3) == "abc");
CHECK(s_ignored2_ascii.substr(s_ignored2_ascii.size() - 4, 3) == "xyz");
CHECK(s_replaced2.substr(1, 3) == "abc");
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
// create and check a JSON string with up to four UTF-8 bytes
void check_utf8string(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
if (++calls % 100000 == 0)
{
std::cout << calls << " of 5517507 UTF-8 strings checked" << std::endl; // NOLINT(performance-avoid-endl)
}
static std::string json_string;
json_string = "\"";
CAPTURE(byte1)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
CAPTURE(byte2)
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
CAPTURE(byte3)
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte4));
}
json_string += "\"";
CAPTURE(json_string)
json _;
if (success_expected)
{
CHECK_NOTHROW(_ = json::parse(json_string));
}
else
{
CHECK_THROWS_AS(_ = json::parse(json_string), json::parse_error&);
}
}
} // namespace
TEST_CASE("Unicode (4/5)" * doctest::skip())
{
SECTION("RFC 3629")
{
/*
RFC 3629 describes in Sect. 4 the syntax of UTF-8 byte sequences as
follows:
A UTF-8 string is a sequence of octets representing a sequence of UCS
characters. An octet sequence is valid UTF-8 only if it matches the
following syntax, which is derived from the rules for encoding UTF-8
and is expressed in the ABNF of [RFC2234].
UTF8-octets = *( UTF8-char )
UTF8-char = UTF8-1 / UTF8-2 / UTF8-3 / UTF8-4
UTF8-1 = %x00-7F
UTF8-2 = %xC2-DF UTF8-tail
UTF8-3 = %xE0 %xA0-BF UTF8-tail / %xE1-EC 2( UTF8-tail ) /
%xED %x80-9F UTF8-tail / %xEE-EF 2( UTF8-tail )
UTF8-4 = %xF0 %x90-BF 2( UTF8-tail ) / %xF1-F3 3( UTF8-tail ) /
%xF4 %x80-8F 2( UTF8-tail )
UTF8-tail = %x80-BF
*/
SECTION("UTF8-4 (xF1-F3 UTF8-tail UTF8-tail UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(true, byte1, byte2, byte3, byte4);
check_utf8dump(true, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: missing fourth byte")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x80 <= byte2 && byte2 <= 0xBF)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: wrong fourth byte")
{
for (int byte1 = 0xF1; byte1 <= 0xF3; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0xBF; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x00; byte4 <= 0xFF; ++byte4)
{
// skip correct fourth byte
if (0x80 <= byte4 && byte4 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
}
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+326
View File
@@ -0,0 +1,326 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale>
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <fstream>
#include <sstream>
#include <iostream>
#include <iomanip>
#include "make_test_data_available.hpp"
#include "test_utils.hpp"
// this test suite uses static variables with non-trivial destructors
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wexit-time-destructors")
namespace
{
extern size_t calls;
size_t calls = 0;
void check_utf8dump(bool success_expected, int byte1, int byte2, int byte3, int byte4);
void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
static std::string json_string;
json_string.clear();
CAPTURE(byte1)
CAPTURE(byte2)
CAPTURE(byte3)
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
json_string += std::string(1, static_cast<char>(byte4));
}
CAPTURE(json_string)
// store the string in a JSON value
static json j;
static json j2;
j = json_string;
j2 = "abc" + json_string + "xyz";
static std::string s_ignored;
static std::string s_ignored2;
static std::string s_ignored_ascii;
static std::string s_ignored2_ascii;
static std::string s_replaced;
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
s_ignored2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::ignore);
s_replaced = j.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
if (success_expected)
{
static std::string s_strict;
// strict mode must not throw if success is expected
s_strict = j.dump();
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
}
else
{
// strict mode must throw if success is not expected
// dump() is nodiscard; the exception is thrown by dump() itself before it would return
CHECK_THROWS_AS(utils::ignore_return_value(j.dump()), json::type_error&);
// ignore and replace must create different dumps
CHECK(s_ignored != s_replaced);
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
}
// check that prefix and suffix are preserved
CHECK(s_ignored2.substr(1, 3) == "abc");
CHECK(s_ignored2.substr(s_ignored2.size() - 4, 3) == "xyz");
CHECK(s_ignored2_ascii.substr(1, 3) == "abc");
CHECK(s_ignored2_ascii.substr(s_ignored2_ascii.size() - 4, 3) == "xyz");
CHECK(s_replaced2.substr(1, 3) == "abc");
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
// create and check a JSON string with up to four UTF-8 bytes
void check_utf8string(bool success_expected, int byte1, int byte2 = -1, int byte3 = -1, int byte4 = -1)
{
if (++calls % 100000 == 0)
{
std::cout << calls << " of 1246225 UTF-8 strings checked" << std::endl; // NOLINT(performance-avoid-endl)
}
static std::string json_string;
json_string = "\"";
CAPTURE(byte1)
json_string += std::string(1, static_cast<char>(byte1));
if (byte2 != -1)
{
CAPTURE(byte2)
json_string += std::string(1, static_cast<char>(byte2));
}
if (byte3 != -1)
{
CAPTURE(byte3)
json_string += std::string(1, static_cast<char>(byte3));
}
if (byte4 != -1)
{
CAPTURE(byte4)
json_string += std::string(1, static_cast<char>(byte4));
}
json_string += "\"";
CAPTURE(json_string)
json _;
if (success_expected)
{
CHECK_NOTHROW(_ = json::parse(json_string));
}
else
{
CHECK_THROWS_AS(_ = json::parse(json_string), json::parse_error&);
}
}
} // namespace
TEST_CASE("Unicode (5/5)" * doctest::skip())
{
SECTION("RFC 3629")
{
/*
RFC 3629 describes in Sect. 4 the syntax of UTF-8 byte sequences as
follows:
A UTF-8 string is a sequence of octets representing a sequence of UCS
characters. An octet sequence is valid UTF-8 only if it matches the
following syntax, which is derived from the rules for encoding UTF-8
and is expressed in the ABNF of [RFC2234].
UTF8-octets = *( UTF8-char )
UTF8-char = UTF8-1 / UTF8-2 / UTF8-3 / UTF8-4
UTF8-1 = %x00-7F
UTF8-2 = %xC2-DF UTF8-tail
UTF8-3 = %xE0 %xA0-BF UTF8-tail / %xE1-EC 2( UTF8-tail ) /
%xED %x80-9F UTF8-tail / %xEE-EF 2( UTF8-tail )
UTF8-4 = %xF0 %x90-BF 2( UTF8-tail ) / %xF1-F3 3( UTF8-tail ) /
%xF4 %x80-8F 2( UTF8-tail )
UTF8-tail = %x80-BF
*/
SECTION("UTF8-4 (xF4 x80-8F UTF8-tail UTF8-tail)")
{
SECTION("well-formed")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x8F; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(true, byte1, byte2, byte3, byte4);
check_utf8dump(true, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: missing second byte")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
check_utf8string(false, byte1);
check_utf8dump(false, byte1);
}
}
SECTION("ill-formed: missing third byte")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x8F; ++byte2)
{
check_utf8string(false, byte1, byte2);
check_utf8dump(false, byte1, byte2);
}
}
}
SECTION("ill-formed: missing fourth byte")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x8F; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
check_utf8string(false, byte1, byte2, byte3);
check_utf8dump(false, byte1, byte2, byte3);
}
}
}
}
SECTION("ill-formed: wrong second byte")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
for (int byte2 = 0x00; byte2 <= 0xFF; ++byte2)
{
// skip correct second byte
if (0x80 <= byte2 && byte2 <= 0x8F)
{
continue;
}
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: wrong third byte")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x8F; ++byte2)
{
for (int byte3 = 0x00; byte3 <= 0xFF; ++byte3)
{
// skip correct third byte
if (0x80 <= byte3 && byte3 <= 0xBF)
{
continue;
}
for (int byte4 = 0x80; byte4 <= 0xBF; ++byte4)
{
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
SECTION("ill-formed: wrong fourth byte")
{
for (int byte1 = 0xF4; byte1 <= 0xF4; ++byte1)
{
for (int byte2 = 0x80; byte2 <= 0x8F; ++byte2)
{
for (int byte3 = 0x80; byte3 <= 0xBF; ++byte3)
{
for (int byte4 = 0x00; byte4 <= 0xFF; ++byte4)
{
// skip correct fourth byte
if (0x80 <= byte4 && byte4 <= 0xBF)
{
continue;
}
check_utf8string(false, byte1, byte2, byte3, byte4);
check_utf8dump(false, byte1, byte2, byte3, byte4);
}
}
}
}
}
}
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP