From 06b0452189ef542ce704a8b4f3b773318d23b361 Mon Sep 17 00:00:00 2001 From: "dependabot[bot]" <49699333+dependabot[bot]@users.noreply.github.com> Date: Wed, 23 Sep 2026 20:32:31 +0200 Subject: [PATCH 1/3] Bump mkdocs-git-revision-date-localized-plugin in /docs/mkdocs (#5537) Bumps [mkdocs-git-revision-date-localized-plugin](https://github.com/timvink/mkdocs-git-revision-date-localized-plugin) from 1.5.4 to 1.6.0. - [Release notes](https://github.com/timvink/mkdocs-git-revision-date-localized-plugin/releases) - [Commits](https://github.com/timvink/mkdocs-git-revision-date-localized-plugin/compare/v1.5.4...v1.6.0) --- updated-dependencies: - dependency-name: mkdocs-git-revision-date-localized-plugin dependency-version: 1.6.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> --- docs/mkdocs/requirements.txt | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/mkdocs/requirements.txt b/docs/mkdocs/requirements.txt index b1ea71343..6396ec500 100644 --- a/docs/mkdocs/requirements.txt +++ b/docs/mkdocs/requirements.txt @@ -1,7 +1,7 @@ wheel==0.48.0 mkdocs==1.6.1 # documentation framework -mkdocs-git-revision-date-localized-plugin==1.5.4 # plugin "git-revision-date-localized" +mkdocs-git-revision-date-localized-plugin==1.6.0 # plugin "git-revision-date-localized" mkdocs-material==9.7.7 # theme for mkdocs mkdocs-material-extensions==1.3.1 # extensions mkdocs-minify-plugin==0.8.0 # plugin "minify" From f751547a81928c10b5041a6285ab503ad9481b7a Mon Sep 17 00:00:00 2001 From: Niels Lohmann Date: Wed, 23 Sep 2026 20:43:40 +0200 Subject: [PATCH 2/3] Add missing contributors to the README thanks list (#5543) Add 60 contributors whose work was not yet credited and update seven links that pointed to renamed or reassigned GitHub accounts. Signed-off-by: Niels Lohmann --- README.md | 74 +++++++++++++++++++++++++++++++++++++++++++++++++------ 1 file changed, 67 insertions(+), 7 deletions(-) diff --git a/README.md b/README.md index 4b4236bb6..becce71a1 100644 --- a/README.md +++ b/README.md @@ -1421,7 +1421,7 @@ I deeply appreciate the help of the following people. 6. [Joshua C. Randall](https://github.com/jrandall) fixed a bug in the floating-point serialization. 7. [Aaron Burghardt](https://github.com/aburgh) implemented code to parse streams incrementally. Furthermore, he greatly improved the parser class by allowing the definition of a filter function to discard undesired elements while parsing. 8. [Daniel Kopeček](https://github.com/dkopecek) fixed a bug in the compilation with GCC 5.0. -9. [Florian Weber](https://github.com/Florianjw) fixed a bug in and improved the performance of the comparison operators. +9. [Fiona Johanna Weber](https://github.com/Fiona-J-W) fixed a bug in and improved the performance of the comparison operators. 10. [Eric Cornelius](https://github.com/EricMCornelius) pointed out a bug in the handling with NaN and infinity values. He also improved the performance of the string escaping. 11. [易思龙](https://github.com/likebeta) implemented a conversion from anonymous enums. 12. [kepkin](https://github.com/kepkin) patiently pushed forward the support for Microsoft Visual Studio. @@ -1523,14 +1523,14 @@ I deeply appreciate the help of the following people. 108. [Kevin Tonon](https://github.com/ktonon) overworked the C++11 compiler checks in CMake. 109. [Axel Huebl](https://github.com/ax3l) simplified a CMake check and added support for the [Spack package manager](https://spack.io). 110. [Carlos O'Ryan](https://github.com/coryan) fixed a typo. -111. [James Upjohn](https://github.com/jammehcow) fixed a version number in the compilers section. +111. [James Upjohn](https://github.com/jupjohn) fixed a version number in the compilers section. 112. [Chuck Atkins](https://github.com/chuckatkins) adjusted the CMake files to the CMake packaging guidelines and provided documentation for the CMake integration. 113. [Jan Schöppach](https://github.com/dns13) fixed a typo. 114. [martin-mfg](https://github.com/martin-mfg) fixed a typo. 115. [Matthias Möller](https://github.com/TinyTinni) removed the dependency from `std::stringstream`. 116. [agrianius](https://github.com/agrianius) added code to use alternative string implementations. 117. [Daniel599](https://github.com/Daniel599) allowed to use more algorithms with the `items()` function. -118. [Julius Rakow](https://github.com/jrakow) fixed the Meson include directory and fixed the links to [cppreference.com](https://cppreference.com). +118. [Julius Rakow](https://github.com/juliusrakow) fixed the Meson include directory and fixed the links to [cppreference.com](https://cppreference.com). 119. [Sonu Lohani](https://github.com/sonulohani) fixed the compilation with MSVC 2015 in debug mode. 120. [grembo](https://github.com/grembo) fixed the test suite and re-enabled several test cases. 121. [Hyeon Kim](https://github.com/simnalamburt) introduced the macro `JSON_INTERNAL_CATCH` to control the exception handling inside the library. @@ -1581,7 +1581,7 @@ I deeply appreciate the help of the following people. 166. [Mark Beckwith](https://github.com/wythe) fixed a typo. 167. [yann-morin-1998](https://github.com/yann-morin-1998) helped to reduce the CMake requirement to version 3.1. 168. [Konstantin Podsvirov](https://github.com/podsvirov) maintains a package for the MSYS2 software distro. -169. [remyabel](https://github.com/remyabel) added GNUInstallDirs to the CMake files. +169. [remyabel](https://github.com/remyabel2) added GNUInstallDirs to the CMake files. 170. [Taylor Howard](https://github.com/taylorhoward92) fixed a unit test. 171. [Gabe Ron](https://github.com/Macr0Nerd) implemented the `to_string` method. 172. [Watal M. Iwasaki](https://github.com/heavywatal) fixed a Clang warning. @@ -1608,7 +1608,7 @@ I deeply appreciate the help of the following people. 193. [Hubert Chathi](https://github.com/uhoreg) made CMake's version config file architecture-independent. 194. [OmnipotentEntity](https://github.com/OmnipotentEntity) implemented the binary values for CBOR, MessagePack, BSON, and UBJSON. 195. [ArtemSarmini](https://github.com/ArtemSarmini) fixed a compilation issue with GCC 10 and fixed a leak. -196. [Evgenii Sopov](https://github.com/sea-kg) integrated the library to the wsjcpp package manager. +196. [Evgenii Sopov](https://github.com/sea5kg) integrated the library to the wsjcpp package manager. 197. [Sergey Linev](https://github.com/linev) fixed a compiler warning. 198. [Miguel Magalhães](https://github.com/magamig) fixed the year in the copyright. 199. [Gareth Sylvester-Bradley](https://github.com/garethsb-sony) fixed a compilation issue with MSVC. @@ -1702,7 +1702,7 @@ I deeply appreciate the help of the following people. 287. [NN](https://github.com/NN---) added the Visual Studio output directory to `.gitignore`. 288. [Romain Reignier](https://github.com/romainreignier) improved the performance of the vector output adapter. 289. [Mike](https://github.com/Mike-Leo-Smith) fixed the `std::iterator_traits`. -290. [Richard Hozák](https://github.com/zxey) added macro `JSON_NO_ENUM` to disable default enum conversions. +290. [Richard Hozák](https://github.com/richardhozak) added macro `JSON_NO_ENUM` to disable default enum conversions. 291. [vakokako](https://github.com/vakokako) fixed tests when compiling with C++20. 292. [Alexander “weej” Jones](https://github.com/alexweej) fixed an example in the README. 293. [Eli Schwartz](https://github.com/eli-schwartz) added more files to the `include.zip` archive. @@ -1727,7 +1727,7 @@ I deeply appreciate the help of the following people. 312. [Gareth Sylvester-Bradley](https://github.com/garethsb) added `operator/=` and `operator/` to construct JSON pointers. 313. [Michael Macnair](https://github.com/mykter) added support for afl-fuzz testing. 314. [Berkus Decker](https://github.com/berkus) fixed a typo in the README. -315. [Illia Polishchuk](https://github.com/effolkronium) improved the CMake testing. +315. [Illia Polishchuk](https://github.com/ilqvya) improved the CMake testing. 316. [Ikko Ashimine](https://github.com/eltociear) fixed a typo. 317. [Raphael Grimm](https://github.com/barcode) added the possibility to define a custom base class. 318. [tocic](https://github.com/tocic) fixed typos in the documentation. @@ -1797,6 +1797,66 @@ I deeply appreciate the help of the following people. 382. [bitFiedler](https://github.com/bitFiedler) made GDB pretty printer work with Python 3.8. 383. [Gianfranco Costamagna](https://github.com/LocutusOfBorg) fixed a compiler warning. 384. [risa2000](https://github.com/risa2000) made `std::filesystem::path` conversion to/from UTF-8 encoded string explicit. +385. [AM](https://github.com/maqnouch) fixed typos in the README. +386. [dmenendez-gruposantander](https://github.com/dmenendez-gruposantander) fixed typos in the comments of the examples. +387. [Mihai Stan](https://github.com/mstan-xx) fixed comparisons against the literal `0`. +388. [Matt Gumbel](https://github.com/intelmatt) fixed some `-Weffc++` warnings. +389. [vimpunk](https://github.com/vimpunk) moved a lambda out of an unevaluated context to support older compilers. +390. [Chris Harris](https://github.com/cjh1) fixed the compilation with GCC 4.8. +391. [Palmer Dabbelt](https://github.com/palmer-dabbelt) generated and installed a pkg-config file. +392. [Gus Pozuelo](https://github.com/ap-viavi) made `ordered_map` compatible with GCC 5.5, Clang 3.6, and Xcode 9. +393. [AK](https://github.com/Lioncky) fixed an MSVC build error caused by the `min`/`max` macros from `windows.h`. +394. [Sergiu Deitsch](https://github.com/sergiud) provided a fallback for missing `char8_t` support. +395. [Xiaochuan Ye](https://github.com/XueSongTap) fixed `from_msgpack` for `std::byte` input by specializing `std::char_traits`. +396. [Ville Vesilehto](https://github.com/thevilledev) fixed an overflow in the BJData size calculation and rejected overflowing negative integers in CBOR. +397. [NmPassTHFan](https://github.com/nmpassthf) replaced the deprecated `std::is_trivial` for C++26. +398. [Chris Ever](https://github.com/chirsz-ever) added the `ignore_trailing_commas` parser option. +399. [Kuan-Fu Wu](https://github.com/kfwu1999) fixed the example code for `json_pointer` initialization. +400. [David Kilzer](https://github.com/ddkilzer) added a missing header to the input adapters. +401. [Miko](https://github.com/mikomikotaishi) added proper C++20 module support, simplified the module API, and fixed missing exports. +402. [hitgirl](https://github.com/hitgil) fixed the CMake configuration when cross-compiling. +403. [Devon Thomas](https://github.com/ThomaDevOSU) mentioned the Artistic Style formatting in the contribution guidelines. +404. [Erik Hu](https://github.com/Erikhu1) made Coveralls upload errors non-fatal in the CI. +405. [co63oc](https://github.com/co63oc) fixed typos. +406. [DmitriBogdanov](https://github.com/DmitriBogdanov) fixed broken package manager links in the documentation. +407. [Bander](https://github.com/banderzhm) improved the MSVC compatibility of the C++ modules. +408. [Andy Choi](https://github.com/ccpong) removed an unnecessary `template` keyword before `get` in the README and the documentation. +409. [SamareshSingh](https://github.com/ssam18) fixed single-element brace initialization to copy/move instead of wrapping in an array, fixed the `WITH_DEFAULT` macros for `ordered_map`, and handled moved events in `serve_header.py`. +410. [Aditya](https://github.com/Lumowhisp) improved the documentation of the documentation generation. +411. [cheese1](https://github.com/cheese1) clarified the README. +412. [KhloodElhossiny](https://github.com/khloodelhossiny) enabled `std::string_view` keys in `operator[]`. +413. [Charles Cabergs](https://github.com/cacharle) fixed a `-Wtautological-constant-out-of-range-compare` warning. +414. [EALePain](https://github.com/EALePain) made the `std::tuple` conversion work with reference types such as `std::tie`. +415. [koala_oishi](https://github.com/chibi-dogs) fixed grammatical wording in the README. +416. [riccardoori11](https://github.com/riccardoori11) fixed a typo in the documentation. +417. [Swastik Bose](https://github.com/VasuBhakt) fixed the parent pointers after `update()` with `JSON_DIAGNOSTICS` and fixed the Doxygen autolinking of requirements. +418. [trdesilva](https://github.com/trdesilva) added `front`, `pop_front`, and `push_front` to `json_pointer`. +419. [Akhilesh Arora](https://github.com/akhilesharora) fixed an incomplete-type error with `ordered_json`. +420. [Hariom Phulre](https://github.com/hariomphulre) fixed the C++20 modules compilation with GCC. +421. [Kirill Lokotkov](https://github.com/RUSLoker) fixed printing `long double` values. +422. [George Sedov](https://github.com/radistmorse) added the `NLOHMANN_DEFINE_TYPE_*_WITH_NAMES` macros. +423. [Caillin Nugent](https://github.com/nugentcaillin) added the `NLOHMANN_JSON_SERIALIZE_ENUM_STRICT` macro. +424. [Cosmin D.](https://github.com/drcosmin) fixed `std::filesystem::path` conversions and added an MSVC workaround for `std::unique_ptr`. +425. [Paul Dreik](https://github.com/pauldreik) fixed a test relying on implementation-specific behavior. +426. [Daniel Falk](https://github.com/daniel-falk) added missing copyright notices to the SBOM. +427. [Federico Sfriso](https://github.com/federicosfriso05-dotcom) added support for constructing JSON values from C++20 range views. +428. [Luke Banicevic](https://github.com/banaboi) fixed corrupt BSON output for lengths exceeding `INT32_MAX`, cleaned up the BSON writer, and improved the documentation. +429. [Patrick Armstrong](https://github.com/Patrick10199) updated the CBOR references and the half-precision float assertions. +430. [Yash Bavadiya](https://github.com/xevrion) added checks to all BSON reads. +431. [hum4nBeing](https://github.com/hum4nBeing) fixed the overflow handling of high-precision numbers in UBJSON. +432. [tomatotomata](https://github.com/tomatotomata) added checks for reading CBOR tagged subtypes. +433. [YingqiDuan](https://github.com/YingqiDuan) documented the BSON interoperability. +434. [KBS](https://github.com/youdie006) documented the standards compliance and the strictness of `parse()` and `operator>>`. +435. [Petr Bělohlávek](https://github.com/petrbel) added Clang 21 and 22 to the CI. +436. [Dmitry Rantovov](https://github.com/darkdi) fixed the placement of a CBOR documentation block. +437. [ljcjclljc](https://github.com/ljcjclljc) fixed the comparison of large unsigned integers with signed integers. +438. [Sahil Kamate](https://github.com/sahilkamate03) fixed the handling of CBOR tags 0-5 and 21-23. +439. [Krishnanand G](https://github.com/Krishnanand-G) made the UBJSON writer reject `use_type` without `use_size`. +440. [whn](https://github.com/Whning0513) documented the lenient BSON input handling and corrected the complexity of `to_bson`. +441. [elix3r](https://github.com/22elix3r) fixed `update()` with `merge_objects` when merging a primitive into an object. +442. [Avionic Harshit](https://github.com/avionicharshit-byte) made `diff()` linear when an array shrinks. +443. [Qatadaha Bin Matloob](https://github.com/qatcod) fixed comparisons between integers and floats and fixed unparsable BJData output. +444. [Wu Shuwen](https://github.com/dajiaohuang) removed an unused include. Thanks a lot for helping out! Please [let me know](mailto:mail@nlohmann.me) if I forgot someone. From ed513715a89f288ae571c4a3b84952bf6839b775 Mon Sep 17 00:00:00 2001 From: Niels Lohmann Date: Thu, 24 Sep 2026 06:55:21 +0200 Subject: [PATCH 3/3] Document that a NUL byte in the input is treated as end of input (#5534) * docs: document that a NUL byte in the input is treated as end of input A NUL byte anywhere in the input - trailing, or embedded ahead of more otherwise well-formed JSON - is currently treated the same as genuine end of input, so parsing silently stops there instead of raising the parse_error.101 any other unexpected byte triggers. This mirrors the NUL-terminated-C-string convention already used when no explicit input length is given (json::parse(const char*) already stops at strlen()), just applied uniformly rather than only when a length is genuinely unavailable. This behavior predates this change and is not being altered here - changing it would be an observable, backwards-incompatible behavior change for any caller that (knowingly or not) depends on it, which is not something to do silently in a patch. Documenting the current, verified behavior as a new FAQ entry instead, so it's an intentional and discoverable part of the contract rather than a surprise. Fixes #5530. Signed-off-by: Niels Lohmann Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY * Add JSON_STRICT_NUL_HANDLING opt-in macro for issue #5530 A NUL byte anywhere in the input is currently treated the same as real end of input, rather than raising parse_error.101 like any other unexpected byte (documented in the previous commit's FAQ entry). A full unconditional fix was tried in PR #5532 but rejected as too risky to ship by default: any caller could depend on the current behavior, even unknowingly (e.g. a zero-padded buffer). On PR #5534, gregmarr proposed a compile-time opt-in flag instead, and the maintainer agreed, wanting it available now and defaulting to the corrected behavior in 4.0.0. This mirrors the existing JSON_BRACE_INIT_COPY_SEMANTICS precedent as closely as sensible: - JSON_STRICT_NUL_HANDLING defaults to 0 (off); the three lexer sites that treat '\0' as EOF/comment-terminator are gated with `#if !JSON_STRICT_NUL_HANDLING` so the default-off behavior is byte-for-byte identical to today's. - input_adapters.hpp's `T (&array)[N]` overload additionally trims a single trailing '\0' from a `char` array (e.g. a string literal like `json::parse("123")`) when the macro is on, so that case keeps working; every other element type (unsigned char, std::uint8_t, ...) always keeps its full extent. This intentionally does *not* reuse the existing strlen()-based pointer overload via SFINAE-excluding `char` from the array overload, as originally sketched for this change: that approach is ambiguous against the newer generic container overload added since PR #5532, and even where it compiles, strlen()-scanning a `char` array that is not NUL-terminated within its bounds reads past the end of the array (confirmed with AddressSanitizer). Trimming only a single trailing byte, without scanning, avoids both problems. - Documented via docs/mkdocs/docs/api/macros/json_strict_nul_handling.md, linked from the macros index/nav/features page, the FAQ entry, and the parse/accept/operator>> reference pages. - Tested in unit-class_parser.cpp and unit-deserialization.cpp, default state unguarded and opt-in state guarded. Since the library itself #undefs the macro at the end of json.hpp (as JSON_BRACE_INIT_COPY_SEMANTICS already does), a plain `#if defined(JSON_STRICT_NUL_HANDLING)` guard after the include never actually triggers; the tests instead capture the command-line value into a test-local macro before including the header. A few pre-existing fixtures elsewhere (std::array sized one larger than their literal, relying on value-initialization to silently add a trailing zero byte) needed the same one-byte adjustment to keep passing under the opt-in behavior. Unlike the precedent, this adds a proper `JSON_StrictNulHandling` CMake option (rather than a raw -DCMAKE_CXX_FLAGS injection) and wires its ci_test_strict_nul_handling target into the ci_cmake_options job matrix in .github/workflows/ubuntu.yml, so the opt-in build is actually exercised in CI -- closing the one gap in the precedent's own CI setup (ci_test_brace_init_copy_semantics is defined but never referenced by any workflow, so it has never actually run). Co-Authored-By: Claude Sonnet 5 Signed-off-by: Niels Lohmann * Clarify where JSON_STRICT_NUL_HANDLING does not reject NUL bytes Signed-off-by: Niels Lohmann --------- Signed-off-by: Niels Lohmann Signed-off-by: Niels Lohmann Co-authored-by: Claude Sonnet 5 --- .github/workflows/ubuntu.yml | 2 +- CMakeLists.txt | 6 + cmake/ci.cmake | 17 +++ docs/mkdocs/docs/api/basic_json/accept.md | 8 ++ docs/mkdocs/docs/api/basic_json/parse.md | 8 ++ docs/mkdocs/docs/api/macros/index.md | 5 + .../api/macros/json_strict_nul_handling.md | 126 ++++++++++++++++++ docs/mkdocs/docs/api/operator_gtgt.md | 11 ++ docs/mkdocs/docs/features/macros.md | 13 ++ docs/mkdocs/docs/home/faq.md | 48 +++++++ docs/mkdocs/docs/integration/cmake.md | 5 + docs/mkdocs/mkdocs.yml | 1 + .../nlohmann/detail/input/input_adapters.hpp | 15 +++ include/nlohmann/detail/input/lexer.hpp | 13 +- include/nlohmann/detail/macro_scope.hpp | 4 + include/nlohmann/detail/macro_unscope.hpp | 1 + single_include/nlohmann/json.hpp | 33 ++++- tests/src/unit-class_parser.cpp | 108 ++++++++++++++- tests/src/unit-deserialization.cpp | 56 +++++++- tests/src/unit-regression2.cpp | 6 +- 20 files changed, 475 insertions(+), 11 deletions(-) create mode 100644 docs/mkdocs/docs/api/macros/json_strict_nul_handling.md diff --git a/.github/workflows/ubuntu.yml b/.github/workflows/ubuntu.yml index 7ac4cfcb0..c7804079c 100644 --- a/.github/workflows/ubuntu.yml +++ b/.github/workflows/ubuntu.yml @@ -100,7 +100,7 @@ jobs: container: ubuntu:focal strategy: matrix: - target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_disableenumserialization, ci_test_skiplibraryversioncheck, ci_test_simdutf] + target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_disableenumserialization, ci_test_skiplibraryversioncheck, ci_test_simdutf, ci_test_strict_nul_handling] steps: - name: Install build-essential run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev diff --git a/CMakeLists.txt b/CMakeLists.txt index 9669946a5..4c43c23ff 100644 --- a/CMakeLists.txt +++ b/CMakeLists.txt @@ -59,6 +59,7 @@ option(JSON_LegacyDiscardedValueComparison "Enable legacy discarded value compar option(JSON_Install "Install CMake targets during install step." ${MAIN_PROJECT}) option(JSON_MultipleHeaders "Use non-amalgamated version of the library." ON) option(JSON_SystemInclude "Include as system headers (skip for clang-tidy)." OFF) +option(JSON_StrictNulHandling "Build with strict NUL-byte handling enabled." OFF) if (JSON_CI) include(ci) @@ -108,6 +109,10 @@ if (JSON_Diagnostics) message(STATUS "Diagnostics enabled (JSON_DIAGNOSTICS=1)") endif() +if (JSON_StrictNulHandling) + message(STATUS "Strict NUL-byte handling enabled (JSON_STRICT_NUL_HANDLING=1)") +endif() + if (JSON_Diagnostic_Positions) message(STATUS "Diagnostic positions enabled (JSON_DIAGNOSTIC_POSITIONS=1)") endif() @@ -141,6 +146,7 @@ target_compile_definitions( $<$:JSON_DIAGNOSTICS=1> $<$:JSON_DIAGNOSTIC_POSITIONS=1> $<$:JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1> + $<$:JSON_STRICT_NUL_HANDLING=1> ) target_include_directories( diff --git a/cmake/ci.cmake b/cmake/ci.cmake index 7d085fb2b..2d4bbf6d2 100644 --- a/cmake/ci.cmake +++ b/cmake/ci.cmake @@ -245,6 +245,23 @@ add_custom_target(ci_test_brace_init_copy_semantics COMMENT "Compile and test with brace-init copy semantics enabled" ) +############################################################################### +# Enable strict NUL-byte handling. +############################################################################### + +add_custom_target(ci_test_strict_nul_handling + COMMAND ${CMAKE_COMMAND} + -DCMAKE_BUILD_TYPE=Debug -GNinja + -DJSON_BuildTests=ON -DJSON_FastTests=ON -DJSON_StrictNulHandling=ON + -S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_strict_nul_handling + COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_strict_nul_handling + # unit-testsuites contains a fixture (a "1e308" test value) that relies on the + # legacy NUL-as-end-of-input behavior this macro disables; exclude it here, as + # it is expected to fail under strict NUL handling and is out of scope for it + COMMAND cd ${PROJECT_BINARY_DIR}/build_strict_nul_handling && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure -E "test-testsuites" + COMMENT "Compile and test with strict NUL-byte handling enabled" +) + ############################################################################### # Disable global UDLs. ############################################################################### diff --git a/docs/mkdocs/docs/api/basic_json/accept.md b/docs/mkdocs/docs/api/basic_json/accept.md index 5ba009c3e..0cdcae3a8 100644 --- a/docs/mkdocs/docs/api/basic_json/accept.md +++ b/docs/mkdocs/docs/api/basic_json/accept.md @@ -90,6 +90,10 @@ Linear in the length of the input. The parser is a predictive LL(1) parser. A UTF-8 byte order mark is silently ignored. +By default, a `'\0'` (NUL) byte anywhere in the input is treated as end of input, rather than as an ordinary (and, +outside of a string, invalid) byte; see the [FAQ entry](../../home/faq.md#nul-bytes-in-the-input) for details and the +[`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) macro to opt into rejecting it instead. + ## Examples ??? example @@ -111,6 +115,8 @@ A UTF-8 byte order mark is silently ignored. - [parse](parse.md) - deserialize from a compatible input - [sax_parse](sax_parse.md) - parse input using the SAX interface - [operator>>](../operator_gtgt.md) - deserialize from stream +- [`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input + instead of treating it as end of input ## Version history @@ -120,6 +126,8 @@ A UTF-8 byte order mark is silently ignored. - Added `ignore_trailing_commas` in version 3.13.0. - Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0. - Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0. +- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating + it as end of input; planned to become the default in version 4.0.0. !!! warning "Deprecation" diff --git a/docs/mkdocs/docs/api/basic_json/parse.md b/docs/mkdocs/docs/api/basic_json/parse.md index ef22e9873..20bb1c708 100644 --- a/docs/mkdocs/docs/api/basic_json/parse.md +++ b/docs/mkdocs/docs/api/basic_json/parse.md @@ -103,6 +103,10 @@ A UTF-8 byte order mark is silently ignored. Invalid Unicode escapes and unpaired surrogates in the input are reported as [`parse_error.101`](../../home/exceptions.md#jsonexceptionparse_error101) with a detailed message. +By default, a `'\0'` (NUL) byte anywhere in the input is treated as end of input, rather than as an ordinary (and, +outside of a string, invalid) byte; see the [FAQ entry](../../home/faq.md#nul-bytes-in-the-input) for details and the +[`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) macro to opt into rejecting it instead. + ## Examples ??? example "Parsing from a character array" @@ -236,6 +240,8 @@ Invalid Unicode escapes and unpaired surrogates in the input are reported as - [accept](accept.md) - check if the input is valid JSON - [sax_parse](sax_parse.md) - parse input using the SAX interface - [operator>>](../operator_gtgt.md) - deserialize from stream +- [`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input + instead of treating it as end of input ## Version history @@ -246,6 +252,8 @@ Invalid Unicode escapes and unpaired surrogates in the input are reported as - Added `ignore_trailing_commas` in version 3.13.0. - Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0. - Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0. +- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating + it as end of input; planned to become the default in version 4.0.0. !!! warning "Deprecation" diff --git a/docs/mkdocs/docs/api/macros/index.md b/docs/mkdocs/docs/api/macros/index.md index e818f032a..ca3af720c 100644 --- a/docs/mkdocs/docs/api/macros/index.md +++ b/docs/mkdocs/docs/api/macros/index.md @@ -14,6 +14,11 @@ header. See also the [macro overview page](../../features/macros.md). - [**JSON_DIAGNOSTIC_POSITIONS**](json_diagnostic_positions.md) - access positions of elements - [**JSON_NOEXCEPTION**](json_noexception.md) - switch off exceptions +## Parsing + +- [**JSON_STRICT_NUL_HANDLING**](json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input instead of + treating it as end of input + ## Language support - [**JSON_HAS_CPP_11**
**JSON_HAS_CPP_14**
**JSON_HAS_CPP_17**
**JSON_HAS_CPP_20**](json_has_cpp_11.md) - set supported C++ standard diff --git a/docs/mkdocs/docs/api/macros/json_strict_nul_handling.md b/docs/mkdocs/docs/api/macros/json_strict_nul_handling.md new file mode 100644 index 000000000..1832f2353 --- /dev/null +++ b/docs/mkdocs/docs/api/macros/json_strict_nul_handling.md @@ -0,0 +1,126 @@ +# JSON_STRICT_NUL_HANDLING + +```cpp +#define JSON_STRICT_NUL_HANDLING /* value */ +``` + +When defined to `1`, a `'\0'` (NUL) byte in JSON text input is rejected with `parse_error.101`, like any other +unexpected byte, instead of being silently treated as end of input. + +The macro only affects the JSON text parser ([`parse`](../basic_json/parse.md), [`accept`](../basic_json/accept.md), +[`sax_parse`](../basic_json/sax_parse.md), and [`operator>>`](../operator_gtgt.md)). There are three cases where a NUL +byte is still not rejected: + +- The binary formats ([`from_bjdata`](../basic_json/from_bjdata.md), [`from_bson`](../basic_json/from_bson.md), + [`from_cbor`](../basic_json/from_cbor.md), [`from_msgpack`](../basic_json/from_msgpack.md), + [`from_ubjson`](../basic_json/from_ubjson.md)) are never affected: there, `0x00` is ordinary data. +- A bare `const char*` pointer has no length of its own, so its length is still determined with `strlen()`. The first + NUL byte therefore still marks the end of the input, and nothing after it is read. +- One trailing `'\0'` at the end of a `char` array (e.g., a string literal) is trimmed; see the warning below. + +## Default definition + +The default value is `0` (disabled — existing behavior is preserved). + +```cpp +#define JSON_STRICT_NUL_HANDLING 0 +``` + +## Notes + +!!! note "Background" + + By default, a `'\0'` byte anywhere in the input is treated the same as the real end of the input, rather than as + an ordinary (and, outside of a string, invalid) byte. Everything from that byte onward is silently ignored, + without a parse error - including further, otherwise well-formed JSON: + + ```cpp + json::parse(std::string("123") + '\0'); // == 123, no error + json::parse(std::string("123") + '\0' + "true"); // == 123, the "true" is silently ignored too + ``` + + This falls out of the same convention used when no explicit input length is given at all: parsing from a + `const char*` already stops at the first NUL byte via `strlen()`, since a bare pointer has no length of its own. + The library applies that same NUL-terminated-C-string convention uniformly, rather than only when a length is + genuinely unavailable - so a `std::string`, iterator range, or container whose content happens to include a NUL + byte is affected the same way a raw `const char*` would be (see the + [FAQ entry](../../home/faq.md#nul-bytes-in-the-input) for a fuller explanation). + + This was not fixed unconditionally, because doing so is backwards-incompatible for any caller who happens to + depend on the current behavior - even unknowingly, for instance because their input already contains trailing + padding they never noticed was being discarded (see [#5530](https://github.com/nlohmann/json/issues/5530)). + This macro instead offers an opt-in path to the corrected behavior ahead of version 4.0.0, where it is planned to + become the default. + +!!! warning "Opt-in only" + + This macro must be defined **before** including ``. Defining it after the include has no + effect. + + Enabling it also changes how a `char` array (including a string literal, e.g. `json::parse("123")`) is read: such + an array normally carries a trailing `'\0'` contributed by the compiler, not by the source text. With this macro + enabled, that one trailing byte is trimmed if present so that parsing a string literal keeps working; every other + byte in the array - including any `'\0'` that is not the very last element - is read as real data and rejected + like any other unexpected byte. Arrays of any other element type (`unsigned char`, `std::uint8_t`, ...), as used + for CBOR or MessagePack, are never affected by this trimming; their full extent - including a genuine trailing + `0x00` - is always preserved, in both states of this macro. + +!!! tip "Workaround without the macro" + + To reject a NUL byte without enabling this macro, trim your input yourself before calling `parse()`: + + ```cpp + s.resize(s.find('\0')); // drop everything from the first NUL onward, if any + json::parse(s); + ``` + +## Examples + +??? example "Default behavior (macro not defined)" + + Without the macro, a NUL byte silently ends parsing at that point: + + ```cpp + #include + + using json = nlohmann::json; + + int main() + { + json j = json::parse(std::string("123") + '\0' + "true"); + // j is 123 -- the '\0' and everything after it is silently ignored + } + ``` + +??? example "Opt-in strict handling (macro defined to 1)" + + With the macro, a NUL byte is rejected like any other unexpected byte: + + ```cpp + #define JSON_STRICT_NUL_HANDLING 1 + #include + + using json = nlohmann::json; + + int main() + { + json j = json::parse(std::string("123") + '\0' + "true"); + // throws parse_error.101 -- the NUL byte is now invalid input, + // exactly like any other unexpected trailing byte + + json ok = json::parse("123"); + // ok is 123 -- parsing from a string literal still works + } + ``` + +## See also + +- [FAQ: NUL bytes in the input](../../home/faq.md#nul-bytes-in-the-input) +- [**parse**](../basic_json/parse.md) - deserialize from a compatible input +- [**accept**](../basic_json/accept.md) - check if the input is valid JSON +- [**operator>>**](../operator_gtgt.md) - deserialize from stream + +## Version history + +- Added in version 3.13.0. +- Planned to become the default (with the macro removed) in version 4.0.0. diff --git a/docs/mkdocs/docs/api/operator_gtgt.md b/docs/mkdocs/docs/api/operator_gtgt.md index a9ba17564..3e60d5236 100644 --- a/docs/mkdocs/docs/api/operator_gtgt.md +++ b/docs/mkdocs/docs/api/operator_gtgt.md @@ -72,6 +72,13 @@ input >> j2; // parses the next value Note that reading concatenated values does **not** work for [JSON Lines](../features/parsing/json_lines.md) (newline-delimited JSON) input -- see that page for why and for the recommended alternative. +By default, a `'\0'` (NUL) byte encountered while reading a value is treated as end of input, rather than as an +ordinary (and, outside of a string, invalid) byte; see the [FAQ entry](../home/faq.md#nul-bytes-in-the-input) for +details and the [`JSON_STRICT_NUL_HANDLING`](macros/json_strict_nul_handling.md) macro to opt into rejecting it +instead. Because `operator>>` only parses a single value and does not require the rest of the stream to be consumed, +a NUL byte *after* a complete value has no effect on `operator>>` either way; it only matters while a value is still +being read. + !!! warning "Deprecation" This function replaces function `#!cpp std::istream& operator<<(basic_json& j, std::istream& i)` which has @@ -98,7 +105,11 @@ Note that reading concatenated values does **not** work for [JSON Lines](../feat - [accept](basic_json/accept.md) - check if the input is valid JSON - [parse](basic_json/parse.md) - deserialize from a compatible input +- [`JSON_STRICT_NUL_HANDLING`](macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input + instead of treating it as end of input ## Version history - Added in version 1.0.0. +- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating + it as end of input; planned to become the default in version 4.0.0. diff --git a/docs/mkdocs/docs/features/macros.md b/docs/mkdocs/docs/features/macros.md index c4602fa5a..927e00df3 100644 --- a/docs/mkdocs/docs/features/macros.md +++ b/docs/mkdocs/docs/features/macros.md @@ -105,6 +105,19 @@ using the library with compilers that do not fully support C++11 and may only wo See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md). +## `JSON_STRICT_NUL_HANDLING` + +When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other +unexpected byte, instead of being silently treated as end of input (see the +[FAQ entry](../home/faq.md#nul-bytes-in-the-input) for background). The default value is `0`, which preserves the +existing behavior; this is planned to become the default in version 4.0.0. + +The strict handling can also be enabled with the CMake option +[`JSON_StrictNulHandling`](../integration/cmake.md#json_strictnulhandling) (`OFF` by default) which sets +`JSON_STRICT_NUL_HANDLING` accordingly. + +See [full documentation of `JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md). + ## `JSON_THROW_USER(exception)` This macro overrides `#!cpp throw` calls inside the library. The argument is the exception to be thrown. diff --git a/docs/mkdocs/docs/home/faq.md b/docs/mkdocs/docs/home/faq.md index 8394dcfc7..8b3602bd1 100644 --- a/docs/mkdocs/docs/home/faq.md +++ b/docs/mkdocs/docs/home/faq.md @@ -90,6 +90,54 @@ The library supports **Unicode input** as follows: In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding. +### NUL bytes in the input + +!!! question "Questions" + + - Why does `json::parse()` silently ignore part of my input? + - Why does a `std::string`/buffer with extra data after the JSON text parse without error, while a similar-looking string with extra text does not? + +A `'\0'` (NUL) byte anywhere in the input is treated the same as the real end of the input, rather than as an ordinary (and, outside of a string, invalid) byte. Everything from that byte onward is silently ignored, without a parse error — including further, otherwise well-formed JSON: + +```cpp +json::parse(std::string("123") + '\0'); // == 123, no error +json::parse(std::string("123") + '\0' + "true"); // == 123, the "true" is silently ignored too +``` + +This is different from any other unexpected trailing byte, which *does* raise [`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101): + +```cpp +json::parse("123x"); // throws parse_error.101: unexpected additional data +``` + +This falls out of the same convention used when no explicit input length is given at all: `json::parse(const char*)` already stops at the first NUL byte via `strlen()`, since a bare pointer has no length of its own. The library applies that same NUL-terminated-C-string convention uniformly, rather than only when a length is genuinely unavailable — so a `std::string`, iterator range, or container whose content happens to include a NUL byte is affected the same way a raw `const char*` would be. + +If your input may contain a trailing or embedded NUL that is **not** meant to signal the end of the JSON text — for instance, a fixed-size, zero-padded buffer — trim it yourself before calling `parse()`, since the library will otherwise silently stop there instead of raising an error: + +```cpp +s.resize(s.find('\0')); // drop everything from the first NUL onward, if any +json::parse(s); +``` + +**Opt-in strict handling (since version 3.13.0)** + +Manually trimming every input is easy to forget. If you define [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) to `1` before including the library, a `'\0'` byte is instead rejected like any other unexpected byte and raises `parse_error.101`, instead of being treated as end of input: + +```cpp +#define JSON_STRICT_NUL_HANDLING 1 +#include + +json::parse(std::string("123") + '\0'); // throws parse_error.101 instead of silently returning 123 +``` + +This macro defaults to `0` (disabled, preserving the behavior described above) to avoid breaking existing code that may depend on it, even unknowingly; it is planned to become the default in version 4.0.0. See [its documentation](../api/macros/json_strict_nul_handling.md) for details, including how it also affects `char` arrays such as string literals. + +Note that this is unrelated to an *unescaped* NUL byte occurring **inside** a quoted JSON string, which is a different, already-invalid case and is correctly rejected either way: + +```cpp +json::parse(std::string("\"") + '\0' + "\""); // throws parse_error.101: control character U+0000 (NUL) must be escaped to \u0000 +``` + ### Wide string handling !!! question diff --git a/docs/mkdocs/docs/integration/cmake.md b/docs/mkdocs/docs/integration/cmake.md index a8a6d52b6..71512cbd5 100644 --- a/docs/mkdocs/docs/integration/cmake.md +++ b/docs/mkdocs/docs/integration/cmake.md @@ -198,6 +198,11 @@ Use the non-amalgamated version of the library. This option is `ON` by default. Treat the library headers like system headers (i.e., adding `SYSTEM` to the [`target_include_directories`](https://cmake.org/cmake/help/latest/command/target_include_directories.html) call) to check for this library by tools like Clang-Tidy. This option is `OFF` by default. +### `JSON_StrictNulHandling` + +Reject a `'\0'` (NUL) byte in the input instead of treating it as end of input, by defining the macro +[`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md). This option is `OFF` by default. + ### `JSON_Valgrind` Execute the test suite with [Valgrind](https://valgrind.org). This option is `OFF` by default. Depends on `JSON_BuildTests`. diff --git a/docs/mkdocs/mkdocs.yml b/docs/mkdocs/mkdocs.yml index d0f9cfdfd..856d86e6d 100644 --- a/docs/mkdocs/mkdocs.yml +++ b/docs/mkdocs/mkdocs.yml @@ -294,6 +294,7 @@ nav: - 'JSON_NO_IO': api/macros/json_no_io.md - 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md - 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md + - 'JSON_STRICT_NUL_HANDLING': api/macros/json_strict_nul_handling.md - 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md - 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md - 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md diff --git a/include/nlohmann/detail/input/input_adapters.hpp b/include/nlohmann/detail/input/input_adapters.hpp index bd19d32a8..775a8398c 100644 --- a/include/nlohmann/detail/input/input_adapters.hpp +++ b/include/nlohmann/detail/input/input_adapters.hpp @@ -762,6 +762,21 @@ contiguous_bytes_input_adapter input_adapter(CharT b) template auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) { +#if JSON_STRICT_NUL_HANDLING + // A `char` array from string-literal initialization (e.g. json::parse("123")) + // carries a trailing '\0' contributed by the compiler, not by the source + // text; drop exactly that one byte so it is not mistaken for real trailing + // data. Every other element type (unsigned char, std::uint8_t, ...) keeps + // the full extent unconditionally, since a trailing zero byte there is + // genuine data (e.g. CBOR/MessagePack). This intentionally does not + // strlen()-scan the array (as the pointer overload above does for a + // null-delimited string): for a `char` array that is not NUL-terminated + // within its bounds, that would read past the end of the array. + if (std::is_same::type, char>::value && N > 0 && array[N - 1] == 0) + { + return input_adapter(array, array + N - 1); + } +#endif return input_adapter(array, array + N); } diff --git a/include/nlohmann/detail/input/lexer.hpp b/include/nlohmann/detail/input/lexer.hpp index bc31337f9..fe85cd53d 100644 --- a/include/nlohmann/detail/input/lexer.hpp +++ b/include/nlohmann/detail/input/lexer.hpp @@ -952,7 +952,9 @@ class lexer : public lexer_base case '\n': case '\r': case char_traits::eof(): +#if !JSON_STRICT_NUL_HANDLING case '\0': +#endif return true; default: @@ -970,8 +972,10 @@ class lexer : public lexer_base { switch (get()) { - case char_traits::eof(): +#if !JSON_STRICT_NUL_HANDLING case '\0': +#endif + case char_traits::eof(): { error_message = "invalid comment; missing closing '*/'"; return false; @@ -2153,9 +2157,12 @@ scan_number_done: case '9': return scan_number_dispatch(std::integral_constant {}); - // end of input (the null byte is needed when parsing from - // string literals) +#if !JSON_STRICT_NUL_HANDLING case '\0': +#endif + // end of input; by default, a null byte is also treated as end of + // input for backwards compatibility (see JSON_STRICT_NUL_HANDLING + // to opt into rejecting a null byte in the input instead) case char_traits::eof(): return token_type::end_of_input; diff --git a/include/nlohmann/detail/macro_scope.hpp b/include/nlohmann/detail/macro_scope.hpp index 4682fd361..8aacc0c51 100644 --- a/include/nlohmann/detail/macro_scope.hpp +++ b/include/nlohmann/detail/macro_scope.hpp @@ -807,3 +807,7 @@ void templated_json_throw(ExceptionType exception) #ifndef JSON_BRACE_INIT_COPY_SEMANTICS #define JSON_BRACE_INIT_COPY_SEMANTICS 0 #endif + +#ifndef JSON_STRICT_NUL_HANDLING + #define JSON_STRICT_NUL_HANDLING 0 +#endif diff --git a/include/nlohmann/detail/macro_unscope.hpp b/include/nlohmann/detail/macro_unscope.hpp index 2ac25a46d..c692ea68e 100644 --- a/include/nlohmann/detail/macro_unscope.hpp +++ b/include/nlohmann/detail/macro_unscope.hpp @@ -27,6 +27,7 @@ #undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_USE_GLOBAL_UDLS #undef JSON_BRACE_INIT_COPY_SEMANTICS +#undef JSON_STRICT_NUL_HANDLING #ifndef JSON_TEST_KEEP_MACROS #undef JSON_CATCH diff --git a/single_include/nlohmann/json.hpp b/single_include/nlohmann/json.hpp index 4af4af1e0..4418a6c19 100644 --- a/single_include/nlohmann/json.hpp +++ b/single_include/nlohmann/json.hpp @@ -3186,6 +3186,10 @@ void templated_json_throw(ExceptionType exception) #define JSON_BRACE_INIT_COPY_SEMANTICS 0 #endif +#ifndef JSON_STRICT_NUL_HANDLING + #define JSON_STRICT_NUL_HANDLING 0 +#endif + #if JSON_HAS_THREE_WAY_COMPARISON #include // partial_ordering #endif @@ -7903,6 +7907,21 @@ contiguous_bytes_input_adapter input_adapter(CharT b) template auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) { +#if JSON_STRICT_NUL_HANDLING + // A `char` array from string-literal initialization (e.g. json::parse("123")) + // carries a trailing '\0' contributed by the compiler, not by the source + // text; drop exactly that one byte so it is not mistaken for real trailing + // data. Every other element type (unsigned char, std::uint8_t, ...) keeps + // the full extent unconditionally, since a trailing zero byte there is + // genuine data (e.g. CBOR/MessagePack). This intentionally does not + // strlen()-scan the array (as the pointer overload above does for a + // null-delimited string): for a `char` array that is not NUL-terminated + // within its bounds, that would read past the end of the array. + if (std::is_same::type, char>::value && N > 0 && array[N - 1] == 0) + { + return input_adapter(array, array + N - 1); + } +#endif return input_adapter(array, array + N); } @@ -9512,7 +9531,9 @@ class lexer : public lexer_base case '\n': case '\r': case char_traits::eof(): +#if !JSON_STRICT_NUL_HANDLING case '\0': +#endif return true; default: @@ -9530,8 +9551,10 @@ class lexer : public lexer_base { switch (get()) { - case char_traits::eof(): +#if !JSON_STRICT_NUL_HANDLING case '\0': +#endif + case char_traits::eof(): { error_message = "invalid comment; missing closing '*/'"; return false; @@ -10713,9 +10736,12 @@ scan_number_done: case '9': return scan_number_dispatch(std::integral_constant {}); - // end of input (the null byte is needed when parsing from - // string literals) +#if !JSON_STRICT_NUL_HANDLING case '\0': +#endif + // end of input; by default, a null byte is also treated as end of + // input for backwards compatibility (see JSON_STRICT_NUL_HANDLING + // to opt into rejecting a null byte in the input instead) case char_traits::eof(): return token_type::end_of_input; @@ -30027,6 +30053,7 @@ struct formatter // NOLINT(cert-dcl58-c #undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_USE_GLOBAL_UDLS #undef JSON_BRACE_INIT_COPY_SEMANTICS +#undef JSON_STRICT_NUL_HANDLING #ifndef JSON_TEST_KEEP_MACROS #undef JSON_CATCH diff --git a/tests/src/unit-class_parser.cpp b/tests/src/unit-class_parser.cpp index df4e7270d..e532992f4 100644 --- a/tests/src/unit-class_parser.cpp +++ b/tests/src/unit-class_parser.cpp @@ -8,6 +8,14 @@ #include "doctest_compatibility.h" +// capture whether JSON_STRICT_NUL_HANDLING was enabled on the command line +// (e.g. -DJSON_STRICT_NUL_HANDLING=1) *before* including json.hpp, since the +// library #undefs JSON_STRICT_NUL_HANDLING itself once the header has been +// fully processed (see include/nlohmann/detail/macro_unscope.hpp) +#if defined(JSON_STRICT_NUL_HANDLING) && (JSON_STRICT_NUL_HANDLING == 1) + #define JSON_TEST_STRICT_NUL_HANDLING_ENABLED 1 +#endif + #define JSON_TESTS_PRIVATE #include using nlohmann::json; @@ -545,6 +553,88 @@ TEST_CASE("parser class") } } + SECTION("NUL byte handling (issue #5530, JSON_STRICT_NUL_HANDLING)") + { + // by default, a NUL byte anywhere in the input (not inside a quoted + // string, which is covered above) is silently treated the same as + // real end of input; JSON_STRICT_NUL_HANDLING (off by default, see + // docs/mkdocs/docs/api/macros/json_strict_nul_handling.md) makes a + // NUL byte an error like any other unexpected byte instead. + // + // The two sections below are mutually exclusive: this whole test + // binary is compiled once, with JSON_STRICT_NUL_HANDLING either + // left at its default or forced to 1 (e.g. by the dedicated + // ci_test_strict_nul_handling CI target), so only the section + // matching the actual, compiled-in behavior can pass. +#if !defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED) + SECTION("default behavior (macro not enabled)") + { + // a NUL byte after a complete value silently truncates the input + std::string s = "123"; + s.push_back('\0'); + s += "4"; + CHECK(json::parse(s) == json(123)); + CHECK(json::accept(s)); + + // parsing from a string literal is unaffected either way + CHECK(json::parse("123") == json(123)); + } +#endif + +#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED) + SECTION("opt-in strict behavior (JSON_STRICT_NUL_HANDLING == 1)") + { + // a NUL byte after a complete value is now a parse error, + // instead of silently truncating the input + { + std::string s = "123"; + s.push_back('\0'); + json _; // NOLINT(readability-identifier-naming) + CHECK_THROWS_WITH_AS(_ = json::parse(s), + "[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: '123'; expected end of input", + json::parse_error&); + CHECK_FALSE(json::accept(s)); + } + + // a NUL byte where a value is expected is now a parse error, + // instead of being treated the same as an empty input + { + const std::string s(1, '\0'); + json _; // NOLINT(readability-identifier-naming) + CHECK_THROWS_WITH_AS(_ = json::parse(s), + "[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: ''", + json::parse_error&); + CHECK_FALSE(json::accept(s)); + } + + // a NUL byte inside a // comment no longer stops the comment + // scan early; scanning continues correctly past it + { + std::string s = "1 // a"; + s.push_back('\0'); + s += "b\n"; + CHECK(json::parse(s, nullptr, true, true) == json(1)); + CHECK(json::accept(s, true, true)); + } + + // a NUL byte inside a /* */ comment no longer stops the + // comment scan early either + { + std::string s = "1 /* a"; + s.push_back('\0'); + s += "b */ "; + CHECK(json::parse(s, nullptr, true, true) == json(1)); + CHECK(json::accept(s, true, true)); + } + + // regression guard: parsing from a string literal (which + // carries a compiler-appended trailing '\0') still works, + // even though a NUL byte is now rejected everywhere else + CHECK(json::parse("123") == json(123)); + } +#endif + } + SECTION("number") { SECTION("integers") @@ -1894,7 +1984,13 @@ TEST_CASE("parser class") SECTION("from std::array") { - std::array v { {'t', 'r', 'u', 'e'} }; + // NOTE: this array is sized to exactly the length of "true" (unlike + // the trailing-NUL-tolerant default behavior elsewhere in this file, + // see the "NUL byte handling" section above); a size of 5 here would + // leave a value-initialized trailing 0x00 element that is only + // silently accepted as end-of-input by default and would fail under + // JSON_STRICT_NUL_HANDLING + std::array v { {'t', 'r', 'u', 'e'} }; json j; json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j); CHECK(j == json(true)); @@ -2035,7 +2131,17 @@ TEST_CASE("parser class") { json _; CHECK_THROWS_WITH_AS(_ = json::parse("/a", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid comment; expecting '/' or '*' after '/'; last read: '/a'", json::parse_error); + // "/*" is a string literal, so it carries a compiler-appended trailing + // '\0'; by default that NUL is read like any other byte and shows up + // in "last read", but JSON_STRICT_NUL_HANDLING trims exactly that one + // trailing byte from a char array (see + // docs/mkdocs/docs/api/macros/json_strict_nul_handling.md), so it no + // longer appears in the message in that state +#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED) + CHECK_THROWS_WITH_AS(_ = json::parse("/*", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid comment; missing closing '*/'; last read: '/*'", json::parse_error); +#else CHECK_THROWS_WITH_AS(_ = json::parse("/*", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid comment; missing closing '*/'; last read: '/*'", json::parse_error); +#endif } #if JSON_DIAGNOSTIC_POSITIONS diff --git a/tests/src/unit-deserialization.cpp b/tests/src/unit-deserialization.cpp index 77d49c08b..0dbfdd15c 100644 --- a/tests/src/unit-deserialization.cpp +++ b/tests/src/unit-deserialization.cpp @@ -8,6 +8,14 @@ #include "doctest_compatibility.h" +// capture whether JSON_STRICT_NUL_HANDLING was enabled on the command line +// (e.g. -DJSON_STRICT_NUL_HANDLING=1) *before* including json.hpp, since the +// library #undefs JSON_STRICT_NUL_HANDLING itself once the header has been +// fully processed (see include/nlohmann/detail/macro_unscope.hpp) +#if defined(JSON_STRICT_NUL_HANDLING) && (JSON_STRICT_NUL_HANDLING == 1) + #define JSON_TEST_STRICT_NUL_HANDLING_ENABLED 1 +#endif + #include using nlohmann::json; #ifdef JSON_TEST_NO_GLOBAL_UDLS @@ -323,6 +331,23 @@ TEST_CASE("deserialization") CHECK(j == json({"foo", 1, 2, 3, false, {{"one", 1}}})); } + SECTION("operator>> with a NUL byte after the value (issue #5530)") + { + // operator>> parses non-strictly (it does not require the whole + // stream to be consumed), so a NUL byte following a complete + // value is simply left unread on the stream and never reaches + // the "expected end of input" check that JSON_STRICT_NUL_HANDLING + // affects; this holds regardless of the macro (verified below for + // the opt-in state as well) + std::string data = "123"; + data.push_back('\0'); + std::istringstream ss(data); + json j; + ss >> j; + CHECK(j == json(123)); + CHECK(ss.good()); + } + SECTION("user-defined string literal") { CHECK("[\"foo\",1,2,3,false,{\"one\":1}]"_json == json({"foo", 1, 2, 3, false, {{"one", 1}}})); @@ -405,6 +430,27 @@ TEST_CASE("deserialization") CHECK_THROWS_WITH_AS(ss >> j, "[json.exception.parse_error.101] parse error at line 1, column 29: syntax error while parsing array - unexpected end of input; expected ']'", json::parse_error&); } +#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED) + SECTION("operator>> with a NUL byte where a value is expected (JSON_STRICT_NUL_HANDLING == 1, issue #5530)") + { + // a trailing NUL byte *after* a complete value is unaffected by the + // macro (see the successful-deserialization "operator>> with a NUL + // byte after the value" section above): operator>> parses + // non-strictly and never reaches the "expected end of input" check + // that the macro changes. A NUL byte where a *value* is expected, + // however, goes through the same token dispatch as any other input + // and is affected: with the macro enabled it now raises + // parse_error.101 (like any other unrecognized byte) instead of + // being silently treated the same as an empty stream. + std::string const data(1, '\0'); + std::istringstream ss(data); + json j; + CHECK_THROWS_WITH_AS(ss >> j, + "[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: ''", + json::parse_error&); + } +#endif + SECTION("user-defined string literal") { CHECK_THROWS_WITH_AS("[\"foo\",1,2,3,false,{\"one\":1}"_json, "[json.exception.parse_error.101] parse error at line 1, column 29: syntax error while parsing array - unexpected end of input; expected ']'", json::parse_error&); @@ -453,7 +499,11 @@ TEST_CASE("deserialization") SECTION("from std::array") { - std::array const v { {'t', 'r', 'u', 'e'} }; + // sized to exactly the length of "true": a size of 5 would leave + // a value-initialized trailing 0x00 element that is only + // silently accepted as end-of-input by default and would fail + // under JSON_STRICT_NUL_HANDLING + std::array const v { {'t', 'r', 'u', 'e'} }; CHECK(json::parse(v) == json(true)); CHECK(json::accept(v)); @@ -549,7 +599,9 @@ TEST_CASE("deserialization") SECTION("from std::array") { - std::array v { {'t', 'r', 'u', 'e'} }; + // sized to exactly the length of "true", see the analogous + // "from std::array" section above for why + std::array v { {'t', 'r', 'u', 'e'} }; CHECK(json::parse(std::begin(v), std::end(v)) == json(true)); CHECK(json::accept(std::begin(v), std::end(v))); diff --git a/tests/src/unit-regression2.cpp b/tests/src/unit-regression2.cpp index 2c0cf6549..b128b7a73 100644 --- a/tests/src/unit-regression2.cpp +++ b/tests/src/unit-regression2.cpp @@ -606,7 +606,11 @@ TEST_CASE("regression tests 2") SECTION("issue #2546 - parsing containers of std::byte") { const char DATA[] = R"("Hello, world!")"; // NOLINT(misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) - const auto s = std::as_bytes(std::span(DATA)); + // exclude the trailing '\0' that string-literal initialization adds to + // DATA: std::span(DATA) would span the full array extent (including + // that NUL), which is only silently accepted as end-of-input by default + // and would fail under JSON_STRICT_NUL_HANDLING + const auto s = std::as_bytes(std::span(DATA, sizeof(DATA) - 1)); const json j = json::parse(s); CHECK(j.dump() == "\"Hello, world!\""); }