Commit Graph
28 Commits
Author SHA1 Message Date
Niels Lohmann 3926fcaac3 Deduplicate serializer dump code; fix stale includes, docs, and lint (#5729)
* Share scalar serialization between dump_internal and dump_value

dump_value()'s cases for string, binary, boolean, number_integer,
number_unsigned, number_float, discarded and null were a byte-for-byte
copy of dump_internal()'s (added together in #5285 for the iterative
fallback path). Any future change to scalar output had to be made in
both places, or the recursive and depth-limited paths would silently
start producing different bytes.

Extract the shared cases into a private dump_scalar() and have both
dump_internal() and dump_value() call it. Output is unchanged: dump(),
dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t
variants are byte-identical over the json_test_data corpus before and
after, and dump() throughput on a scalar-heavy document is unaffected.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop serializer.hpp's dependency on binary_writer.hpp

The only use of binary_writer in serializer.hpp was
binary_writer<BasicJsonType, char>::to_char_type() to write the
U+FFFD replacement character's three bytes. With CharType=char this
is an identity conversion, so the include of binary_writer.hpp (and
transitively binary_reader.hpp) pulled in a large, unrelated header
for a no-op call.

Write the three bytes directly instead. serializer.hpp compiles
standalone with -Wall -Wextra -Werror, with and without
-funsigned-char, and unit-serialization's error_handler_t::replace
cases (with and without ensure_ascii) still pass. Moving
binary_writer's to_char_type/to_msgpack_length to its private section
is left as an optional follow-up.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale #include lines in the output headers

output_adapters.hpp included <algorithm> and <iterator> for std::copy
and std::back_inserter, which have not been used there since #3569
(2022). serializer.hpp included <algorithm> for std::reverse (also
unused), <cmath> for labs/isnan/signbit (only std::isfinite is used)
and <utility> for std::move (nothing from <utility> is used there),
while using std::next without including <iterator> at all, relying on
getting it transitively through output_adapters.hpp's own stale
<iterator>.

Drop the unused includes, add <iterator> for std::next, and correct
the remaining include comments. Both headers still compile standalone
with -Wall -Wextra -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides

output_vector_adapter, output_stream_adapter and output_string_adapter
declared their write_characters(const CharType*, std::size_t) override
JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with
a null pointer and length 0 for an empty string or binary value; the
type-erased call path only stayed silent under UBSan because the
static callee at those call sites is the unattributed virtual base.
A nonnull attribute on a definition lets GCC and Clang assume the
parameter is non-null inside the function body even when the call is
virtual, so this was latent undefined behavior, not just style.

Drop the attribute from the three overrides and document the
(nullptr, 0) contract on output_adapter_protocol::write_characters.
unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise
empty binary/string payloads through the stream and vector/string
adapters) pass under -fsanitize=address,undefined,nonnull-attribute.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update stale serializer doc comments to match the current implementation

dump_internal()'s doc block still described the pre-#5285/#5449
implementation: an escape_string() function that does not exist
(the function is dump_escaped), integer conversion "implicitly via
operator<<" (dump_integer actually uses a digit-pair lookup table),
and floating-point conversion via "%g" (IEEE-754 types go through
to_chars, others through snprintf). dump_value()'s comment said
elements are pushed for dump_internal to walk, but it is
dump_iteratively() that walks the stack. dump_escaped(), dump_integer()
and dump_float() each said they write "to output stream @a o", which
has not been true since the writer moved to write_buffer.

Doc-only change; no behavior, API or ABI impact.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge duplicate byte-to-hex helper and drop stale '| 0' promotions

serializer::hex_bytes() and binary_writer::hex_byte() had identical
bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it
from both. Also drop the `| 0` at the two serializer call sites
(hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 (7440786b8)
added it so that `ss << std::hex << (byte | 0)` printed a number
rather than a char with the old stringstream writer; the int result
just narrows back to uint8_t now, so it was a no-op.

Behavior is unchanged: unit-serialization, unit-bon8 (whose
type_error.316 messages exercise this code) and unit-diagnostics
pass, and both headers still compile standalone with -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Trim two stale lint suppressions in serializer.hpp

dump_integer()'s `auto buffer_ptr = number_buffer.begin();` carried
NOLINT entries for cppcoreguidelines-pro-type-vararg and hicpp-vararg,
left over from the snprintf-based implementation (#3088); there is no
variadic call on that line, so keep only the qualified-auto
suppressions it actually needs. remove_sign()'s assert checked
`x < 0 && x < (std::numeric_limits<number_integer_t>::max)()) `with a
NOLINT(misc-redundant-expression) to hide it; the second conjunct is
always true once x < 0, and has been since 6ce2f35ba (2019), so
reduce the assert to `x < 0` and drop the suppression instead of
masking it.

Both are documentation-only changes to assertions/suppressions, not
behavior. The to_chars.hpp `#if 0` branch this item also flagged is
left alone, next to draft PR #5634's pending hunk.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the Hoehrmann SPDX copyright line to string_utils.hpp

serializer.hpp carried the SPDX-FileCopyrightText line for Björn
Hoehrmann's UTF-8 decoder, but the decoder (decode() and the utf8d
table) has lived in string_utils.hpp since #5185 (d19f7f5dc);
serializer.hpp now only calls decode(). Move the copyright line to
where the code it covers actually is.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop calling std::localeconv() on every dump()

The serializer constructor snapshotted std::localeconv() into a
locale_chars member on every dump(), even though the only reader is
dump_float(number_float_t, std::false_type)'s snprintf path, taken
only for a number_float_t that is neither IEEE single nor double.
localeconv() is not required to be thread-safe with setlocale(), so
every dump() paid for and raced on a lookup that almost never mattered.

Remove locale_chars and the locale member. Right before the
thousands-separator/decimal-point fixups in the snprintf path, read
std::localeconv() into local thousands_sep/decimal_point variables
(null-checked, first byte only, as before) - the same way
lexer::get_decimal_point() already does since #5597. Output is
unchanged unless the locale changes during a single dump(); in that
case the fixups now match what snprintf just produced, instead of a
value snapshotted before the call.

Overlaps draft PR #5608, which touches the same constructor and
dump_float() lines to move this code into a new
dump_float_snprintf(); this lands the lookup change now as #5709 asks,
and #5608 can do the lookup inside dump_float_snprintf() when it
rebases.

Verification: the full json_test_data corpus (742 files, dump(),
dump(4) and dump(-1,' ',true)) is byte-identical to before the change
under the C locale. Added a test pinning the new per-conversion
lookup: it switches LC_NUMERIC mid-dump() (via a streambuf that
switches on its first write, after the serializer's write buffer has
been flushed once but before a later float is converted) and checks
the decimal point is still normalized using the locale active at
conversion time. On a platform where long double is IEEE-754 double
(e.g. 64-bit Arm), dump_float() takes the locale-independent
to_chars() path and the test is a no-op there; it is meaningful on a
platform where long double is extended precision (most x86 targets).

Part of #5709 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:31:42 +02:00
Niels LohmannandClaude 1054b2097e Speed up binary writing: value-type output sink + byte-swap number encoding (#5286)
* Devirtualize binary_writer via a value-type output sink

to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson wrote every byte through
output_adapter_t, a shared_ptr<output_adapter_protocol> whose
write_character/write_characters are virtual. Unlike the lexer (templated
on a concrete InputAdapterType), the binary writer never got that
treatment, so binary output paid a vtable lookup per byte and a
make_shared per call.

Template binary_writer on an OutputSinkType and give it two concrete,
non-virtual sinks:

- output_vector_sink: appends straight into a std::vector (push_back /
  insert), used by the vector-returning to_* convenience functions. No
  vtable, no shared_ptr; the writes inline.
- output_adapter_sink: forwards to a type-erased output_adapter_t, so the
  existing to_*(j, output_adapter) overloads (streams, strings, custom
  adapters) keep working exactly as before -- one virtual call each,
  unchanged.

binary_writer keeps a convenience constructor taking output_adapter_t
(building the default output_adapter_sink), so the adapter overloads are
untouched; only the convenience functions switch to the vector sink. The
friend declaration and the basic_json binary_writer alias gain the new
(defaulted) template parameter.

Output is byte-for-byte identical: verified across ~3000 randomized
values plus curated edge cases (all scalar widths, strings with invalid
UTF-8, binary, nested arrays/objects) for CBOR, MessagePack, UBJSON (both
size/type settings), BJData, and BSON, plus the output_adapter path, in
C++11/17/20. Warning-clean under clang -Weverything and the gcc pedantic
set; clang-tidy clean on the changed headers; make check-amalgamation
clean.

Throughput (g++ -O3, vs develop): scalar-dense binary output such as
integer arrays ~1.4x; many small to_cbor calls ~1.04x (DOM traversal
bound); string/blob-heavy output unchanged (already bulk-bound). No
workload regressed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI failures from binary_writer output-sink change

Four CI jobs failed on the initial commit; all are addressed here without
changing any output (binary encodings remain byte-for-byte identical to
develop across the differential corpus):

1. ci_test_gcc / cuda (-Werror=duplicated-branches): for number_float_t ==
   float, static_cast<float>(n) is the identity, so write_compact_float's
   two branches are intentionally identical. Once the concrete vector sink
   is inlined, GCC constant-folds and diagnoses this (the type-erased path
   hid it behind a non-inlined virtual call). Silence -Wduplicated-branches
   for GCC (clang has no such warning) alongside the existing -Wfloat-equal
   pragma.

2. ci_static_analysis_clang (UBSan nonnull-attribute): binary_writer passes
   a null pointer with length 0 for empty strings/binary. output_vector_sink
   / output_adapter_sink declared write_characters JSON_HEDLEY_NON_NULL, so
   the sanitizer flagged the (harmless) zero-length call once the sink was
   called directly rather than through the attribute-free virtual base. Drop
   the attribute from both sinks, matching the pre-existing behavior.

3. ci_cpplint (build/include_what_you_use): output_adapter_sink uses
   std::move; add #include <utility>.

4. ci_cuda_example (nvcc 11.8): NVCC's front end rejects the default
   template argument on the binary_writer alias template. Revert the alias
   to its original single-parameter form (relying on binary_writer's own
   defaulted OutputSinkType) and spell out the full type in the vector-sink
   convenience functions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Encode big-endian numbers with a byte swap instead of std::reverse

write_number() reordered multi-byte numbers for the big-endian formats
(CBOR/MessagePack/UBJSON) with std::reverse over the byte array. GCC
lowered only some sizes to a bswap; clang kept a scalar byte shuffle
(0 bswap instructions in the CBOR number path). Replace the reverse with
size-dispatched __builtin_bswap16/32/64 helpers (portable shift fallback
for other compilers; std::reverse retained for exotic sizes such as a
long double number_float_t).

Codegen: the CBOR number path now emits bswap on both compilers
(gcc 2 -> 16, clang 0 -> 4). Output is byte-for-byte identical to the
previous implementation across the binary differential corpus.

Throughput (isolated vs the std::reverse version, best of 9):
  CBOR int64 array   gcc +7%   clang +10%
  CBOR uint16 array  gcc +27%  clang flat

Modest but consistent on number-dense encodings; negligible on
string/blob-heavy output, as expected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Reserve output capacity up front for binary serialization

The vector-returning to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson grew
the output buffer purely by geometric reallocation. Reserving an estimate
up front avoids the early reallocations, which is the dominant per-byte
cost for array/object-heavy output.

The estimate (binary_reserve_hint) is deliberately conservative and safe
against untrusted input: it consults only the top-level element count
(O(1), no walk of the DOM), guards the multiplication against overflow,
and clamps the result to a fixed 1 MiB ceiling, so a large or hostile DOM
can never force an oversized allocation here. The buffer still grows
geometrically past the hint, so an underestimate only costs a few later
reallocations; scalars/strings/binary are written in one shot and get no
hint. Reserving capacity does not change the bytes produced.

Throughput (g++/clang -O3, vs the previous commit):
  cbor int array     +10% / +13%
  cbor object array  +20% / +38%

Output is byte-for-byte identical to develop across the binary
differential corpus.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review findings on the binary writer output sinks

- binary_reserve_hint(): the 4-bytes-per-element estimate over-reserved by up
  to 4x for arrays of small scalars (CBOR encodes 0..23 in one byte), and the
  returned vector kept that capacity. Make the hint a strict lower bound on the
  encoded size instead, which also removes the 1 MiB clamp whose branch no test
  could reach (the largest container in the suite has 65793 elements).

- Guard the -Wduplicated-branches pragma with __GNUC__ >= 7. The warning does
  not exist before GCC 7, so naming it made GCC 4.8/4.9/5/6 - which the CI
  matrix still builds - warn under -Wpragmas on every including translation
  unit, breaking downstream -Werror builds.

- Constrain the adapter constructor of binary_writer with the enable_if its
  documentation already claimed, so a writer over some other sink type is no
  longer advertised as constructible from an output adapter.

- Let output_vector_adapter wrap output_vector_sink rather than duplicating the
  append logic, so the type-erased and templated paths share one implementation.

- Collapse the three copies of the memcpy/byte_swap/memcpy dance into a single
  byte_swap_buffer() helper, and add the MSVC _byteswap_* intrinsics so MSVC no
  longer falls back to the scalar shuffle this change exists to eliminate.

- Add a vector_writer() helper for the five vector-returning to_* overloads
  instead of spelling out the writer type at each call site, and drop a dead
  default member initializer on output_adapter_sink.

- New tests: the vector sink and the adapter sink must produce identical bytes
  for every format (the two to_* overloads no longer delegate to each other and
  could otherwise drift), and binary_reserve_hint() must never exceed the size
  actually written.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Route the -Wduplicated-branches pragma through Hedley

Match #5485, which moved the binary writer's hand-rolled diagnostic
pragmas onto JSON_HEDLEY_PRAGMA (merged into develop while this branch
was open). The devirtualization's -Wduplicated-branches suppression in
write_compact_float was the one raw '#pragma GCC diagnostic' left; it
now uses JSON_HEDLEY_PRAGMA like the adjacent -Wfloat-equal line, still
guarded to GCC >= 7 and non-clang (the warning exists only there).

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-23 08:59:46 +02:00
Niels Lohmann 515d994acb 📄 adjust year (#5044)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-01-01 20:00:39 +01:00
Niels Lohmann 54be9b04f0 📄 update REUSE (#4960) 2025-10-23 06:56:36 +02:00
Niels Lohmann 1705bfe914 🔖 set version to 3.12.0 (#4727)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2025-04-11 10:41:14 +02:00
Niels Lohmann f06604fce0 Bump the copyright years (#4606)
* 📄 bump the copyright years

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 📄 bump the copyright years

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* 📄 bump the copyright years

Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
2025-01-19 17:04:17 +01:00
Niels Lohmann 1b9a9d1f21 Update licenses (#4521)
* 📄 update licenses

* 📄 update licenses
2024-11-29 17:38:42 +01:00
Niels Lohmann 9cca280a4d JSON for Modern C++ 3.11.3 (#4222) 2023-11-28 22:36:31 +01:00
Niels Lohmann f56c6e2e30 Update documentation for the next release (#4216) 2023-11-26 15:51:19 +01:00
Niels Lohmann 9d69186291 🔖 set version to 3.11.2 2022-08-12 15:04:06 +02:00
Niels Lohmann f2020da0dd 🔖 set version to 3.11.1 2022-08-01 23:27:58 +02:00
Niels Lohmann ce0e13ccea 🔖 set version to 3.11.0 2022-07-31 23:19:06 +02:00
Florian Albrechtskirchinger d909f80960 Add versioned, ABI-tagged inline namespace and namespace macros (#3590)
* Add versioned inline namespace

Add a versioned inline namespace to prevent ABI issues when linking code
using multiple library versions.

* Add namespace macros

* Encode ABI information in inline namespace

Add _diag suffix to inline namespace if JSON_DIAGNOSTICS is enabled, and
_ldvcmp suffix if JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON is enabled.

* Move ABI-affecting macros into abi_macros.hpp

* Move std_fs namespace definition into std_fs.hpp

* Remove std_fs namespace from unit test

* Format more files in tests directory

* Add unit tests

* Update documentation

* Fix GDB pretty printer

* fixup! Add namespace macros

* Derive ABI prefix from NLOHMANN_JSON_VERSION_*
2022-07-30 21:59:13 +02:00
Niels LohmannandFlorian Albrechtskirchinger 527da54dcb Use REUSE framework (#3546)
* 📄 add licenses

* 👷 add REUSE compliance check

* 📝 add badge for REUSE

Co-authored-by: Florian Albrechtskirchinger <falbrechtskirchinger@gmail.com>
2022-07-20 12:38:07 +02:00
Romain Reignier d4daaa897f Optimize output vector adapter write (#3569)
* Add benchmark for cbor serialization and cbor binary data serialization

* Use std::vector::insert method instead of std::copy in output_vector_adapter::write_characters

This change increases a lot the performance when writing lots of binary data.
2022-07-08 08:12:00 +02:00
Niels Lohmann 0b345b20c8 Allow allocators for output_vector_adapter (#2989)
* ♻️ allow allocators for vectors

* ✅ add regression tests
2021-09-12 18:55:47 +02:00
David Pfahler aa849a2275 Merge branch 'nlohmann:develop' into without-io 2021-06-14 08:22:49 +02:00
David Pfahler ae9bbbc941 include io only if JSON_NO_IO is not set for #2728 2021-05-31 14:26:45 +02:00
David Pfahler 1a1381f071 Fixes #2728
includes some macros to be defined for using without file io.
2021-04-21 10:24:01 +02:00
Niels Lohmann 6f551930e5 🚨 add new CI and fix warnings (#2561)
* ⚗️ move CI targets to CMake
* ♻️ add target for cpplint
* ♻️ add target for self-contained binaries
* ♻️ add targets for iwyu and infer
* 🔊 add version output
* ♻️ add target for oclint
* 🚨 fix warnings
* ♻️ rename targets
* ♻️ use iwyu properly
* 🚨 fix warnings
* ♻️ use iwyu properly
* ♻️ add target for benchmarks
* ♻️ add target for CMake flags
* 👷 use GitHub Actions
* ⚗️ try to install Clang 11
* ⚗️ try to install GCC 11
* ⚗️ try to install Clang 11
* ⚗️ try to install GCC 11
* ⚗️ add clang analyze target
* 🔥 remove Google Benchmark
* ⬆️ Google Benchmark 1.5.2
* 🔥 use fetchcontent
* 🐧 add target to download a Linux version of CMake
* 🔨 fix dependency
* 🚨 fix includes
* 🚨 fix comment
* 🔧 adjust flags for GCC 11.0.0 20210110 (experimental)
* 🐳 user Docker image to run CI
* 🔧 add target for Valgrind
* 👷 add target for Valgrind tests
* ⚗️ add Dart
* ⏪ remove Dart
* ⚗️ do not call ctest in test subdirectory
* ⚗️ download test data explicitly
* ⚗️ only execute Valgrind tests
* ⚗️ fix labels
* 🔥 remove unneeded jobs
* 🔨 cleanup
* 🐛 fix OCLint call
* ✅ add targets for offline and git-independent tests
* ✅ add targets for C++ language versions and reproducible tests
* 🔨 clean up
* 👷 add CI steps for cppcheck and cpplint
* 🚨 fix warnings from Clang-Tidy
* 👷 add CI steps for Clang-Tidy
* 🚨 fix warnings
* 🔧 select proper binary
* 🚨 fix warnings
* 🚨 suppress some unhelpful warnings
* 🚨 fix warnings
* 🎨 fix format
* 🚨 fix warnings
* 👷 add CI steps for Sanitizers
* 🚨 fix warnings
* ⚡ add optimization to sanitizer build
* 🚨 fix warnings
* 🚨 add missing header
* 🚨 fix warnings
* 👷 add CI step for coverage
* 👷 add CI steps for disabled exceptions and implicit conversions
* 🚨 fix warnings
* 👷 add CI steps for checking indentation
* 🐛 fix variable use
* 💚 fix build
* ➖ remove CircleCI
* 👷 add CI step for diagnostics
* 🚨 fix warning
* 🔥 clean Travis
2021-03-24 07:15:18 +01:00
Niels Lohmann 90798caa62 🚚 rename Hedley macros 2019-07-01 22:37:30 +02:00
Niels Lohmann 897362191d 🔨 add NLOHMANN_JSON prefix and undef macros 2019-07-01 22:24:39 +02:00
Niels Lohmann 1720bfedd1 ⚗️ add Hedley annotations 2019-06-30 22:14:02 +02:00
Niels Lohmann 4e765596f7 🔨 small improvements 2018-10-27 18:31:03 +02:00
Niels Lohmann 858e75c4df 🚨 fixed some clang-tidy warnings 2018-10-07 18:39:18 +02:00
Vitaliy Manushkin 830f3e5290 forward alternative string class from output_adapter to output_string_adapter 2018-03-10 23:58:16 +03:00
Vitaliy Manushkin faccc37d0d dump to alternate implementation of string, as defined in basic_json template 2018-03-10 17:19:28 +03:00
Théo DELRIEU 14cd019861 fix cmake install directory (for real this time)
* Rename 'develop' folder to 'include/nlohmann'
* Rename 'src' folder to 'single_include/nlohmann'
* Use <nlohmann/*> headers in sources and tests
* Change amalgamate config file
2018-02-01 11:06:51 +01:00