Commit Graph
12 Commits
Author SHA1 Message Date
5379e04ce4 Throw type_error.321 when serializing discarded values to binary formats (#5761)
* Throw type_error.321 when serializing discarded values to binary formats

The CBOR, MessagePack, UBJSON, BJData, and BSON writers silently
skipped the payload of a value_t::discarded value nested in an array
or object, while still writing its slot in the element/member count
(and, for BSON, its entry header), producing a binary document whose
declared size does not match what was actually written.

Throw type_error.321 instead, for a discarded value anywhere in the
tree, including at the top level.

Rewritten from the original PR against the current (non-recursive
option aside) binary_writer.hpp, which has changed substantially since
this was first proposed; the out_of_range.412 MessagePack size check
and unrelated test reformatting from that PR are dropped as out of
scope here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document and test type_error.321 for discarded binary values

Add docs for the new exception (home/exceptions.md and the Exceptions
sections of to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson) and test
coverage for a discarded value nested in an array or object, nested
deeper, and (for UBJSON/BJData) inside an optimized same-type array,
for each of CBOR, MessagePack, UBJSON, BJData, and BSON. Adjust the
three pre-existing "discarded" tests that asserted the old silent
behavior (empty/short output) to expect type_error.321 instead.

Co-authored-by: ameliabarnabyhub <312084480+ameliabarnabyhub@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ameliabarnabyhub <ameliabarnabyhub@users.noreply.github.com>
Co-authored-by: ameliabarnabyhub <312084480+ameliabarnabyhub@users.noreply.github.com>
2026-10-06 07:33:34 +02:00
Niels Lohmann 73e9eae3c1 Add an error_handler parameter for UTF-8 to the binary readers and writers (#5746)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 12:13:49 +02:00
Niels Lohmann f56b418c56 Follow each binary format's UTF-8 rule: strict writers (CBOR/UBJSON/BJData/BSON), lenient readers (#5741)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 12:13:48 +02:00
Niels Lohmann 63c10a51fc Review and extend the documentation, and check it in CI (#5638)
* Review and extend the documentation, and check it in CI

A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.

Fixes:
- wrong signatures and version histories (operator!= C++20 member,
  binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
  example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
  suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed

Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
  Mermaid diagrams, new API pages for json_pointer::operator<=> and
  byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release

Checks:
- strict documentation build (broken links/anchors fail it); CI and
  the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
  header links, docset index coverage; its unused-example check works
  again
- all examples produce the same output on every platform

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep the customer links that could not be fixed

A dead link on the customers page is still the evidence of where the
use of the library was documented. Keep the original URLs of the entries
without a working replacement (Marne, Cisco Webex Desk Camera, Philips
Hue, CyberArk) and exclude exactly these URLs from the link check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the duplicate-key recipe's claim about SAX positions

The SAX interface's key() receives no position either; only parse_error()
does. Also note that the recipe does not report the path to the repeated
key (see discussion #5085).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Say the library is available as a single header and mention json_fwd.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct documentation errors found while hunting for bugs

- patch/patch_inplace: list the JSON pointer errors parse_error.106-109
  and out_of_range.402/404, and quote the actual parse_error.105 message.
- unflatten: list parse_error.106/107/108 and out_of_range.404.
- to_bson: list out_of_range.415 (binary subtype above 255) and note
  that 412 and 415 are new in 3.13.0.
- to_string: state that string_t must be convertible to std::string, also
  in the StringType requirements table.
- JSON Lines: a `while (input >> j)` loop also throws after the last value
  for concatenated JSON values; show a loop that works for both.
- BON8: a string gets 0xFF only if nothing follows it in the message; a
  string at the end of an array or object is ended by 0xFE.
- custom_string_type.hpp: add operator+=(char), which the "Always
  required" list asks for (json_pointer::to_string, flatten, unflatten,
  and diff did not compile), and an ADL int_to_string for diff and items.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cache the release headers with functools.lru_cache

Codacy (Pylint) flagged the mutable default argument that header() used
as its cache. functools.lru_cache keeps the same memoization without it.
The script's output is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:32:15 +02:00
Niels Lohmann 5bd766aa50 Move to_bson's binary subtype check into calc_bson_sizes (#5703)
to_bson() rejected a binary value's subtype above 255 (out_of_range.415)
in write_bson_binary(), which only has the binary_t, not the basic_json
value that holds it, so the exception was created with no JSON_DIAGNOSTICS
context even though the equivalent to_msgpack() check names the value's
path. The check also ran after the document size, all preceding elements,
and this element's header and length had already reached the output
adapter, so a caller-provided std::vector or std::string ended up holding
a truncated document.

calc_bson_sizes() already walks every value before anything is written,
to size embedded documents and arrays and to reject invalid keys
(out_of_range.409) up front. The subtype check now runs there instead,
in calc_bson_binary_size(), which is given the basic_json value so the
exception can use it as context. The now-redundant check in
write_bson_binary() is removed, since calc_bson_sizes() always throws
first if any binary value in the document has an oversized subtype.

Fixes #5675.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:41 +02:00
Niels Lohmann 1e101ecac1 Add BON8 support (#2998)
* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 16:56:21 +02:00
Niels Lohmann 95e9a5931c Write BSON in linear time, without recursing per nesting level (#5553)
* Write BSON in linear time, without recursing per nesting level

to_bson() had two problems with nested values:

- It recursed once per nesting level, so a value nested deeply enough -
  100,000 levels on an 8 MiB stack - exhausted the call stack and
  terminated the process, although parse() accepts such values without
  complaint.
- BSON prefixes every document and array with its length. The writer
  computed that length by walking the entire value below it, again for
  every nested document it wrote, which made serializing O(size x depth).
  A 200-level document took 30 ms instead of 1.

Both passes are now iterative, and each length is computed exactly once:

- calc_bson_sizes() computes the length of every document and array in
  one pass, each from the lengths of its entries, into a table ordered
  the way they are written.
- write_bson_document() then writes the document, taking each length from
  the table.

Everything observable is unchanged, as a differential test against
develop confirms byte for byte:

- The same bytes are written.
- A key containing U+0000 still throws out_of_range.409 for the same
  first key, with the same diagnostics path, before anything is written.
- A document too large for BSON still throws out_of_range.412 before
  anything is written.
- A binary subtype above 255 still throws out_of_range.415 after the
  same partial output.

Only the enclosing objects and arrays are kept on a stack, so a flat
document allocates nothing for it. Measured against develop (clang -O3,
median of 201 runs): flat objects unchanged, flat arrays 37% faster (the
array length was computed twice), a nested 3,000-object document 2x
faster, a 200-level document 33x faster.

to_bson.md documented the quadratic complexity since #5334; it is linear
again.

Fixes #5392 for BSON, and #5308.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not require a default-constructible string_t in the BSON writer

GCC 4.9 and MSVC rejected the test's huge_string_t, which has no default
constructor; develop never default-constructed string_t here either.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Let the BSON index-name helper only fill its output parameter

It returned a reference to the string it filled, so callers held a second
name for index_name. Addresses review feedback.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 21:56:39 +02:00
whn 1ac268d409 docs: correct to_bson complexity (#5334)
Signed-off-by: whn <142425816+Whning0513@users.noreply.github.com>
2026-08-28 13:27:18 +01:00
Luke Banicevic 2e23687092 to_bson() silently emits corrupt documents when a length exceeds INT32_MAX (#5314) 2026-07-28 09:08:57 +00:00
Niels Lohmann c5b2b26fdc 📝 fix docs (#5217) 2026-06-29 22:15:18 +02:00
Niels Lohmann 4424a0fcc1 📝 update documentation (#4723)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2025-04-05 18:54:35 +02:00
Niels Lohmann b21c345179 Reorganize directories (#3462)
* 🚚 move files
* 🚚 rename doc folder to docs
* 🚚 rename test folder to tests
2022-05-01 09:41:50 +02:00