Read binary strings/blobs in bulk chunks with a memcpy fast path

The binary reader read CBOR/MessagePack/BSON/UBJSON strings and byte
arrays one byte at a time via get()/push_back(), and the iterator input
adapter's get_elements() fallback was itself a per-byte loop, so even
contiguous inputs never benefited from a block copy.

Two changes:

1. iterator_input_adapter::get_elements() gains a contiguous fast path
   that copies the whole requested range with std::memcpy. Contiguity is
   detected via std::is_pointer (all standards) and, in C++20,
   std::contiguous_iterator (so std::vector/std::string iterators also
   qualify). Non-contiguous iterators keep the element-by-element loop.

2. get_string()/get_binary() now share get_bytes(), which reads into the
   result in bounded chunks through get_elements() instead of byte by
   byte. Capping the chunk size preserves the deliberate "do not
   reserve(len) for an untrusted length" DoS protection while turning the
   inner loop into block copies. The min(chunk_size, len) computation is
   width-safe so narrow length types (e.g. MessagePack ext-8's uint8_t)
   cannot truncate chunk_size to zero.

Microbenchmark (2 MiB string + 2 MiB blob + 2000x1 KiB strings, Apple
clang, -O2):

  C++20  from_cbor(vector)   20.1 ms -> 1.0 ms  (~20x)
         from_cbor(pointer)  18.9 ms -> 1.0 ms  (~19x)
  C++17  from_cbor(pointer)  18.7 ms -> 1.0 ms  (~19x, memcpy)
         from_cbor(vector)   18.7 ms -> 4.0 ms  (~4.6x, tight loop)

Behavior is unchanged: truncated input still throws parse_error.110 at
the same byte offset, and all binary-format unit tests pass.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Niels Lohmann
2026-07-04 00:00:12 +02:00
co-authored by Claude Opus 4.8
parent c363dc3e4d
commit 9eee7a7b9b
4 changed files with 216 additions and 46 deletions
+40
View File
@@ -2778,3 +2778,43 @@ TEST_CASE("Tagged values")
CHECK(!jb["binary"].get_binary().has_subtype());
}
}
TEST_CASE("CBOR large strings and binaries (chunked reader)")
{
// The binary reader reads strings and byte arrays in bounded chunks; make
// sure roundtripping is correct for lengths around and beyond the internal
// chunk size (4096 bytes), for both vector (iterator) and pointer inputs.
for (const std::size_t len :
{
std::size_t{0}, std::size_t{1}, std::size_t{4095}, std::size_t{4096},
std::size_t{4097}, std::size_t{8192}, std::size_t{100000}
})
{
CAPTURE(len);
// text string
const json j_string = std::string(len, 'x');
const std::vector<std::uint8_t> v_string = json::to_cbor(j_string);
CHECK(json::from_cbor(v_string) == j_string);
// pointer input exercises the std::memcpy fast path
CHECK(json::from_cbor(reinterpret_cast<const char*>(v_string.data()),
reinterpret_cast<const char*>(v_string.data()) + v_string.size()) == j_string);
// byte string
const json j_binary = json::binary(std::vector<std::uint8_t>(len, 0xCD));
const std::vector<std::uint8_t> v_binary = json::to_cbor(j_binary);
CHECK(json::from_cbor(v_binary) == j_binary);
CHECK(json::from_cbor(reinterpret_cast<const char*>(v_binary.data()),
reinterpret_cast<const char*>(v_binary.data()) + v_binary.size()) == j_binary);
// a truncated payload must still be reported as an error, never crash
// or loop, regardless of the (large) announced length
if (len > 16)
{
std::vector<std::uint8_t> truncated = v_string;
truncated.resize(truncated.size() - 8);
json _;
CHECK_THROWS_AS(_ = json::from_cbor(truncated), json::parse_error);
}
}
}