Keep the first of duplicate keys in the object hash index again

Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a69542046; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann committed 2026-10-09 23:33:31 +02:00
1 parent 42b5e4ac34
commit 61e3d97bfd
6 files changed
+27 -27

No files matched your search

@@ -116,6 +116,7 @@ document.
member in order and so keep the *last* value for a repeated key, so `#!cpp v["a"]` and
`#!cpp v.materialize()["a"]` can differ. To get the value `parse()` would give, use
[`materialize()`](materialize.md) or iterate the members with [`items()`](items.md) and keep the last match.
The hash index of a larger object (128 members or more) leads to the first member of a key as well.
[`begin()`](begin.md)/[`end()`](end.md) and [`items()`](items.md) iterate over *all* members, including
duplicates, in document order. See [Duplicate keys](../../features/json_view.md#duplicate-keys) and
[`size()`](size.md#notes).
+4 -3
View File
@@ -165,9 +165,10 @@ document: `#!cpp auto v = json_document::parse(text).root();` does not compile.
**Why the first?** A lookup can stop as soon as it finds a match. Returning the last member would force every lookup
to scan all members of the object, even when the key is found at the very first one: this made lookups in small
objects 1.6 to 3.4 times slower. Other zero-copy parsers that index the source text, such as yyjson and simdjson, also
return the first member. RFC 8259 only says that names within an object SHOULD be unique and that the behavior of a
receiver that sees duplicates is unpredictable, so neither choice is wrong.
objects 1.6 to 3.4 times slower. The hash index of objects with 128 members or more leads to the first member of a
key as well. Other zero-copy parsers that index the source text, such as yyjson and simdjson, also return the first
member. RFC 8259 only says that names within an object SHOULD be unique and that the behavior of a receiver that sees
duplicates is unpredictable, so neither choice is wrong.
**What stays the same as `parse()`?** [`materialize()`](../api/basic_json_view/materialize.md) and
[`get<std::map<...>>()`](../api/basic_json_view/get.md) replay every member in order, so they keep the *last* value
+1 -1
View File
@@ -215,7 +215,7 @@ packet-beta
- **Large objects** (128 members or more) get a hash index after parsing
([`detail/view/object_index.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/view/object_index.hpp)):
an open-addressing table whose slots hold the distance from the object's node to a key's node, so that a lookup does
not compare every key. Of duplicate keys the table leads to the last, as a linear search does. The object's `extra`
not compare every key. Of duplicate keys the table leads to the first, as a linear search does. The object's `extra`
holds the number of its table. Only 65,535 tables fit into `extra`; objects beyond them are searched linearly.
For example, `#!json {"a": [1, 2.5]}` becomes five nodes. Each node's elements follow it, and `next` leads from an