Update the discard_number_values comments to the current number path

The comments explaining the accept() shortcut in convert_number() and
the member documentation still argued in terms of strtoull()/strtoll()
and errno, which #5283 replaced with convert_integer(), and pointed at
scan_number() instead of convert_number(). They also did not say that
scan_number_bulk_contiguous() converts integers itself, so the shortcut
is only reached for input without bulk access, with
JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back.

Rewrite both comments to describe the digit-count check in front of
convert_integer(), keeping the 18-digit bound and the json_sax_acceptor
argument. The stale <cstdlib> comment is left for after #5616, which
edits that include block.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-30 09:47:01 +02:00
parent 50d8e71b9f
commit 5e41bf284a
2 changed files with 54 additions and 84 deletions
+27 -42
View File
@@ -1494,45 +1494,30 @@ scan_number_done:
*/ */
token_type convert_number(token_type number_type, std::size_t mantissa_end) token_type convert_number(token_type number_type, std::size_t mantissa_end)
{ {
// If the caller does not need the converted value (only whether the // accept() only needs to know whether the input is valid, so it sets
// input is syntactically valid; see json_sax_acceptor/accept()), an // discard_number_values (see json.hpp), and an integer token whose
// unsigned/integer token can be reported without calling // digit count shows that it fits is reported without calling
// strtoull()/strtoll() at all, *provided* we can already tell from // convert_integer(). A number with up to 18 digits always fits into
// the digit count alone that the conversion cannot overflow 64 bits. // both std::uint64_t and std::int64_t (18 nines is about 1e18, below
// Such tokens are always finite and are accepted unconditionally by // INT64_MAX, which is about 9.2e18). Longer tokens take the exact path
// the parser regardless of their actual value (parser::sax_parse_internal() // below, including the fallback to floating point when the value does
// never checks finiteness for value_unsigned/value_integer), so the // not fit.
// classification below is all that is needed.
// //
// A decimal number with up to 18 digits is always representable in // With a narrower number_unsigned_t/number_integer_t (e.g.
// both std::uint64_t and std::int64_t (18 nines is ~1e18, well below // std::uint32_t), the exact path would reclassify some of these tokens
// both UINT64_MAX ~1.8e19 and INT64_MAX ~9.2e18), so strtoull()/strtoll() // as (finite) floats, while this check reports integers. That does not
// could not have set errno to ERANGE for it. Numbers with more digits // change the result of accept(): it always parses through
// (rare in practice) fall through to the exact code below, unchanged, // json_sax_acceptor, whose number callbacks discard their argument and
// so their handling -- including reclassification to value_float when // return true, and the parser rejects neither integers nor finite
// the value overflows 64 bits, and rejection when it is not even // floats. value_unsigned/value_integer are left unset here, so a caller
// finite as a double -- is bit-for-bit identical to before this // that reads the converted value must not set discard_number_values.
// optimization.
// //
// Note this reasons about std::uint64_t/std::int64_t, not about // On contiguous input, scan_number_bulk_contiguous() converts integer
// number_unsigned_t/number_integer_t (BasicJsonType's own, possibly // tokens itself and does not pass them to this function, unless
// narrower, template parameters -- e.g. std::uint32_t). That is fine // JSON_DIAGNOSTIC_POSITIONS is enabled. This check is therefore only
// *only* because discard_number_values is exclusively set by // reached for input without bulk access (e.g. streams), with
// accept() (see json.hpp), and accept() always parses through the // JSON_DIAGNOSTIC_POSITIONS, or when scan_number_bulk_contiguous()
// library's own json_sax_acceptor -- never a user-supplied SAX // falls back to scan_number().
// consumer -- whose number_unsigned()/number_integer()/number_float()
// callbacks unconditionally discard their argument and return true.
// So for every caller that can reach this branch, neither the token
// classification below nor the eventual (possibly narrowed, and on
// this fast path left stale/unset) value_unsigned/value_integer is
// ever consulted -- an unsigned/integer token is accepted outright,
// and even a >18-digit token that this fast path deliberately falls
// through for is, once reclassified to value_float, still finite
// (and thus accepted) for any digit count that fits in number_unsigned_t
// or number_integer_t regardless of that type's width. If this
// function is ever taught to run with discard_number_values true for
// a caller that *does* read the converted value, this reasoning (and
// the fast path below) would need to be revisited.
if (discard_number_values) if (discard_number_values)
{ {
constexpr std::size_t safe_digit_count = 18; constexpr std::size_t safe_digit_count = 18;
@@ -2325,11 +2310,11 @@ scan_number_done:
/// the position of the decimal point in token_buffer /// the position of the decimal point in token_buffer
std::size_t decimal_point_position = std::string::npos; std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the /// whether the caller only needs the token types and never looks at the
/// token classification and never looks at the converted numeric value; /// converted numeric values; set only by accept(), which parses through
/// when set, scan_number() may skip strtoull()/strtoll() for /// json_sax_acceptor. When set, convert_number() skips converting integer
/// value_unsigned/value_integer tokens whose digit count guarantees they /// tokens whose digit count guarantees that they fit into 64 bits (see
/// fit into 64 bits (see scan_number()) /// there)
const bool discard_number_values = false; const bool discard_number_values = false;
}; };
+27 -42
View File
@@ -10586,45 +10586,30 @@ scan_number_done:
*/ */
token_type convert_number(token_type number_type, std::size_t mantissa_end) token_type convert_number(token_type number_type, std::size_t mantissa_end)
{ {
// If the caller does not need the converted value (only whether the // accept() only needs to know whether the input is valid, so it sets
// input is syntactically valid; see json_sax_acceptor/accept()), an // discard_number_values (see json.hpp), and an integer token whose
// unsigned/integer token can be reported without calling // digit count shows that it fits is reported without calling
// strtoull()/strtoll() at all, *provided* we can already tell from // convert_integer(). A number with up to 18 digits always fits into
// the digit count alone that the conversion cannot overflow 64 bits. // both std::uint64_t and std::int64_t (18 nines is about 1e18, below
// Such tokens are always finite and are accepted unconditionally by // INT64_MAX, which is about 9.2e18). Longer tokens take the exact path
// the parser regardless of their actual value (parser::sax_parse_internal() // below, including the fallback to floating point when the value does
// never checks finiteness for value_unsigned/value_integer), so the // not fit.
// classification below is all that is needed.
// //
// A decimal number with up to 18 digits is always representable in // With a narrower number_unsigned_t/number_integer_t (e.g.
// both std::uint64_t and std::int64_t (18 nines is ~1e18, well below // std::uint32_t), the exact path would reclassify some of these tokens
// both UINT64_MAX ~1.8e19 and INT64_MAX ~9.2e18), so strtoull()/strtoll() // as (finite) floats, while this check reports integers. That does not
// could not have set errno to ERANGE for it. Numbers with more digits // change the result of accept(): it always parses through
// (rare in practice) fall through to the exact code below, unchanged, // json_sax_acceptor, whose number callbacks discard their argument and
// so their handling -- including reclassification to value_float when // return true, and the parser rejects neither integers nor finite
// the value overflows 64 bits, and rejection when it is not even // floats. value_unsigned/value_integer are left unset here, so a caller
// finite as a double -- is bit-for-bit identical to before this // that reads the converted value must not set discard_number_values.
// optimization.
// //
// Note this reasons about std::uint64_t/std::int64_t, not about // On contiguous input, scan_number_bulk_contiguous() converts integer
// number_unsigned_t/number_integer_t (BasicJsonType's own, possibly // tokens itself and does not pass them to this function, unless
// narrower, template parameters -- e.g. std::uint32_t). That is fine // JSON_DIAGNOSTIC_POSITIONS is enabled. This check is therefore only
// *only* because discard_number_values is exclusively set by // reached for input without bulk access (e.g. streams), with
// accept() (see json.hpp), and accept() always parses through the // JSON_DIAGNOSTIC_POSITIONS, or when scan_number_bulk_contiguous()
// library's own json_sax_acceptor -- never a user-supplied SAX // falls back to scan_number().
// consumer -- whose number_unsigned()/number_integer()/number_float()
// callbacks unconditionally discard their argument and return true.
// So for every caller that can reach this branch, neither the token
// classification below nor the eventual (possibly narrowed, and on
// this fast path left stale/unset) value_unsigned/value_integer is
// ever consulted -- an unsigned/integer token is accepted outright,
// and even a >18-digit token that this fast path deliberately falls
// through for is, once reclassified to value_float, still finite
// (and thus accepted) for any digit count that fits in number_unsigned_t
// or number_integer_t regardless of that type's width. If this
// function is ever taught to run with discard_number_values true for
// a caller that *does* read the converted value, this reasoning (and
// the fast path below) would need to be revisited.
if (discard_number_values) if (discard_number_values)
{ {
constexpr std::size_t safe_digit_count = 18; constexpr std::size_t safe_digit_count = 18;
@@ -11417,11 +11402,11 @@ scan_number_done:
/// the position of the decimal point in token_buffer /// the position of the decimal point in token_buffer
std::size_t decimal_point_position = std::string::npos; std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the /// whether the caller only needs the token types and never looks at the
/// token classification and never looks at the converted numeric value; /// converted numeric values; set only by accept(), which parses through
/// when set, scan_number() may skip strtoull()/strtoll() for /// json_sax_acceptor. When set, convert_number() skips converting integer
/// value_unsigned/value_integer tokens whose digit count guarantees they /// tokens whose digit count guarantees that they fit into 64 bits (see
/// fit into 64 bits (see scan_number()) /// there)
const bool discard_number_values = false; const bool discard_number_values = false;
}; };