mirror of
https://github.com/nlohmann/json.git
synced 2026-10-05 14:21:05 +01:00
Handle numbers that do not fit narrow number types in the binary readers (#5607)
* Handle numbers that do not fit narrow number types in the binary readers With custom number types narrower than the values in a binary document, for example basic_json<..., std::int32_t, std::uint32_t, float>, every binary reader (CBOR, MessagePack, UBJSON, BJData, BSON, BON8) passed the decoded number to the SAX interface with an implicit conversion: the integer 5000000000 silently became 705032704, and a finite double such as 1e300 became infinity. The lexer handles the same values in JSON text: an integer that fits neither integer type is stored as number_float_t, and a finite number that overflows number_float_t is rejected with out_of_range.406. Pass every number read from binary input through three helpers that apply the lexer's rules: - emit_signed(): number_integer_t, else number_unsigned_t for a non-negative value, else number_float_t - emit_unsigned(): number_unsigned_t, else number_float_t - emit_float(): out_of_range.406 if a finite value overflows number_float_t; infinity and NaN are passed on For consistency, a CBOR negative integer below the range of number_integer_t is now stored as number_float_t, like a too small integer in JSON text, instead of being rejected with parse_error.112. With the default number types, this is the only change in behavior. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix MSVC and clang 3.5 in the narrow number type test MSVC types 3000000000 and 5000000000 as unsigned long, so json(-3000000000) triggered C4146 (unary minus on an unsigned type), which /WX turns into an error. Use LL literals, as elsewhere in the tests. clang 3.5 cannot convert the lambdas in the braced initializer of the format table to function pointers. Use named functions instead. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Check integer-to-float fallbacks for overflow in the binary readers emit_signed, emit_unsigned, and the CBOR negative integer fallback now pass their number_float_t fallback through emit_float, so a value that overflows number_float_t is rejected with out_of_range.406 like a floating-point value, instead of silently becoming infinity. This only matters for a number_float_t that cannot represent 2^64, such as a half-precision type. The CBOR value -1 - n is computed as long double so that emit_float sees a finite value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use the input_format member instead of passing the format to binary_reader helpers The helpers (get_number, get_to, get_string, get_binary, get_bytes, emit_signed, emit_unsigned, emit_float, unexpect_eof, exception_message) are members of binary_reader, which already stores the format it was constructed with, so the parameter was redundant. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -55,6 +55,10 @@ This implementation does exactly follow this approach, as it uses double precisi
|
||||
smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally
|
||||
and be serialized to `null`.
|
||||
|
||||
During deserialization (from JSON text or any of the binary formats), a finite number that does not fit into
|
||||
`number_float_t` is rejected with [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406), for
|
||||
example a double-precision number in a binary format when `number_float_t` is `#!cpp float`.
|
||||
|
||||
### Storage
|
||||
|
||||
Floating-point number values are stored directly inside a `basic_json` type.
|
||||
|
||||
@@ -47,8 +47,9 @@ With the default values for `NumberIntegerType` (`std::int64_t`), the default va
|
||||
|
||||
When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and
|
||||
the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of
|
||||
range will yield over/underflow when used in a constructor. During deserialization, too large or small integer numbers
|
||||
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
|
||||
range will yield over/underflow when used in a constructor. During deserialization (from JSON text or any of the binary
|
||||
formats), too large or small integer numbers will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md)
|
||||
or [`number_float_t`](number_float_t.md).
|
||||
|
||||
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
|
||||
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
|
||||
|
||||
@@ -48,8 +48,9 @@ With the default values for `NumberUnsignedType` (`std::uint64_t`), the default
|
||||
|
||||
When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and
|
||||
the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow
|
||||
when used in a constructor. During deserialization, too large or small integer numbers will automatically be stored
|
||||
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
|
||||
when used in a constructor. During deserialization (from JSON text or any of the binary formats), too large or small
|
||||
integer numbers will automatically be stored as [`number_integer_t`](number_integer_t.md) or
|
||||
[`number_float_t`](number_float_t.md).
|
||||
|
||||
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
|
||||
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
|
||||
|
||||
@@ -168,9 +168,9 @@ The library maps CBOR types to JSON value types as follows:
|
||||
!!! warning "Negative integer overflow"
|
||||
|
||||
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
|
||||
result to fit into `number_integer_t` (`std::int64_t` by default), parsing fails with a
|
||||
[`parse_error.112`](../../home/exceptions.md#jsonexceptionparse_error112) exception rather than overflowing
|
||||
silently.
|
||||
result to fit into `number_integer_t` (`std::int64_t` by default), the result is stored as `number_float_t`, like
|
||||
a too small integer in JSON text. For example, `-18446744073709551616` (`0x3B` followed by eight `0xFF` bytes) is
|
||||
stored as `-1.8446744073709552e+19`.
|
||||
|
||||
!!! warning "Object keys"
|
||||
|
||||
|
||||
@@ -331,9 +331,6 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
|
||||
[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1
|
||||
```
|
||||
```
|
||||
[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow
|
||||
```
|
||||
```
|
||||
[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)
|
||||
```
|
||||
|
||||
@@ -863,13 +860,18 @@ The JSON Patch operations 'remove' and 'add' cannot be applied to the root eleme
|
||||
|
||||
### json.exception.out_of_range.406
|
||||
|
||||
A parsed number could not be stored as without changing it to NaN or INF.
|
||||
A parsed number could not be stored without changing it to NaN or INF. For the binary formats, this happens when a
|
||||
finite floating-point number does not fit into [`number_float_t`](../api/basic_json/number_float_t.md), for example a
|
||||
double-precision number when `number_float_t` is `#!cpp float`.
|
||||
|
||||
!!! failure "Example message"
|
||||
!!! failure "Example messages"
|
||||
|
||||
```
|
||||
number overflow parsing '10E1000'
|
||||
```
|
||||
```
|
||||
[json.exception.out_of_range.406] syntax error while parsing CBOR value: number overflow
|
||||
```
|
||||
|
||||
### json.exception.out_of_range.407
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
+283
-211
File diff suppressed because it is too large
Load Diff
@@ -11,7 +11,12 @@
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
|
||||
#include <cmath>
|
||||
#include <fstream>
|
||||
#include <limits>
|
||||
#include <map>
|
||||
#include <string>
|
||||
#include <vector>
|
||||
#include "make_test_data_available.hpp"
|
||||
|
||||
TEST_CASE("Binary Formats" * doctest::skip())
|
||||
@@ -224,3 +229,139 @@ TEST_CASE("Binary Formats" * doctest::skip())
|
||||
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450));
|
||||
}
|
||||
}
|
||||
|
||||
namespace
|
||||
{
|
||||
// the binary formats as function pointers for "Binary formats with narrow number types";
|
||||
// named functions rather than lambdas, because clang 3.5 cannot convert a lambda
|
||||
// to a function pointer in the braced initializer of the format table
|
||||
using narrow_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int32_t, std::uint32_t, float>;
|
||||
using bytes = std::vector<std::uint8_t>;
|
||||
|
||||
bytes encode_cbor(const json& j)
|
||||
{
|
||||
return json::to_cbor(j);
|
||||
}
|
||||
narrow_json decode_cbor(const bytes& v, bool allow_exceptions)
|
||||
{
|
||||
return narrow_json::from_cbor(v, true, allow_exceptions);
|
||||
}
|
||||
|
||||
bytes encode_msgpack(const json& j)
|
||||
{
|
||||
return json::to_msgpack(j);
|
||||
}
|
||||
narrow_json decode_msgpack(const bytes& v, bool allow_exceptions)
|
||||
{
|
||||
return narrow_json::from_msgpack(v, true, allow_exceptions);
|
||||
}
|
||||
|
||||
bytes encode_ubjson(const json& j)
|
||||
{
|
||||
return json::to_ubjson(j);
|
||||
}
|
||||
narrow_json decode_ubjson(const bytes& v, bool allow_exceptions)
|
||||
{
|
||||
return narrow_json::from_ubjson(v, true, allow_exceptions);
|
||||
}
|
||||
|
||||
bytes encode_bjdata(const json& j)
|
||||
{
|
||||
return json::to_bjdata(j);
|
||||
}
|
||||
narrow_json decode_bjdata(const bytes& v, bool allow_exceptions)
|
||||
{
|
||||
return narrow_json::from_bjdata(v, true, allow_exceptions);
|
||||
}
|
||||
|
||||
// BSON can only store numbers as object members
|
||||
bytes encode_bson(const json& j)
|
||||
{
|
||||
return json::to_bson(json{{"a", j}});
|
||||
}
|
||||
narrow_json decode_bson(const bytes& v, bool allow_exceptions)
|
||||
{
|
||||
const auto result = narrow_json::from_bson(v, true, allow_exceptions);
|
||||
return result.is_discarded() ? result : result.at("a");
|
||||
}
|
||||
|
||||
bytes encode_bon8(const json& j)
|
||||
{
|
||||
return json::to_bon8(j);
|
||||
}
|
||||
narrow_json decode_bon8(const bytes& v, bool allow_exceptions)
|
||||
{
|
||||
return narrow_json::from_bon8(v, true, allow_exceptions);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
TEST_CASE("Binary formats with narrow number types")
|
||||
{
|
||||
// Numbers that do not fit the number types are handled like the lexer
|
||||
// handles them in JSON text: an integer that fits neither integer type is
|
||||
// stored as a floating-point number, and a finite floating-point number
|
||||
// that overflows number_float_t is rejected with out_of_range.406.
|
||||
struct binary_format
|
||||
{
|
||||
const char* name;
|
||||
bytes (*encode)(const json&);
|
||||
narrow_json (*decode)(const bytes&, bool);
|
||||
};
|
||||
|
||||
const std::vector<binary_format> formats =
|
||||
{
|
||||
{"CBOR", encode_cbor, decode_cbor},
|
||||
{"MessagePack", encode_msgpack, decode_msgpack},
|
||||
{"UBJSON", encode_ubjson, decode_ubjson},
|
||||
{"BJData", encode_bjdata, decode_bjdata},
|
||||
{"BSON", encode_bson, decode_bson},
|
||||
{"BON8", encode_bon8, decode_bon8},
|
||||
};
|
||||
|
||||
for (const auto& format : formats)
|
||||
{
|
||||
const std::string name = format.name;
|
||||
INFO("format := ", name);
|
||||
const auto roundtrip = [&format](const json & j)
|
||||
{
|
||||
return format.decode(format.encode(j), true);
|
||||
};
|
||||
|
||||
// integers that fit keep their type
|
||||
CHECK(roundtrip(json(-5)).is_number_integer());
|
||||
CHECK(roundtrip(json(-5)).get<std::int32_t>() == -5);
|
||||
CHECK(roundtrip(json(3000000000u)).is_number_unsigned());
|
||||
CHECK(roundtrip(json(3000000000u)).get<std::uint32_t>() == 3000000000u);
|
||||
|
||||
// integers that fit neither integer type are stored as float
|
||||
CHECK(roundtrip(json(5000000000u)).is_number_float());
|
||||
CHECK(roundtrip(json(5000000000u)).get<float>() == 5000000000.0f);
|
||||
if (name != "BON8") // BON8 cannot encode integers above INT64_MAX
|
||||
{
|
||||
CHECK(roundtrip(json(10000000000000000000u)).is_number_float());
|
||||
CHECK(roundtrip(json(10000000000000000000u)).get<float>() == 10000000000000000000.0f);
|
||||
}
|
||||
CHECK(roundtrip(json(-3000000000LL)).is_number_float());
|
||||
CHECK(roundtrip(json(-3000000000LL)).get<float>() == -3000000000.0f);
|
||||
CHECK(roundtrip(json(-5000000000LL)).is_number_float());
|
||||
CHECK(roundtrip(json(-5000000000LL)).get<float>() == -5000000000.0f);
|
||||
|
||||
// floating-point numbers that fit
|
||||
CHECK(roundtrip(json(1.5)).get<float>() == 1.5f);
|
||||
const auto just_above_max = std::nextafter(static_cast<double>((std::numeric_limits<float>::max)()),
|
||||
std::numeric_limits<double>::infinity());
|
||||
CHECK(roundtrip(json(just_above_max)).get<float>() == (std::numeric_limits<float>::max)());
|
||||
|
||||
// infinity and NaN are passed on
|
||||
CHECK(std::isinf(roundtrip(json(std::numeric_limits<double>::infinity())).get<float>()));
|
||||
CHECK(std::isnan(roundtrip(json(std::numeric_limits<double>::quiet_NaN())).get<float>()));
|
||||
|
||||
// finite floating-point numbers that overflow number_float_t are rejected
|
||||
const std::string message = "[json.exception.out_of_range.406] syntax error while parsing " + name
|
||||
+ " value: number overflow";
|
||||
CHECK_THROWS_WITH_AS(roundtrip(json(1e300)), message.c_str(), narrow_json::out_of_range&);
|
||||
CHECK_THROWS_WITH_AS(roundtrip(json(-1e300)), message.c_str(), narrow_json::out_of_range&);
|
||||
CHECK(format.decode(format.encode(json(1e300)), false).is_discarded());
|
||||
}
|
||||
}
|
||||
|
||||
+16
-14
@@ -3185,7 +3185,8 @@ TEST_CASE("Tagged values")
|
||||
// CBOR encodes negative integers as: result = -1 - n
|
||||
// For type 0x3B, n is an 8-byte uint64_t. Valid range for n with
|
||||
// the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1].
|
||||
// When n > INT64_MAX, the result exceeds int64_t range and is rejected.
|
||||
// When n > INT64_MAX, the result exceeds int64_t range and is stored
|
||||
// as a floating-point number, as the lexer does for JSON text.
|
||||
|
||||
SECTION("n = 0 is valid (result = -1)")
|
||||
{
|
||||
@@ -3206,33 +3207,34 @@ TEST_CASE("Tagged values")
|
||||
CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)());
|
||||
}
|
||||
|
||||
SECTION("n = INT64_MAX + 1 is rejected (overflow)")
|
||||
SECTION("n = INT64_MAX + 1 is stored as float")
|
||||
{
|
||||
// n = INT64_MAX + 1 (0x8000000000000000)
|
||||
// result = -1 - n = -9223372036854775809, which exceeds int64_t range
|
||||
// result = -1 - n = -9223372036854775809, which exceeds int64_t range;
|
||||
// the nearest double is -9223372036854775808.0
|
||||
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
|
||||
json _;
|
||||
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
|
||||
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
|
||||
json::parse_error);
|
||||
const auto result = json::from_cbor(input);
|
||||
CHECK(result.is_number_float());
|
||||
CHECK(result.get<double>() == -9223372036854775808.0);
|
||||
CHECK(result == json::parse("-9223372036854775809"));
|
||||
}
|
||||
|
||||
SECTION("n = UINT64_MAX is rejected (overflow)")
|
||||
SECTION("n = UINT64_MAX is stored as float")
|
||||
{
|
||||
// n = UINT64_MAX (0xFFFFFFFFFFFFFFFF)
|
||||
// result = -1 - n = -18446744073709551616, which exceeds int64_t range
|
||||
const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
|
||||
json _;
|
||||
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
|
||||
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
|
||||
json::parse_error);
|
||||
const auto result = json::from_cbor(input);
|
||||
CHECK(result.is_number_float());
|
||||
CHECK(result.get<double>() == -18446744073709551616.0);
|
||||
CHECK(result == json::parse("-18446744073709551616"));
|
||||
}
|
||||
|
||||
SECTION("overflow with allow_exceptions=false returns discarded")
|
||||
SECTION("overflow with allow_exceptions=false is not an error")
|
||||
{
|
||||
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
|
||||
const auto result = json::from_cbor(input, true, false);
|
||||
CHECK(result.is_discarded());
|
||||
CHECK(result.is_number_float());
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user