* Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests
The strongest correctness checks for the UBJSON and BJData writers lived
only in the OSS-Fuzz drivers: anything from_ubjson()/from_bjdata()
returns must serialize with every option combination, parse back, and
re-serialize stably. Those checks only run at OSS-Fuzz, so regressions
surfaced days later as external reports - the same BJData assert pair
was reported five times over three years, and #5494's harness change
was followed by OSS-Fuzz 563659413 within a day.
Add "UBJSON round-trip invariants" and "BJData round-trip invariants"
test cases that run the drivers' checks on a fixed, deterministic corpus
(tests/src/round_trip_corpus.hpp): integer and float boundaries,
non-finite numbers, strings, binary values, optimized containers, deep
nesting, the JData annotated-array matrix, and seeded random containers.
They also check two properties the drivers do not: the first round trip
preserves the value, and re-serializing reproduces the exact bytes. For
BJData both exclude values containing a binary value, which is read back
as an array of integers unless it was written as a Draft 3 optimized
binary array; this carve-out is now documented in bjdata.md. Run against
the headers before #5542, the BJData test fails, including on the shape
from OSS-Fuzz 563659413.
Also document how OSS-Fuzz reports are handled (reference them as
"OSS-Fuzz: <id>", turn the reproducer into a unit test, keep drivers and
unit tests in sync) in tests/fuzzing.md, and link it from the PR
template and the quality assurance page.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add the OSS-Fuzz reproducers for 474400817 and 474480402 as unit tests
Following the convention added to tests/fuzzing.md, the reproducers of
the two BJData fuzzer asserts tracked since January are now unit tests:
- 474400817 (assert(false)): an empty object _ArraySize_ was written as
the ND-array header length, which from_bjdata() could not read back.
Fixed by #5455.
- 474480402 (to_bjdata(j2, false, false) == vec2): a one-byte Draft 3
binary array is written in Draft 2 mode as a uint8 array and then
re-serialized with the int8 marker. This is the documented exception to
byte stability, not a library bug; OSS-Fuzz closed it after #5494
relaxed the harness to value stability. The test pins the exact bytes
so the exception stays deliberate.
The 563659413 reproducer is already a unit test (#5542). A comment also
ties the existing UBJSON excessive-count test to the timeout OSS-Fuzz
reported for that shape (testcase 6347769435193344).
OSS-Fuzz: 474400817
OSS-Fuzz: 474480402
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix GCC -Weffc++ and -Wuseless-cast warnings in the round-trip corpus
Initialize the atoms in the member initialization list, and drop the cast of
the generator's result, which already is std::size_t on 64-bit Linux.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_bjdata_ndarray() encoded a JData-annotated object as a BJData
ND-array whenever its dimensions' product matched _ArrayData_.size(),
which lost information in two ways:
- _ArrayData_ was never required to be an array. null has size 0, any
other scalar has size 1, and iterating an object visits its values, so
e.g. {"_ArraySize_":[1],"_ArrayData_":5} was written as the array [5],
and an object _ArrayData_ came back as an array.
- The reader only restores an annotated object from an ND-array with at
least two non-zero dimensions that is not a 1xN row vector; an empty,
1-D, row-vector, or zero-sized shape is read back as a plain array. The
writer nonetheless emitted ND-array headers for these shapes, so the
annotation was silently dropped.
OSS-Fuzz issue 563659413 hit this in parse_bjdata_fuzzer: an empty binary
_ArraySize_ is written as a plain object and read back as an empty array,
after which {"_ArrayType_":"int16","_ArraySize_":[],"_ArrayData_":null}
was encoded as the ND-array header "[$I#[]" and re-read as [], failing the
harness's value-stability check.
Such objects now fall back to a plain object encoding, which round-trips.
Genuine ND-arrays (two or more positive dimensions, not a 1xN row vector)
are encoded exactly as before. Existing fallback tests that used 1-D
shapes are moved to 2-D shapes so they keep exercising the check they
were written for, and the BJData documentation is updated.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fall back to plain-object encoding when _ArrayType_ is not a string
write_bjdata_ndarray() looked up _ArrayType_ by calling get<string_t>()
directly, which throws type_error.302 when the annotation is not a
string (e.g. a number, null, boolean, array, or object). Per the
documented BJData ndarray contract, an object only qualifies for the
compact ndarray encoding if _ArrayType_ names a known type; anything
else must fall back to plain-object encoding, the same way an unknown
type-name string already does.
Add an is_string() check before the get<string_t>() call so a
non-string _ArrayType_ takes the existing "unrecognized type name"
fallback path instead of throwing.
Fixes#5398.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Relax the BJData fuzzer's round-trip check from byte-exact to value-exact
Fixing #5398 lets to_bjdata() proceed past the object it used to reject,
which exposed a pre-existing, unrelated round-trip quirk to the fuzzer:
a binary_t value serialized through the non-optimized ("$U#"-less)
array encoding is parsed back as a plain array of numbers, since
from_bjdata() has no way to tell "array of uint8 numbers" apart from
"array of bytes" without that optimized header. Re-serializing that
plain array then goes through the generic smallest-type writer, which
- unrelated to this PR, and long predating it - prefers the 'i' (int8)
marker over 'U' (uint8) for values that fit both, so the re-encoded
bytes can differ from the original even though both decode to the same
value.
This is not introduced by the #5398 fix; the same divergence reproduces
from a bare json::binary_t value with no _ArrayType_ annotation
involved at all, on the commit immediately preceding it. A general fix
would mean changing the shared UBJSON/BJData smallest-type selection
that hundreds of existing tests pin to 'i' for small positive
integers, which is out of scope and too risky for this PR.
Update fuzzer-parse_bjdata.cpp's round-trip assertions to check that
re-serializing is value-stable (from_bjdata(to_bjdata(j)) == j) rather
than byte-exact, matching the guarantee BJData actually provides, and
add a regression test in unit-bjdata.cpp using the exact OSS-Fuzz input
that documents the behavior.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare dump()s instead of json values in the BJData fuzzer's round-trip check
The value-stability assertion added to fix the earlier OSS-Fuzz crash
(json::from_bjdata(to_bjdata(j2)) == j2) itself broke on a NaN payload:
IEEE 754 NaN is never equal to itself, so operator== reports two
structurally-identical trees containing a non-finite double as
different -- not a round-trip bug, just NaN's ordinary
non-reflexivity. dump() serializes any non-finite double the same
deterministic way (as JSON null, since JSON cannot represent NaN or
Infinity), so comparing dumps is stable under exactly the values that
break operator==.
Verified against both the original OSS-Fuzz crash input and the new
one (0x68 0x68 0x7c, which decodes to a NaN), plus a local 2.5M-case
random-input sweep with no failures.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix to_bjdata() emitting the Draft-3-only 'B' marker in default Draft-2 mode
_ArrayType_ = "byte" mapped unconditionally to the BJData type marker
'B', regardless of the requested bjdata_version. 'B' is defined only by
BJData Draft 3; with the default version (draft2), this produced a
stream that is invalid for Draft 2 and, unlike every other
_ArrayType_, round-tripped back as a binary value instead of the
original annotated object.
Only accept "byte" / emit 'B' when bjdata_version selects Draft 3.
Under Draft 2, fall back to the same plain-object encoding used
elsewhere in this function for other invalid-annotation cases, so the
value round-trips correctly.
Fixes#5404.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Future-proof the Draft-3-only 'B' marker gate
@gregmarr pointed out that dtype == 'B' && bjdata_version != draft3
only future-proofs by accident, since bjdata_version_t currently has
exactly two values. Compare with < instead, so a later draft that
keeps the 'B' marker valid does not need this gate revisited.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Reserve capped array capacity for definite-length binary arrays
CBOR, MessagePack, and the optimized [$type#count UBJSON/BJData form all
pass an exact element count to sax->start_array(len), but
json_sax_dom_parser::start_array() (and the callback variant) only used
len for an overflow check against max_size() and never reserved the
underlying vector, so each element triggered a reallocation cascade via
emplace_back().
Reserve upfront, but cap the reservation at 16384 elements: max_size()
for a std::vector is far larger than any realistic input, so an
unbounded reserve(len) would let a crafted/truncated header (e.g. CBOR
0x9A + a huge uint32 count with no data) trigger a multi-gigabyte
allocation attempt instead of the normal graceful parse_error. With the
cap, a hostile length still fails fast with the existing parse_error,
while realistic arrays get a single up-front allocation.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make the huge-claimed-length DoS regression tests portable across size_t widths
On a platform where size_t is narrower than 64 bits (e.g. 32-bit mingw/msvc
x86), the previously-hardcoded huge test lengths either collide with that
platform's unknown_size() sentinel (CBOR/MessagePack, both using exactly
SIZE_MAX) or exceed the platform's smaller vector<json>::max_size()
(UBJSON/BJData's 0x7FFFFFFF), so the header is now rejected outright
(out_of_range.408) instead of being accepted and only found short of data
(parse_error.110). Both are safe, bounded rejections of the hostile input;
the property under test -- no attempt to allocate space for billions of
elements -- holds either way. Accept both outcomes instead of pinning the
64-bit-only exact result.
Also fixed an unrelated clang-tidy finding (google-readability-casting) on
the functional-style std::size_t(...) casts in the neighboring "arrays of
various sizes" section.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix remaining CI failures in the huge-claimed-length DoS regression tests
- Apply the same google-readability-casting fix (std::size_t{N} instead
of std::size_t(N)) to the "arrays of various sizes" section in
unit-msgpack.cpp, unit-ubjson.cpp, and unit-bjdata.cpp; only
unit-cbor.cpp had been fixed previously, since clang-tidy's build
didn't get far enough to report the other three in the same pass.
- json_sax_dom_parser::start_array()'s max_size() check calls JSON_THROW
directly rather than going through sax->parse_error(), so unlike the
scanner's own "not enough data" parse_error it is not gated by
allow_exceptions=false. On a platform where a header's claimed count
exceeds max_size() (e.g. 32-bit, for UBJSON/BJData's 0x7FFFFFFF test
value), from_ubjson/from_bjdata(input, true, false) can therefore still
throw instead of returning a discarded value. Make that assertion
tolerant of either outcome, same as the main exception-catching check
above it.
- Guard all four "a huge claimed length..." SECTIONs with
#if !defined(JSON_NOEXCEPTION), matching this test suite's existing
convention for exception-dependent tests: under JSON_NOEXCEPTION,
JSON_THROW never produces a catchable C++ exception at all (it aborts
the process), so a section that relies on try/catch to distinguish
between two acceptable outcomes cannot be expressed under that build
configuration regardless of platform.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use (std::min)(len, reserve_cap) instead of a ternary in start_array()
Addresses review feedback from @gregmarr on PR #5476.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_ubjson_size_type() takes an inside_ndarray parameter saying whether it is
being called for an ndarray's dimension vector, where another ndarray is not
allowed. It then seeded the flag it passes down to get_ubjson_size_value()
with `false` rather than with that parameter, and only consulted
inside_ndarray afterwards, on the '$' branch.
So on the '#' branch nothing stopped the descent: every "#[" pair of an input
like "[" followed by "#[#[#[..." opened another dimension vector, several
native stack frames deeper each time, and the recursion was only reported on
the way back out. 100,000 pairs crash the process. This is #5104 again, in a
path that has nothing to do with containers.
Seed the flag with inside_ndarray, which is what get_ubjson_size_value()
documents it wants: "for input, `true` means already inside an ndarray vector
or ndarray dimension is not allowed". The nested '[' is then refused where it
is read, so the length of the chain no longer matters.
Both post-checks gain `&& !inside_ndarray`, because an ndarray was found
*here* only if the flag flipped -- get_ubjson_size_value() only ever returns
`true` when its initial value was `false`, as its documentation says. With
that, the "ndarray can not be recursive" branch is unreachable: a recursive
ndarray is now caught one level earlier, and reported as "ndarray dimensional
vector is not allowed" like every other nested dimension vector.
Three existing expectations move accordingly (vR2, vR4, vR6). All three now
fail earlier, and all three now report the same error that vR1, vR5 and vH
already reported for the same shape, which is the more consistent outcome.
Everything else is unchanged: valid 1D and 2D ndarrays, optimized containers
and plain arrays produce identical results, and unit-ubjson is untouched.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_bjdata_ndarray() validated that each _ArrayData_ element matched
the number kind (integer vs. float) named by _ArrayType_, but not its
range. An element that did not fit the target C++ type (e.g. 256 for
"uint8") was silently wrapped by the static_cast used to write it, or,
for "single", silently overflowed to infinity.
Range-check each element against the type named by _ArrayType_ before
writing it, reusing the existing fallback path that already encodes
the annotated object as a plain object for other invalid-annotation
cases in this function.
Fixes#5403.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix to_bjdata() emitting unparsable output when _ArraySize_ is not an array
write_bjdata_ndarray() never checked that _ArraySize_ is an array. The shape
is written verbatim as the header length, so a null shape emitted 'Z' and an
object shape emitted '{' after the '#', neither of which from_bjdata()
accepts, and the round-trip guarantee in the BJData docs was broken.
Both slipped through the existing validation: for null, empty() is true so
the element count starts at 0 and the per-dimension loop never runs, and for
an object the loop walks its values, which can satisfy the non-negative
integer check. When _ArrayData_ then matched that count, the writer took the
ndarray path.
Require the shape to be an array, so anything else falls back to a plain
object encoding that round-trips, as the fallback rule in the docs already
specifies.
Signed-off-by: qatcod <79017227+qatcod@users.noreply.github.com>
* Document that _ArraySize_ must be an array in the ndarray requirements
The list at bjdata.md is the exhaustive set of conditions for the ndarray
encoding, but it only implied this one through 'every entry of'.
Signed-off-by: qatcod <79017227+qatcod@users.noreply.github.com>
---------
Signed-off-by: qatcod <79017227+qatcod@users.noreply.github.com>
* Throw other_error.502 when UBJSON use_type is set without use_size
Fixes#5321
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Scope UBJSON use_type check to container branches and expand tests
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Re-amalgamate single_include/json.hpp
The previous commit updated the split headers but the amalgamated
file didn't go back through astyle before I committed it, so CI's
amalgamation check caught formatting drift in json_fwd.hpp and a
few noexcept clauses in basic_json, plus one doc example. None of
it touches the UBJSON logic. Applied the patch CI generated to
bring single_include back in sync.
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
---------
Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
* Do not write BJData ndarrays whose size overflows std::size_t
write_bjdata_ndarray() multiplied the _ArraySize_ dimensions into a
std::size_t without checking for overflow. A product that wraps around
to a value that happens to match the size of _ArrayData_ passed the
length check, and the writer emitted an ndarray header announcing an
element count that cannot be represented:
{"_ArrayType_":"uint8","_ArraySize_":[9223372036854775808,2],"_ArrayData_":[]}
was encoded as 5b 24 55 23 5b 4d 00 00 00 00 00 00 00 80 69 02 5d, an
ndarray of 2^64 elements followed by no data. Reading that back throws
out_of_range.408 ("excessive ndarray size caused overflow"), so to_bjdata
produced output that from_bjdata rejects. This is reachable by parsing
untrusted JSON and re-encoding it as BJData.
Mirror the overflow check the binary reader already performs, and also
reject a single dimension that does not fit into std::size_t, which the
previous cast silently truncated where std::size_t is narrower than 64
bits. Such objects now fall back to a plain object encoding, which is
what the surrounding type and length validation already does for
annotations it cannot represent, and they round-trip unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document when to_bjdata converts a JData annotation to an ND-array
The BJData page described the 1-D vector case as the only situation in
which an object carrying _ArrayType_/_ArraySize_/_ArrayData_ is not
written as a compact ND-array. The writer has always had several other
fallbacks -- an unknown _ArrayType_, a dimension that is not a
non-negative integer, an _ArrayData_ whose length does not match the
product of the dimensions, and elements that are not numbers of the
annotated kind -- all of which cause the value to be serialized as a
regular JSON object instead.
Spell out the conditions, including the size-overflow check added in the
preceding commit, so the documented behavior matches the implementation.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This adds a std::isfinite check to the UBJSON floating-point parsing path, throwing out_of_range.406 on overflow. This makes the UBJSON parser's behavior consistent with the normal JSON parser. Fixes#5322.
Signed-off-by: AJ369ninja <abhishek.j@iitg.ac.in>
Co-authored-by: AJ369ninja <abhishek.j@iitg.ac.in>
* validate ndarray element types in write_bjdata_ndarray
Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
* read ndarray elements through get<> instead of a fixed union member
_ArrayType_ names the wire type, not how the value is stored: parsing
keeps a non-negative integer as number_unsigned while the C++ API keeps
an int literal as number_integer. Selecting the union member from the
type marker therefore reads the inactive alternative for one of the two,
so read through get<> instead, which dispatches on the active member.
Also reject a negative _ArraySize_ entry, which is not a usable
dimension, and cover the parse-built path in the tests.
Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
---------
Signed-off-by: Angadi Yashaswini <angadi@digiscrypt.com>
* Add iterator+sentinel tests and docs for binary deserializers
This commit extends the C++20 ranges support (iterator+sentinel pairs) to the
binary format deserializers from_cbor, from_msgpack, from_ubjson, from_bjdata,
and from_bson, matching what was already done for parse(), accept(), and
sax_parse().
Changes:
- Add istreambuf_sentinel helper to test_utils.hpp for EOF detection in tests
- Add 5 new test cases that read binary files directly via
std::istreambuf_iterator<char> + sentinel, without pre-buffering
- Update documentation for all 5 from_* functions to document overload (3)
with SentinelType parameter
- All tests pass; verified against existing test suite data
- Fix potential buffer over-read warning in heterogeneous iterator test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge iterator+sentinel overloads and fix ambiguity/CI issues
Address PR review feedback and CI failures:
- Merge the separate same-type and sentinel-type iterator overloads of
parse(), accept(), sax_parse(), and the five from_* binary deserializers
into a single overload with SentinelType defaulted to IteratorType,
as suggested in review. Applied the same simplification to the
detail::input_adapter() free functions.
- Fix a latent ambiguity: some compilers (e.g. GCC 4.8) unreliably SFINAE
the operator!= detection for std::nullptr_t against container/string
types, making calls like parse(s, nullptr, ...) ambiguous with the
compatible-input overload. can_compare_ne now explicitly excludes
std::nullptr_t as a SentinelType.
- Use a named enable_if_t template parameter instead of an unnamed
function parameter for the SFINAE guard, fixing a clang-tidy
hicpp-named-parameter/readability-named-parameter failure.
- Update parse.md, accept.md, sax_parse.md, and the five from_*.md pages
to document the merged overload instead of separate (2)/(3) overloads,
also fixing an over-160-char line that broke the documentation
style_check CI job.
- Rework the BSON iterator+sentinel test to parse a BSON file already
present in the test suite instead of writing/deleting a temp file.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Wunneeded-internal-declaration for CustomSentinel in test
CustomSentinel lives in an anonymous namespace (internal linkage), and
the library's parse loop only ever evaluates the iterator-first
direction (it != last), so the reversed-order friend operator!= was
never referenced. Clang's -Weverything flags such unused internal
declarations as an error. Drop the unused overload; the used direction
is enough to satisfy can_compare_ne's either-order detection.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy hicpp-named-parameter and misc-const-correctness
- Drop the unused reversed-order operator!= overload from
utils::istreambuf_sentinel (only iterator != sentinel is ever
evaluated) and name the remaining friend's sentinel parameter, fixing
hicpp-named-parameter/readability-named-parameter.
- Mark the istreambuf_iterator first/last helper variable const in the
five binary-format sentinel tests, fixing misc-const-correctness.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy misc-const-correctness in heterogeneous sentinel test
json_str is only read via .data()/.size() and never reassigned, so
clang-tidy correctly flags it as const-able. Verified against the exact
CI job (silkeh/clang:dev, ci_clang_tidy target) by running clang-tidy
directly on this file plus the five binary-format sentinel tests
touched by prior commits; all are now clean.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Adds pre-multiplication overflow detection to catch cases where dimension
products would exceed size_t max. The previous check only detected when
overflow resulted in exactly 0 or SIZE_MAX, missing other cases.
Retains the original post-multiplication check for backward compatibility.
Adds tests verifying overflow detection with dimensions (2^32+1)×(2^32),
which previously overflowed silently to 2^32.
This prevents custom SAX handlers from receiving incorrect array sizes
that could lead to buffer overflows.
Signed-off-by: Ville Vesilehto <ville@vesilehto.fi>
* BJData dimension length can not be string_t::npos, fix#3541
* handle error messages on 32bit machine
* add explanation to why size can not be string_t::npos
* add test cases to 32bit unit test
Co-authored-by: Florian Albrechtskirchinger <falbrechtskirchinger@gmail.com>
* Fix ndarray dimension signness, fix ndarray length overflow, close#3519
* detect size overflow in ubjson and bjdata
* force reformatting
* Fix MSVC compiler warning
* Add value_in_range_of trait
* Use value_in_range_of trait
* Correct 408 parse_errors to out_of_range
* Add 32bit unit test
The test can be enabled by setting JSON_32bitTest=ON.
* Exclude unreachable lines from coverage
Certain lines are unreachable in 64bit builds.
Co-authored-by: Qianqian Fang <fangqq@gmail.com>
* Discard optimized containers with negative counts in UBJSON/BJData (#3491,#3492,#3490)
* fix msvc error
* update unit tests for negative sized containers
* use a loop to test 0 ndarray dimension
* throw an error when count is negative, merge CHECK_THROW_AS and _WITH with _WITH_AS
* change bjdata ndarray flag to detect negative size, fix https://github.com/nlohmann/json/issues/3475
* fix CI error
* fix CI on 32bit windows
* remove platform specific out_of_range error messages
* Incorporate suggestions from @nlohmann and @falbrechtskirchinger
* fix CI errors
* add coverage
* fix sax event order
* fix coverage