* Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests
The strongest correctness checks for the UBJSON and BJData writers lived
only in the OSS-Fuzz drivers: anything from_ubjson()/from_bjdata()
returns must serialize with every option combination, parse back, and
re-serialize stably. Those checks only run at OSS-Fuzz, so regressions
surfaced days later as external reports - the same BJData assert pair
was reported five times over three years, and #5494's harness change
was followed by OSS-Fuzz 563659413 within a day.
Add "UBJSON round-trip invariants" and "BJData round-trip invariants"
test cases that run the drivers' checks on a fixed, deterministic corpus
(tests/src/round_trip_corpus.hpp): integer and float boundaries,
non-finite numbers, strings, binary values, optimized containers, deep
nesting, the JData annotated-array matrix, and seeded random containers.
They also check two properties the drivers do not: the first round trip
preserves the value, and re-serializing reproduces the exact bytes. For
BJData both exclude values containing a binary value, which is read back
as an array of integers unless it was written as a Draft 3 optimized
binary array; this carve-out is now documented in bjdata.md. Run against
the headers before #5542, the BJData test fails, including on the shape
from OSS-Fuzz 563659413.
Also document how OSS-Fuzz reports are handled (reference them as
"OSS-Fuzz: <id>", turn the reproducer into a unit test, keep drivers and
unit tests in sync) in tests/fuzzing.md, and link it from the PR
template and the quality assurance page.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add the OSS-Fuzz reproducers for 474400817 and 474480402 as unit tests
Following the convention added to tests/fuzzing.md, the reproducers of
the two BJData fuzzer asserts tracked since January are now unit tests:
- 474400817 (assert(false)): an empty object _ArraySize_ was written as
the ND-array header length, which from_bjdata() could not read back.
Fixed by #5455.
- 474480402 (to_bjdata(j2, false, false) == vec2): a one-byte Draft 3
binary array is written in Draft 2 mode as a uint8 array and then
re-serialized with the int8 marker. This is the documented exception to
byte stability, not a library bug; OSS-Fuzz closed it after #5494
relaxed the harness to value stability. The test pins the exact bytes
so the exception stays deliberate.
The 563659413 reproducer is already a unit test (#5542). A comment also
ties the existing UBJSON excessive-count test to the timeout OSS-Fuzz
reported for that shape (testcase 6347769435193344).
OSS-Fuzz: 474400817
OSS-Fuzz: 474480402
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix GCC -Weffc++ and -Wuseless-cast warnings in the round-trip corpus
Initialize the atoms in the member initialization list, and drop the cast of
the generator's result, which already is std::size_t on 64-bit Linux.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The fuzzer drivers check their round trips with assert(), which NDEBUG
compiles away. The OSS-Fuzz build keeps assertions on today, but nothing
pins that: a build change that adds NDEBUG would silently turn every
round-trip check into a mere "does not crash" check. Each driver now
stops the build with an #error instead, and includes <cassert> itself
rather than relying on json.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fall back to plain-object encoding when _ArrayType_ is not a string
write_bjdata_ndarray() looked up _ArrayType_ by calling get<string_t>()
directly, which throws type_error.302 when the annotation is not a
string (e.g. a number, null, boolean, array, or object). Per the
documented BJData ndarray contract, an object only qualifies for the
compact ndarray encoding if _ArrayType_ names a known type; anything
else must fall back to plain-object encoding, the same way an unknown
type-name string already does.
Add an is_string() check before the get<string_t>() call so a
non-string _ArrayType_ takes the existing "unrecognized type name"
fallback path instead of throwing.
Fixes#5398.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Relax the BJData fuzzer's round-trip check from byte-exact to value-exact
Fixing #5398 lets to_bjdata() proceed past the object it used to reject,
which exposed a pre-existing, unrelated round-trip quirk to the fuzzer:
a binary_t value serialized through the non-optimized ("$U#"-less)
array encoding is parsed back as a plain array of numbers, since
from_bjdata() has no way to tell "array of uint8 numbers" apart from
"array of bytes" without that optimized header. Re-serializing that
plain array then goes through the generic smallest-type writer, which
- unrelated to this PR, and long predating it - prefers the 'i' (int8)
marker over 'U' (uint8) for values that fit both, so the re-encoded
bytes can differ from the original even though both decode to the same
value.
This is not introduced by the #5398 fix; the same divergence reproduces
from a bare json::binary_t value with no _ArrayType_ annotation
involved at all, on the commit immediately preceding it. A general fix
would mean changing the shared UBJSON/BJData smallest-type selection
that hundreds of existing tests pin to 'i' for small positive
integers, which is out of scope and too risky for this PR.
Update fuzzer-parse_bjdata.cpp's round-trip assertions to check that
re-serializing is value-stable (from_bjdata(to_bjdata(j)) == j) rather
than byte-exact, matching the guarantee BJData actually provides, and
add a regression test in unit-bjdata.cpp using the exact OSS-Fuzz input
that documents the behavior.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare dump()s instead of json values in the BJData fuzzer's round-trip check
The value-stability assertion added to fix the earlier OSS-Fuzz crash
(json::from_bjdata(to_bjdata(j2)) == j2) itself broke on a NaN payload:
IEEE 754 NaN is never equal to itself, so operator== reports two
structurally-identical trees containing a non-finite double as
different -- not a round-trip bug, just NaN's ordinary
non-reflexivity. dump() serializes any non-finite double the same
deterministic way (as JSON null, since JSON cannot represent NaN or
Infinity), so comparing dumps is stable under exactly the values that
break operator==.
Verified against both the original OSS-Fuzz crash input and the new
one (0x68 0x68 0x7c, which decodes to a NaN), plus a local 2.5M-case
random-input sweep with no failures.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>