Repair complete items in binary formats when parse_error() returns true (#3989)

When the SAX parser asks to recover, the binary readers now repair an
item whose end is known and read on after it, as RFC 8949, Section 5.3
describes for CBOR:

- CBOR: tags are ignored, and simple values other than false, true, and
  null become null (RFC 8949, Section 6.1); a negative integer below the
  range of number_integer_t becomes the nearest floating-point number.
- Strings that are not valid UTF-8 get U+FFFD for each ill-formed
  sequence, as in JSON text; so does a UBJSON/BJData char above 0x7F.
- UBJSON/BJData high-precision numbers keep their longest valid
  beginning (via the lexer's recover_token()), or become infinity.
- Members whose key is not a string are skipped (CBOR, MessagePack,
  BON8), like members without a key in JSON text.
- BSON elements of types the library does not read (ObjectId, datetime,
  decimal128, ...) become null; a string without its terminator and a
  document whose size does not match are kept.

Where the end of an item is unknown, reading stops as before, except
that BSON skips to the end of the document, whose size it knows.

The value read before such an error is now completed by the reader from
its container stack, as the JSON parser does, instead of by a proxy SAX
parser, which is removed. Like the parser, binary_reader gets an
AllowRecovery template parameter, so that from_*() compile without the
new code.

Tests: a table of repairs, numbers out of range, errors that stop, and
a sweep over changed and removed bytes of eight encodings that checks
balanced events and that the first error is the one from_*() reports.
All fuzzers now run a recovering checker; the binary ones also check
that it reports an error exactly when from_*() fails.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-27 21:40:59 +02:00
parent c8735246d0
commit 437a95cfdb
16 changed files with 2925 additions and 673 deletions
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_cbor(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_cbor() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::cbor).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_cbor(vec1);
assert(recovered_without_errors);
try
{
@@ -60,6 +70,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -68,6 +79,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use