* Write BSON in linear time, without recursing per nesting level
to_bson() had two problems with nested values:
- It recursed once per nesting level, so a value nested deeply enough -
100,000 levels on an 8 MiB stack - exhausted the call stack and
terminated the process, although parse() accepts such values without
complaint.
- BSON prefixes every document and array with its length. The writer
computed that length by walking the entire value below it, again for
every nested document it wrote, which made serializing O(size x depth).
A 200-level document took 30 ms instead of 1.
Both passes are now iterative, and each length is computed exactly once:
- calc_bson_sizes() computes the length of every document and array in
one pass, each from the lengths of its entries, into a table ordered
the way they are written.
- write_bson_document() then writes the document, taking each length from
the table.
Everything observable is unchanged, as a differential test against
develop confirms byte for byte:
- The same bytes are written.
- A key containing U+0000 still throws out_of_range.409 for the same
first key, with the same diagnostics path, before anything is written.
- A document too large for BSON still throws out_of_range.412 before
anything is written.
- A binary subtype above 255 still throws out_of_range.415 after the
same partial output.
Only the enclosing objects and arrays are kept on a stack, so a flat
document allocates nothing for it. Measured against develop (clang -O3,
median of 201 runs): flat objects unchanged, flat arrays 37% faster (the
array length was computed twice), a nested 3,000-object document 2x
faster, a 200-level document 33x faster.
to_bson.md documented the quadratic complexity since #5334; it is linear
again.
Fixes#5392 for BSON, and #5308.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Do not require a default-constructible string_t in the BSON writer
GCC 4.9 and MSVC rejected the test's huge_string_t, which has no default
constructor; develop never default-constructed string_t here either.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Let the BSON index-name helper only fill its output parameter
It returned a reference to the string it filled, so callers held a second
name for index_name. Addresses review feedback.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)
from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).
Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.
There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* ⚡ validate only newly read bytes of binary-format strings
get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Reject MessagePack/BSON binary subtypes that don't fit their wire format
Both formats store byte_container_with_subtype's subtype (a uint64_t)
in a single byte. The writers cast to std::int8_t/std::uint8_t without
a range check, so subtypes above 255 were silently truncated modulo
256 instead of raising an error. Throw out_of_range.413 instead when
the subtype exceeds the representable range of 0-255.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the new binary-subtype regression test out of unit-regression2.cpp
unit-regression2.cpp is already at the edge of what the MinGW linker
can relocate; adding this test's ~26 lines tips test-regression2_cpp20
(clang, Windows) over into "relocation truncated to fit:
IMAGE_REL_AMD64_REL32 against `.rdata'" (see 8ce64b9c1 / b82717c8a for
the same failure mode). Split the test along format lines instead:
MessagePack assertions move to unit-msgpack.cpp, BSON assertions to
unit-bson.cpp. The CBOR round-trip guard is dropped as redundant --
unit-cbor.cpp's "Tagged values" section already round-trips subtypes
up to 8589934590, far past the 70000 checked here.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BSON documents without recursing per nesting level
An embedded document (record type 0x03) or array (0x04) was read by calling
back into the document reader, which read its element list, which called the
element reader again for the next embedded one. The native call stack
therefore grew with the nesting depth of the input, and about seven bytes buy
a level, so a document of a few hundred kilobytes crashes the process
(#5104). This is the last of the four binary formats to still do that.
Apply the same shape as the other three: open_bson_document() reads the size
prefix and opens the document, parse_bson_element_internal() calls it for both
record types instead of recursing, and parse_bson_internal() loops over the
element list of whichever document is innermost, closing it when its
terminator is reached and resuming the one below.
check_bson_document_size() is unchanged, and so is when it runs: a document is
still measured from the byte before its size prefix to the byte after its
terminator, and still reported before the end event. The frame carries those
two values, which is what a per-document check needs once the reads are
interleaved rather than nested. Nothing else about the element reader changes.
unit-bson passes unchanged. Round trips through to_bson of nested objects,
arrays, arrays of objects and mixed nesting are identical to the previous
commit, as are the errors for a truncated document, an unsupported record
type, a negative size and a size that does not match, including their byte
offsets. A 30,000-level document built by to_bson is now read to completion
where it used to crash.
Note for sequencing: #5185 changes parse_bson_internal(), the element list and
the array reader, which are the functions this commit restructures. It should
land first; this commit then keeps its checks and moves them onto the loop.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make parse_bson_internal's end-of-document top a copy, not a reference
Same issue as the CBOR and UBJSON/BJData readers: top aliased
container_stack.back() and was read (top.is_object) right after
container_stack.pop_back() ended its lifetime. A copy stays valid
regardless of what happens to the stack; nothing here mutates the live
entry, so no field needs to go through container_stack.back() directly.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add test coverage for documented lenient BSON input handling
Issue #5333 documented three intentionally-lenient behaviors of the BSON
reader (any non-zero byte accepted as a boolean `true`, BSON array element
keys not validated against the required decimal sequence, and the payload
of binary subtype 0x02 "old binary" returned as-is including its inner
length prefix), but none of them was pinned by a test, so a future change
could silently regress the documented behavior.
Also add coverage for the out_of_range.412 length-overflow check
(shared by binary, string, and (sub-)document BSON length fields) for
the string and document cases; only the binary case was previously
tested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix 32-bit overflow in huge_string_t BSON length-overflow tests
huge_string_t doubles as basic_json's StringType, so it is used not only
for the JSON string value under test but also for object keys (e.g. "s",
"nested"). Making size() unconditionally lie about being huge therefore
inflated the keys' reported sizes as well, pushing the running totals
computed while walking the BSON document (calc_bson_object_size and
friends in binary_writer.hpp) past what a 32-bit std::size_t can hold.
On 64-bit platforms this happens to still produce a working (if
needlessly large) result, but on 32-bit platforms (e.g. the mingw x86 CI
job) the size_t arithmetic silently wraps around: for the "document" test
this merely surfaces the wrong number in the exception message, but for
the "string" test the wrapped total happens to fall back under
INT32_MAX, so the intended out_of_range.412 guard is skipped entirely and
the code goes on to actually write ~2 GiB worth of characters from the
key's real, tiny buffer - which is what raised the reported
"vector::_M_range_insert" exception instead of a controlled 412.
Make the fake-huge size opt-in via huge_string_t::as_huge() and only
apply it to the string value under test, leaving keys at their real
(small) size. This keeps every intermediate size well within 32-bit
size_t range on any platform, matching how huge_binary_t already avoids
the same trap (it is only ever used as the BSON value type, never as a
key). Expected out_of_range.412 messages are updated accordingly.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- warn about BSON marker 0x11 interoperability in both directions
- explain subtype-less binary normalization to subtype 0x00
- add a round-trip test for binary values without a subtype
Signed-off-by: YingqiDuan <141370165+YingqiDuan@users.noreply.github.com>
* Add iterator+sentinel tests and docs for binary deserializers
This commit extends the C++20 ranges support (iterator+sentinel pairs) to the
binary format deserializers from_cbor, from_msgpack, from_ubjson, from_bjdata,
and from_bson, matching what was already done for parse(), accept(), and
sax_parse().
Changes:
- Add istreambuf_sentinel helper to test_utils.hpp for EOF detection in tests
- Add 5 new test cases that read binary files directly via
std::istreambuf_iterator<char> + sentinel, without pre-buffering
- Update documentation for all 5 from_* functions to document overload (3)
with SentinelType parameter
- All tests pass; verified against existing test suite data
- Fix potential buffer over-read warning in heterogeneous iterator test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge iterator+sentinel overloads and fix ambiguity/CI issues
Address PR review feedback and CI failures:
- Merge the separate same-type and sentinel-type iterator overloads of
parse(), accept(), sax_parse(), and the five from_* binary deserializers
into a single overload with SentinelType defaulted to IteratorType,
as suggested in review. Applied the same simplification to the
detail::input_adapter() free functions.
- Fix a latent ambiguity: some compilers (e.g. GCC 4.8) unreliably SFINAE
the operator!= detection for std::nullptr_t against container/string
types, making calls like parse(s, nullptr, ...) ambiguous with the
compatible-input overload. can_compare_ne now explicitly excludes
std::nullptr_t as a SentinelType.
- Use a named enable_if_t template parameter instead of an unnamed
function parameter for the SFINAE guard, fixing a clang-tidy
hicpp-named-parameter/readability-named-parameter failure.
- Update parse.md, accept.md, sax_parse.md, and the five from_*.md pages
to document the merged overload instead of separate (2)/(3) overloads,
also fixing an over-160-char line that broke the documentation
style_check CI job.
- Rework the BSON iterator+sentinel test to parse a BSON file already
present in the test suite instead of writing/deleting a temp file.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Wunneeded-internal-declaration for CustomSentinel in test
CustomSentinel lives in an anonymous namespace (internal linkage), and
the library's parse loop only ever evaluates the iterator-first
direction (it != last), so the reversed-order friend operator!= was
never referenced. Clang's -Weverything flags such unused internal
declarations as an error. Drop the unused overload; the used direction
is enough to satisfy can_compare_ne's either-order detection.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy hicpp-named-parameter and misc-const-correctness
- Drop the unused reversed-order operator!= overload from
utils::istreambuf_sentinel (only iterator != sentinel is ever
evaluated) and name the remaining friend's sentinel parameter, fixing
hicpp-named-parameter/readability-named-parameter.
- Mark the istreambuf_iterator first/last helper variable const in the
five binary-format sentinel tests, fixing misc-const-correctness.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy misc-const-correctness in heterogeneous sentinel test
json_str is only read via .data()/.size() and never reassigned, so
clang-tidy correctly flags it as const-able. Verified against the exact
CI job (silkeh/clang:dev, ci_clang_tidy target) by running clang-tidy
directly on this file plus the five binary-format sentinel tests
touched by prior commits; all are now clean.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>