mirror of
https://github.com/nlohmann/json.git
synced 2026-10-02 06:15:46 +01:00
* Run the README test case in JSON_FastTests jobs The "README" test case was marked doctest::skip() when the tests moved from Catch to doctest in 2019, where it replaced Catch's hidden tag. It is not slow (17 assertions, about 0.00 s), but cmake/test.cmake only passes --no-skip when JSON_FastTests is off, so the per-compiler ci_test_*_cxxNN matrix, macOS, Windows Release/ARM, icpc, icpx and nvhpc compiled the README examples without running them. Drop the skip decorator so every job runs the case. Test-only change. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove stale clang ranges guards in unit-iterators2.cpp The "algorithms" and "views" sections were guarded by clang/libstdc++ checks written for a clang 15 (04/2022) bug. The first guard's condition contradicts its own comment: it skips clang+libc++ and keeps clang+libstdc++. Both sections already sit inside `#if JSON_HAS_RANGES`, which macro_scope.hpp excludes for the toolchains these guards targeted, so the inner guards never let the sections run on the platforms they meant to protect and are redundant on the rest. Verified locally with Apple clang 21/libc++ and clang 16.0.6/libstdc++ 12 (Docker): both pass all 1355 assertions with the guards removed. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix copy-pasted CBOR half-float checks; enable stale encode checks In the RFC 8949 Appendix A test case, the decode checks for 5.960464477539063e-8 (0xf9 0x00 0x01) and 0.00006103515625 (0xf9 0x04 0x00) were copy-pasted from the neighboring -4.0 example, so those two half-float byte sequences were never actually decoded and checked, and -4.0 was checked three times instead. The two float32 encode checks for 100000.0 and 3.4028234663852886e+38 were commented out before the writer supported emitting float32 and are now verified to match byte for byte, so they are enabled. The remaining commented-out half-precision to_cbor checks are collapsed into a single explanatory comment, since the writer never emits half-precision floats. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Assert on the result of STL container conversions in tests The "object-like STL containers" and "array-like STL containers" sections converted json values into std::map, unordered_map, multimap, unordered_multimap, list, forward_list, array, valarray, vector, deque, set and unordered_set and discarded the result, so these ~60 conversions only proved that the code compiles and does not throw; a conversion that dropped or reordered elements would still pass. Bind each result and compare it against the expected container. Also fix a copy-paste slip in the deque section (`j2.get<std::deque<double>>()` instead of j3, so j3's doubles were never converted to a deque), and remove the dead `// CHECK(m5["one"] == "eins")` comments that referred to a variable that did not exist by asserting the equivalent through the bound result. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate SaxCountdown and other test helpers across formats SaxCountdown was copied byte-for-byte into six binary-format test files (unit-cbor.cpp, unit-msgpack.cpp, unit-ubjson.cpp, unit-bjdata.cpp, unit-bon8.cpp, unit-bson.cpp), about 370 redundant lines. Move it into tests/src/sax_countdown.hpp (namespace utils, alongside test_utils.hpp and round_trip_corpus.hpp) and include it from all six. trait_test_arg and the "value_in_range_of trait" TEST_CASE_TEMPLATE_DEFINE were duplicated between unit-32bit.cpp and unit-bjdata.cpp; the trait is a detail/meta trait, not specific to either file. Move it into tests/src/value_in_range_of_test.hpp; unit-32bit.cpp keeps its own include, since JSON_32bitTest=ONLY builds only that file. Each file keeps its own TEST_CASE_TEMPLATE_INVOKE list. sax_no_exception and the "issue #2824" section were duplicated in unit-regression2.cpp and unit-disabled_exceptions.cpp. Drop the copy from unit-regression2.cpp; unit-disabled_exceptions.cpp already covers the no-exceptions case that #2824 was about, and ci_test_noexceptions reruns it. No behavior change. Verified by building and running unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bon8, unit-bson, unit-32bit, unit-regression2 and unit-disabled_exceptions against include/ (clang++ -std=c++11, ASan/UBSan where applicable); assertion counts are unchanged from before the refactor. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the unreferenced vendored libFuzzer tests/thirdparty/Fuzzer (155 files, ~776 KB of vendored Apache-2.0 LLVM code from the 2016 OSS-Fuzz import) is not referenced by any CMakeLists, Makefile or workflow: the fuzz drivers link against -fsanitize=fuzzer or the repo's own tests/src/fuzzer-driver_afl.cpp. Its vendored README only points at llvm.org's own libFuzzer docs. Being dead code, it also adds noise to the flawfinder code-scanning workflow, which scans the whole tree. Remove the directory and its .reuse/dep5 entry. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove unreferenced 2016 benchmark and fuzz reports tests/reports (1.6 MB) holds AFL status pages and plots from 2016-08-29 and 2016-10-02, and a nativejson-benchmark snapshot from 2016 with links to rawgit.com, which shut down in 2019. Nothing references this directory: no doc, README section, script or workflow points at it, and it describes a ten-years-old, pre-2.0 snapshot of the library. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Run the CBOR, MessagePack, BSON and BON8 round-trip invariants in CI tests/src/round_trip_corpus.hpp exists so that the byte-stability invariant the fuzzer drivers check also runs on a fixed corpus in CI, instead of only at OSS-Fuzz. So far only the UBJSON and BJData drivers had a matching unit test; the CBOR, MessagePack, BSON and BON8 drivers assert the same invariant (assert(to_X(j2) == vec)) but nothing ran it outside OSS-Fuzz. Add "<FORMAT> round-trip invariants" test cases to unit-cbor.cpp, unit-msgpack.cpp, unit-bson.cpp and unit-bon8.cpp, modeled on the UBJSON case: seed j1 from the corpus (skipping values that do not survive the format's own round trip, as the fuzzer drivers only ever see values from_X() actually produced), then require from_X(to_X(j1)) not to throw and check to_X(j2) == to_X(j1). BSON only serializes objects, so non-object corpus values are skipped. Update the comments in round_trip_corpus.hpp and tests/fuzzing.md to name all six formats. The stream-versus-contiguous check in the BON8 driver is left out, as #5601 reworks it. A local probe confirms no violations on the current corpus (CBOR 3849 checked, MessagePack 3909, BSON 2958, BON8 3841 - matching the counts already recorded for this probe in the issue). Closes #5714 item 1. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix fuzzer driver step lists to match the checks the code performs The header comment of six of the seven binary-format fuzzer drivers listed an invariant the code does not check: CBOR, MessagePack, BSON and BON8 said "assert(j1 == j2)", but the code checks byte stability, assert(to_X(j2) == vec). UBJSON and BJData still described the old "assert(j1 == j2/j3/j4)" byte-exact check from before PR #5494 replaced it with a use_size/use_type-aware round trip (UBJSON) and a value-stability check (BJData); BJData's added paragraph already explained the new check, but the step list above it did not. Also remove a dead branch in fuzzer-parse_bson.cpp: from_bson() is called with allow_exceptions = true, so it throws instead of returning a discarded value, and the "if (j1.is_discarded()) return 0;" guard could never trigger. Drop the unused <iostream> include from all seven drivers and <sstream> from all but fuzzer-parse_bon8.cpp, which is the only one that uses std::istringstream. Overlaps #5601, which edits all seven drivers in the same hunks. Closes #5714 item 4. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Silence the CMP0169 deprecation in cmake_fetch_content, fix stale guards tests/cmake_fetch_content/project calls the single-argument FetchContent_Populate(json) after FetchContent_Declare(), which CMake 3.30 deprecated as CMP0169. Since the project declares cmake_minimum_required(VERSION 3.11...3.14), the policy stays unset, so every configure with a current CMake prints the deprecation warning. The test is kept on purpose: it is the only coverage of the FetchContent_Populate + add_subdirectory pattern for CMake 3.11-3.13 users, which the docs still describe as supported. Explicitly set CMP0169 to OLD, with a comment explaining why. Also fix two stale version guards: - tests/cmake_fetch_content/CMakeLists.txt guarded the test with VERSION_GREATER "3.11.0", which is dead now that tests/CMakeLists.txt requires CMake 3.13. - tests/cmake_fetch_content2/CMakeLists.txt guarded with VERSION_GREATER "3.14.0", which skips exactly 3.14.0, the first version with FetchContent_MakeAvailable. Change it to VERSION_GREATER_EQUAL "3.14". Verified locally: `ctest -R cmake_fetch_content` passes with CMake 4.1, and the CMP0169 deprecation warning that appeared before this change is gone. Closes #5714 item 5. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make the CMake integration-test wrappers consistent The six tests/cmake_* integration-test wrappers had drifted: - Only cmake_import and cmake_import_minver forwarded -A "${CMAKE_GENERATOR_PLATFORM}" to the inner configure, and none forwarded -T "${CMAKE_GENERATOR_TOOLSET}". The Windows workflow configures the outer build with -A Win32 -T ClangCL, so without forwarding, the inner projects of cmake_add_subdirectory, cmake_fetch_content, cmake_fetch_content2 and cmake_target_include_directories built with the generator defaults instead of matching the outer build's platform and toolset. Forward both consistently from all six wrappers. - cmake_fetch_content and cmake_fetch_content2 passed -Dnlohmann_json_source to their inner projects, which never read it (CMake warns "manually-specified variables were not used"); the inner projects fetch their own copy of the library instead. Drop it. - tests/CMakeLists.txt set JSON_FORCED_GLOBAL_COMPILE_OPTIONS from the matching environment variable but never read the cache variable again; the lines right below it read $ENV{JSON_FORCED_GLOBAL_COMPILE_OPTIONS} directly, like the LINK_OPTIONS counterpart already does. Remove the dead set(). This changes which platform and toolset the Win32 and ClangCL CI jobs build the four newly-forwarding wrappers' inner projects with, which may surface new failures there; CI has to confirm those jobs. Verified locally with Ninja (empty -A ""/-T "" is accepted): all 12 cmake_* tests still pass, and the inner fetch_content configures no longer warn about the unused variable. Closes #5714 item 6. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Turn the #972 fifo_map regression test into a real test The #972 regression test in unit-regression1.cpp only built a my_json array from a string literal (the original crash) and had no CHECK, so the fifo_map object type it exists to demonstrate was never exercised. Meanwhile the docs recommend fifo_map for keeping object keys in insertion order (object_order.md, template_parameters.md), and nothing tested that recommendation. Extend the section: after the original array assignment, parse an object with my_json::parse() (not via the "..."_json UDL, which returns a plain nlohmann::json and would exercise the cross-basic_json conversion constructor instead of the parser's own key insertion - and, as tried locally, does not keep fifo order for this stateful comparator) and check that dump() keeps insertion order, and that it survives erase() and inserting a new key. Also narrow thirdparty/fifo_map off the include path of every other test-* target: it was a PUBLIC include directory of test_main, even though unit-regression1.cpp is its only user. Add a small fifo_map_include INTERFACE library with that include directory and attach it to test-regression1 only via json_test_set_test_options(). Verified locally (test-regression1_cpp11, default build and -fsanitize=address,undefined): the new checks pass; `git grep fifo_map tests` still only finds unit-regression1.cpp and the vendored header. Closes #5714 item 7. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Document the vendored doctest.h patch; fix stale doctest_compatibility.h comments tests/thirdparty/doctest/doctest.h is doctest 2.4.12, imported in #4771. Two weeks later, #4801 hand-edited translateActiveException() to declare "String res;" inside the translator loop instead of before it, so a translator that does not match does not leave a previous translator's result in "res" for the next iteration to see. Nothing recorded this, so re-vendoring doctest.h from upstream would silently drop the fix. Add a comment at the patched site naming the version, the PR and the reason, so a future re-vendor knows to re-apply it. Also fix two stale comments in doctest_compatibility.h: - The DOCTEST_THREAD_LOCAL comment referenced Xcode 6/7, which is no longer supported; reword it to explain why the define must stay regardless (it keeps doctest's own thread_local usage out of the way of the same Clang/MinGW crash that JSON_NO_THREAD_LOCAL works around in the library, see ci_test_no_thread_local). - The <iosfwd> include's comment justified it with tests that define "private" as "public"; no test under tests/src does that any more (removed by #2352). Reword the comment instead of dropping the include, since confirming it is safe to drop needs the full CI matrix including MSVC 2015+. Verified locally that tests/src/unit-readme.cpp still builds and passes 17/17 with these headers. Closes #5714 item 9. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop compiling unit-wstring.cpp out entirely on classic ICC tests/src/unit-wstring.cpp wrapped the whole file in #ifndef __INTEL_COMPILER, with the comment "ICPC errors out on multibyte character sequences in source files". The ci_icpc job (intel/oneapi-hpckit:2023.2.1) still exists, so that job ran none of the wstring/u16string/u32string input adapter tests, including the malformed-input checks #5704 (open) extends. Only 9 lines contained non-ASCII bytes: the three *_is_utf16()/ *_is_utf32() probe functions, and three std::wstring/u16string/ u32string literals plus their narrow-string dump() expectations. Rewrite all of them with \u/\U escapes in the wide/u16/u32 literals and \x escapes (split into separate string-literal tokens so a following byte is never read as part of the same hex escape, e.g. "\xE1\x83\x85" "a") in the narrow ones. Remove the #ifndef __INTEL_COMPILER/#endif guard along with it. The *_is_utf16()/*_is_utf32() probes compared a raw multibyte literal against an escape-based one to detect a compiler that misreads the source file's encoding; with no raw literals left to misread, the comparison is now tautological, so drop the probes and the "if" guards around each SECTION's body instead of leaving them in as dead checks. The same non-ASCII-in-source-and-in-a-narrow-comparison pattern existed once more in unit-deserialization.cpp's "Using _json with char8_t literals #4945" test: a raw emoji character in a u8R"(...)" literal, guarded by a check_utf8() that returned false for ICC (same reason) and for Windows without the active UTF-8 code page. Rewrite the literal with a \U escape and compare it against a \x-escaped expectation instead of a second raw literal, and drop check_utf8() and the now-unused <windows.h> include along with the guard. Verified locally (clang, -std=c++11 and -std=c++20, -fsanitize=address,undefined, and a plain build): test-wstring keeps 18/18 assertions and unit-deserialization keeps 466/466 (c++11) and 477/477 (c++20) assertions, matching this branch before the change exactly - no coverage was gained or lost, only the source-encoding dependency was removed. ci_icpc has to confirm classic ICC actually builds and passes test-wstring now; if it does not, that is a real finding, not a reason to restore the guard. Overlaps #5704 (open), which edits unit-wstring.cpp inside the previously-guarded region (an include near the top, checks in the invalid-string sections, and a new section at the end). Closes #5713 item 5. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
1897 lines
71 KiB
C++
1897 lines
71 KiB
C++
// __ _____ _____ _____
|
|
// __| | __| | | | JSON for Modern C++ (supporting code)
|
|
// | | |__ | | | | | | version 3.12.0
|
|
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
//
|
|
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
// SPDX-License-Identifier: MIT
|
|
|
|
#include "doctest_compatibility.h"
|
|
|
|
#include <nlohmann/json.hpp>
|
|
using nlohmann::json;
|
|
|
|
#include <cstdint>
|
|
#include <fstream>
|
|
#include <limits>
|
|
#include <sstream>
|
|
#include <vector>
|
|
#include "make_test_data_available.hpp"
|
|
#include "round_trip_corpus.hpp"
|
|
#include "test_utils.hpp"
|
|
#include "sax_countdown.hpp"
|
|
using utils::SaxCountdown;
|
|
|
|
namespace
|
|
{
|
|
// a binary container that reports a size beyond INT32_MAX without allocating
|
|
// that much memory, so the BSON length overflow can be tested cheaply
|
|
class huge_binary_t : public std::vector<std::uint8_t>
|
|
{
|
|
public:
|
|
using std::vector<std::uint8_t>::vector;
|
|
|
|
size_type size() const noexcept // NOLINT(readability-convert-member-functions-to-static)
|
|
{
|
|
// one byte more than the BSON length field can represent
|
|
return static_cast<size_type>((std::numeric_limits<std::int32_t>::max)()) + 1;
|
|
}
|
|
};
|
|
|
|
using huge_binary_json = nlohmann::basic_json <
|
|
std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t,
|
|
double, std::allocator, nlohmann::adl_serializer, huge_binary_t, void >;
|
|
|
|
// a string type that can be made to report a size beyond INT32_MAX without
|
|
// allocating that much memory, so BSON length overflow can be tested for
|
|
// strings and (embedded) documents as well, following the same idea as
|
|
// huge_binary_t.
|
|
//
|
|
// Unlike huge_binary_t (which is only ever used as the BSON *value* type),
|
|
// this type doubles as basic_json's StringType and is therefore also used
|
|
// for *object keys* (e.g. "s" or "nested" below). Only the designated test
|
|
// value is meant to lie about its size - if every huge_string_t (including
|
|
// keys) reported a huge size, the running totals computed while walking the
|
|
// BSON document (see calc_bson_sizes in binary_writer.hpp)
|
|
// would need more than 32 bits, and on platforms where std::size_t is only
|
|
// 32 bits wide that arithmetic would silently wrap around, producing wrong
|
|
// (or even unguarded) lengths. The fake size is therefore opt-in via
|
|
// as_huge(), and plain strings - in particular object keys - keep reporting
|
|
// their real, small size.
|
|
class huge_string_t : public std::string
|
|
{
|
|
public:
|
|
using std::string::string;
|
|
huge_string_t(const std::string& s) : std::string(s) {} // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
|
|
|
|
// returns a copy of @a s whose size() pretends to be huge
|
|
static huge_string_t as_huge(const std::string& s)
|
|
{
|
|
huge_string_t result(s);
|
|
result.pretend_huge = true;
|
|
return result;
|
|
}
|
|
|
|
size_type size() const noexcept
|
|
{
|
|
if (pretend_huge)
|
|
{
|
|
// one byte more than the BSON length field can represent
|
|
return static_cast<size_type>((std::numeric_limits<std::int32_t>::max)()) + 1;
|
|
}
|
|
return std::string::size();
|
|
}
|
|
|
|
private:
|
|
bool pretend_huge = false;
|
|
};
|
|
|
|
using huge_string_json = nlohmann::basic_json <
|
|
std::map, std::vector, huge_string_t, bool, std::int64_t, std::uint64_t,
|
|
double, std::allocator, nlohmann::adl_serializer, std::vector<std::uint8_t>, void >;
|
|
} // namespace
|
|
|
|
TEST_CASE("BSON")
|
|
{
|
|
SECTION("individual values not supported")
|
|
{
|
|
SECTION("null")
|
|
{
|
|
json const j = nullptr;
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is null", json::type_error&);
|
|
}
|
|
|
|
SECTION("boolean")
|
|
{
|
|
SECTION("true")
|
|
{
|
|
json const j = true;
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is boolean", json::type_error&);
|
|
}
|
|
|
|
SECTION("false")
|
|
{
|
|
json const j = false;
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is boolean", json::type_error&);
|
|
}
|
|
}
|
|
|
|
SECTION("number")
|
|
{
|
|
json const j = 42;
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is number", json::type_error&);
|
|
}
|
|
|
|
SECTION("float")
|
|
{
|
|
json const j = 4.2;
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is number", json::type_error&);
|
|
}
|
|
|
|
SECTION("string")
|
|
{
|
|
json const j = "not supported";
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is string", json::type_error&);
|
|
}
|
|
|
|
SECTION("array")
|
|
{
|
|
json const j = std::vector<int> {1, 2, 3, 4, 5, 6, 7};
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is array", json::type_error&);
|
|
}
|
|
}
|
|
|
|
SECTION("keys containing code-point U+0000 cannot be serialized to BSON")
|
|
{
|
|
json const j =
|
|
{
|
|
{ std::string("en\0try", 6), true }
|
|
};
|
|
#if JSON_DIAGNOSTICS
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.out_of_range.409] (/en) BSON key cannot contain code point U+0000 (at byte 2)", json::out_of_range&);
|
|
#else
|
|
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.out_of_range.409] BSON key cannot contain code point U+0000 (at byte 2)", json::out_of_range&);
|
|
#endif
|
|
}
|
|
|
|
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
|
|
{
|
|
// out_of_range.412 is thrown from a single shared helper
|
|
// (to_bson_length) that guards the BSON length fields of binary
|
|
// values, strings, and (embedded) documents alike
|
|
SECTION("binary")
|
|
{
|
|
huge_binary_json j;
|
|
j["b"] = huge_binary_json::binary(huge_binary_t{});
|
|
|
|
CHECK_THROWS_WITH_AS(huge_binary_json::to_bson(j), "[json.exception.out_of_range.412] BSON length 2147483661 exceeds maximum of 2147483647", huge_binary_json::out_of_range&);
|
|
}
|
|
|
|
SECTION("string")
|
|
{
|
|
huge_string_json j;
|
|
j["s"] = huge_string_t::as_huge("value");
|
|
|
|
CHECK_THROWS_WITH_AS(huge_string_json::to_bson(j), "[json.exception.out_of_range.412] BSON length 2147483661 exceeds maximum of 2147483647", huge_string_json::out_of_range&);
|
|
}
|
|
|
|
SECTION("document")
|
|
{
|
|
// an oversized string nested one level deep makes the
|
|
// *embedded* document's own length exceed INT32_MAX as well
|
|
huge_string_json nested;
|
|
nested["s"] = huge_string_t::as_huge("value");
|
|
huge_string_json j;
|
|
j["nested"] = nested;
|
|
|
|
CHECK_THROWS_WITH_AS(huge_string_json::to_bson(j), "[json.exception.out_of_range.412] BSON length 2147483674 exceeds maximum of 2147483647", huge_string_json::out_of_range&);
|
|
}
|
|
}
|
|
|
|
SECTION("string length must be at least 1")
|
|
{
|
|
// from https://bugs.chromium.org/p/oss-fuzz/issues/detail?id=11175
|
|
std::vector<std::uint8_t> const v =
|
|
{
|
|
0x20, 0x20, 0x20, 0x20,
|
|
0x02,
|
|
0x00,
|
|
0x00, 0x00, 0x00, 0x80
|
|
};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 10: syntax error while parsing BSON string: string length must be at least 1, is -2147483648", json::parse_error&);
|
|
}
|
|
|
|
SECTION("objects")
|
|
{
|
|
SECTION("empty object")
|
|
{
|
|
json const j = json::object();
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x05, 0x00, 0x00, 0x00, // size (little endian)
|
|
// no entries
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with bool")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", true }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x0D, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x08, // entry: boolean
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x01, // value = true
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with bool")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", false }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x0D, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x08, // entry: boolean
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x00, // value = false
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with bool from a non-0/1 byte (lenient parsing)")
|
|
{
|
|
// documented lenient behavior (see gh-5333): any non-zero byte
|
|
// is accepted as `true`, not just 0x01
|
|
std::vector<std::uint8_t> const input =
|
|
{
|
|
0x0D, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x08, // entry: boolean
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x02, // value = 0x02 (neither 0x00 nor 0x01)
|
|
0x00 // end marker
|
|
};
|
|
|
|
const json expected = { { "entry", true } };
|
|
CHECK(json::from_bson(input) == expected);
|
|
}
|
|
|
|
SECTION("non-empty object with double")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", 4.2 }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x14, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x01, /// entry: double
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0xcd, 0xcc, 0xcc, 0xcc, 0xcc, 0xcc, 0x10, 0x40,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with string")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", "bsonstr" }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x18, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x02, /// entry: string (UTF-8)
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x08, 0x00, 0x00, 0x00, 'b', 's', 'o', 'n', 's', 't', 'r', '\x00',
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with null member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", nullptr }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x0C, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x0A, /// entry: null
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with integer (32-bit) member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", std::int32_t{0x12345678} }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x10, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, /// entry: int32
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x78, 0x56, 0x34, 0x12,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with integer (64-bit) member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", std::int64_t{0x1234567804030201} }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x14, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x12, /// entry: int64
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x01, 0x02, 0x03, 0x04, 0x78, 0x56, 0x34, 0x12,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with negative integer (32-bit) member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", std::int32_t{-1} }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x10, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, /// entry: int32
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0xFF, 0xFF, 0xFF, 0xFF,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with negative integer (64-bit) member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", std::int64_t{-1} }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x10, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, /// entry: int32
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0xFF, 0xFF, 0xFF, 0xFF,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with unsigned integer (64-bit) member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", std::uint64_t{0x1234567804030201} }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x14, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x12, /// entry: int64
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x01, 0x02, 0x03, 0x04, 0x78, 0x56, 0x34, 0x12,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with small unsigned integer member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", std::uint64_t{0x42} }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x10, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, /// entry: int32
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x42, 0x00, 0x00, 0x00,
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with object member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", json::object() }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x11, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x03, /// entry: embedded document
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x05, 0x00, 0x00, 0x00, // size (little endian)
|
|
// no entries
|
|
0x00, // end marker (embedded document)
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with array member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", json::array() }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x11, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x04, /// entry: embedded document
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x05, 0x00, 0x00, 0x00, // size (little endian)
|
|
// no entries
|
|
0x00, // end marker (embedded document)
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with non-empty array member")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", json::array({1, 2, 3, 4, 5, 6, 7, 8}) }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x49, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x04, /// entry: embedded document
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x3D, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, '0', 0x00, 0x01, 0x00, 0x00, 0x00,
|
|
0x10, '1', 0x00, 0x02, 0x00, 0x00, 0x00,
|
|
0x10, '2', 0x00, 0x03, 0x00, 0x00, 0x00,
|
|
0x10, '3', 0x00, 0x04, 0x00, 0x00, 0x00,
|
|
0x10, '4', 0x00, 0x05, 0x00, 0x00, 0x00,
|
|
0x10, '5', 0x00, 0x06, 0x00, 0x00, 0x00,
|
|
0x10, '6', 0x00, 0x07, 0x00, 0x00, 0x00,
|
|
0x10, '7', 0x00, 0x08, 0x00, 0x00, 0x00,
|
|
0x00, // end marker (embedded document)
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("array elements with non-conforming keys (lenient parsing)")
|
|
{
|
|
// documented lenient behavior (see gh-5333): BSON array element
|
|
// keys are not checked against the required decimal sequence
|
|
// "0", "1", "2", ... - elements are taken in encoded order
|
|
std::vector<std::uint8_t> const input =
|
|
{
|
|
0x26, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x04, 'e', 'n', 't', 'r', 'y', '\x00', // entry: embedded array
|
|
|
|
0x1A, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, '5', 0x00, 0x0A, 0x00, 0x00, 0x00, // key "5" (bogus) -> 10
|
|
0x10, 'x', 0x00, 0x14, 0x00, 0x00, 0x00, // key "x" (non-numeric) -> 20
|
|
0x10, '1', 0x00, 0x1E, 0x00, 0x00, 0x00, // key "1" (out of order) -> 30
|
|
0x00, // end marker (embedded array)
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const json expected = { { "entry", json::array({10, 20, 30}) } };
|
|
CHECK(json::from_bson(input) == expected);
|
|
}
|
|
|
|
SECTION("non-empty object with binary member")
|
|
{
|
|
const size_t N = 10;
|
|
const auto s = std::vector<std::uint8_t>(N, 'x');
|
|
json const j =
|
|
{
|
|
{ "entry", json::binary(s, 0) }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x1B, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x05, // entry: binary
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x0A, 0x00, 0x00, 0x00, // size of binary (little endian)
|
|
0x00, // Generic binary subtype
|
|
0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78,
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("non-empty object with binary member without subtype")
|
|
{
|
|
const size_t N = 10;
|
|
const auto s = std::vector<std::uint8_t>(N, 'x');
|
|
json const j =
|
|
{
|
|
{ "entry", json::binary(s) }
|
|
};
|
|
|
|
CHECK(!j.at("entry").get_binary().has_subtype());
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x1B, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x05, // entry: binary
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x0A, 0x00, 0x00, 0x00, // size of binary (little endian)
|
|
0x00, // Generic binary subtype
|
|
0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78,
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip adds the generic binary subtype
|
|
const auto roundtrip = json::from_bson(result);
|
|
CHECK(roundtrip != j);
|
|
CHECK(roundtrip.at("entry").get_binary().has_subtype());
|
|
CHECK(roundtrip.at("entry").get_binary().subtype() == 0);
|
|
CHECK(json::from_bson(result, true, false) == roundtrip);
|
|
}
|
|
|
|
SECTION("non-empty object with binary member with subtype")
|
|
{
|
|
// an MD5 hash
|
|
const std::vector<std::uint8_t> md5hash = {0xd7, 0x7e, 0x27, 0x54, 0xbe, 0x12, 0x37, 0xfe, 0xd6, 0x0c, 0x33, 0x98, 0x30, 0x3b, 0x8d, 0xc4};
|
|
json const j =
|
|
{
|
|
{ "entry", json::binary(md5hash, 5) }
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
0x21, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x05, // entry: binary
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x10, 0x00, 0x00, 0x00, // size of binary (little endian)
|
|
0x05, // MD5 binary subtype
|
|
0xd7, 0x7e, 0x27, 0x54, 0xbe, 0x12, 0x37, 0xfe, 0xd6, 0x0c, 0x33, 0x98, 0x30, 0x3b, 0x8d, 0xc4,
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
|
|
SECTION("binary member with subtype 0x02 (old binary) keeps its inner length prefix (lenient parsing)")
|
|
{
|
|
// documented lenient behavior (see gh-5333): the payload for
|
|
// binary subtype 0x02 ("old binary") is returned as-is,
|
|
// including its own inner 4-byte length prefix; it is not
|
|
// stripped or reinterpreted
|
|
std::vector<std::uint8_t> const input =
|
|
{
|
|
0x17, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x05, 'e', 'n', 't', 'r', 'y', '\x00', // entry: binary
|
|
|
|
0x06, 0x00, 0x00, 0x00, // size of binary (little endian)
|
|
0x02, // "old binary" subtype
|
|
0x02, 0x00, 0x00, 0x00, // inner length prefix (part of the old-binary payload)
|
|
0x68, 0x69, // payload ('h', 'i')
|
|
|
|
0x00 // end marker
|
|
};
|
|
|
|
// the inner length prefix is part of the (unmodified) payload
|
|
const std::vector<std::uint8_t> expected_payload = {0x02, 0x00, 0x00, 0x00, 0x68, 0x69};
|
|
const json expected = { { "entry", json::binary(expected_payload, 0x02) } };
|
|
CHECK(json::from_bson(input) == expected);
|
|
}
|
|
|
|
SECTION("Some more complex document")
|
|
{
|
|
json const j =
|
|
{
|
|
{"double", 42.5},
|
|
{"entry", 4.2},
|
|
{"number", 12345},
|
|
{"object", {{ "string", "value" }}}
|
|
};
|
|
|
|
std::vector<std::uint8_t> const expected =
|
|
{
|
|
/*size */ 0x4f, 0x00, 0x00, 0x00,
|
|
/*entry*/ 0x01, 'd', 'o', 'u', 'b', 'l', 'e', 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x40, 0x45, 0x40,
|
|
/*entry*/ 0x01, 'e', 'n', 't', 'r', 'y', 0x00, 0xcd, 0xcc, 0xcc, 0xcc, 0xcc, 0xcc, 0x10, 0x40,
|
|
/*entry*/ 0x10, 'n', 'u', 'm', 'b', 'e', 'r', 0x00, 0x39, 0x30, 0x00, 0x00,
|
|
/*entry*/ 0x03, 'o', 'b', 'j', 'e', 'c', 't', 0x00,
|
|
/*entry: obj-size */ 0x17, 0x00, 0x00, 0x00,
|
|
/*entry: obj-entry*/0x02, 's', 't', 'r', 'i', 'n', 'g', 0x00, 0x06, 0x00, 0x00, 0x00, 'v', 'a', 'l', 'u', 'e', 0,
|
|
/*entry: obj-term.*/0x00,
|
|
/*obj-term*/ 0x00
|
|
};
|
|
|
|
const auto result = json::to_bson(j);
|
|
CHECK(result == expected);
|
|
|
|
// roundtrip
|
|
CHECK(json::from_bson(result) == j);
|
|
CHECK(json::from_bson(result, true, false) == j);
|
|
}
|
|
}
|
|
|
|
SECTION("Examples from https://bsonspec.org/faq.html")
|
|
{
|
|
SECTION("Example 1")
|
|
{
|
|
std::vector<std::uint8_t> input = {0x16, 0x00, 0x00, 0x00, 0x02, 'h', 'e', 'l', 'l', 'o', 0x00, 0x06, 0x00, 0x00, 0x00, 'w', 'o', 'r', 'l', 'd', 0x00, 0x00};
|
|
const json parsed = json::from_bson(input);
|
|
const json expected = {{"hello", "world"}};
|
|
CHECK(parsed == expected);
|
|
const auto dumped = json::to_bson(parsed);
|
|
CHECK(dumped == input);
|
|
CHECK(json::from_bson(dumped) == expected);
|
|
}
|
|
|
|
SECTION("Example 2")
|
|
{
|
|
std::vector<std::uint8_t> input = {0x31, 0x00, 0x00, 0x00, 0x04, 'B', 'S', 'O', 'N', 0x00, 0x26, 0x00, 0x00, 0x00, 0x02, 0x30, 0x00, 0x08, 0x00, 0x00, 0x00, 'a', 'w', 'e', 's', 'o', 'm', 'e', 0x00, 0x01, 0x31, 0x00, 0x33, 0x33, 0x33, 0x33, 0x33, 0x33, 0x14, 0x40, 0x10, 0x32, 0x00, 0xc2, 0x07, 0x00, 0x00, 0x00, 0x00};
|
|
const json parsed = json::from_bson(input);
|
|
const json expected = {{"BSON", {"awesome", 5.05, 1986}}};
|
|
CHECK(parsed == expected);
|
|
const auto dumped = json::to_bson(parsed);
|
|
CHECK(dumped == input);
|
|
CHECK(json::from_bson(dumped) == expected);
|
|
}
|
|
}
|
|
}
|
|
|
|
TEST_CASE("regression test - BSON binary subtype rejects a value that doesn't fit a single byte")
|
|
{
|
|
json const doc255 = {{"b", json::binary({1, 2}, 255)}};
|
|
CHECK(json::from_bson(json::to_bson(doc255))["b"].get_binary().subtype() == 255);
|
|
|
|
CHECK_THROWS_AS(json::to_bson(json{{"b", json::binary({1, 2}, 256)}}), json::out_of_range);
|
|
#if JSON_DIAGNOSTICS
|
|
CHECK_THROWS_WITH_AS(json::to_bson(json {{"b", json::binary({1, 2}, 300)}}), "[json.exception.out_of_range.415] (/b) subtype 300 is too large for the BSON binary subtype (max 255)", json::out_of_range);
|
|
#else
|
|
CHECK_THROWS_WITH_AS(json::to_bson(json {{"b", json::binary({1, 2}, 300)}}), "[json.exception.out_of_range.415] subtype 300 is too large for the BSON binary subtype (max 255)", json::out_of_range);
|
|
#endif
|
|
}
|
|
|
|
TEST_CASE("BSON input/output_adapters")
|
|
{
|
|
const json json_representation =
|
|
{
|
|
{"double", 42.5},
|
|
{"entry", 4.2},
|
|
{"number", 12345},
|
|
{"object", {{ "string", "value" }}}
|
|
};
|
|
|
|
const std::vector<std::uint8_t> bson_representation =
|
|
{
|
|
/*size */ 0x4f, 0x00, 0x00, 0x00,
|
|
/*entry*/ 0x01, 'd', 'o', 'u', 'b', 'l', 'e', 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x40, 0x45, 0x40,
|
|
/*entry*/ 0x01, 'e', 'n', 't', 'r', 'y', 0x00, 0xcd, 0xcc, 0xcc, 0xcc, 0xcc, 0xcc, 0x10, 0x40,
|
|
/*entry*/ 0x10, 'n', 'u', 'm', 'b', 'e', 'r', 0x00, 0x39, 0x30, 0x00, 0x00,
|
|
/*entry*/ 0x03, 'o', 'b', 'j', 'e', 'c', 't', 0x00,
|
|
/*entry: obj-size */ 0x17, 0x00, 0x00, 0x00,
|
|
/*entry: obj-entry*/0x02, 's', 't', 'r', 'i', 'n', 'g', 0x00, 0x06, 0x00, 0x00, 0x00, 'v', 'a', 'l', 'u', 'e', 0,
|
|
/*entry: obj-term.*/0x00,
|
|
/*obj-term*/ 0x00
|
|
};
|
|
|
|
json j2;
|
|
CHECK_NOTHROW(j2 = json::from_bson(bson_representation));
|
|
|
|
// compare parsed JSON values
|
|
CHECK(json_representation == j2);
|
|
|
|
SECTION("roundtrips")
|
|
{
|
|
SECTION("std::ostringstream")
|
|
{
|
|
std::basic_ostringstream<char> ss;
|
|
json::to_bson(json_representation, ss);
|
|
const json j3 = json::from_bson(ss.str());
|
|
CHECK(json_representation == j3);
|
|
}
|
|
|
|
SECTION("std::string")
|
|
{
|
|
std::string s;
|
|
json::to_bson(json_representation, s);
|
|
const json j3 = json::from_bson(s);
|
|
CHECK(json_representation == j3);
|
|
}
|
|
|
|
SECTION("std::vector")
|
|
{
|
|
std::vector<std::uint8_t> v;
|
|
json::to_bson(json_representation, v);
|
|
const json j3 = json::from_bson(v);
|
|
CHECK(json_representation == j3);
|
|
}
|
|
}
|
|
}
|
|
|
|
|
|
TEST_CASE("Incomplete BSON Input")
|
|
{
|
|
SECTION("Incomplete BSON Input 1")
|
|
{
|
|
std::vector<std::uint8_t> const incomplete_bson =
|
|
{
|
|
0x0D, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x08, // entry: boolean
|
|
'e', 'n', 't' // unexpected EOF
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 9: syntax error while parsing BSON cstring: unexpected end of input", json::parse_error&);
|
|
|
|
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(0);
|
|
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("Incomplete BSON Input 2")
|
|
{
|
|
std::vector<std::uint8_t> const incomplete_bson =
|
|
{
|
|
0x0D, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x08, // entry: boolean, unexpected EOF
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 6: syntax error while parsing BSON cstring: unexpected end of input", json::parse_error&);
|
|
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(0);
|
|
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("Incomplete BSON Input 3")
|
|
{
|
|
std::vector<std::uint8_t> const incomplete_bson =
|
|
{
|
|
0x41, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x04, /// entry: embedded document
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0x35, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x10, 0x00, 0x01, 0x00, 0x00, 0x00,
|
|
0x10, 0x00, 0x02, 0x00, 0x00, 0x00
|
|
// missing input data...
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 28: syntax error while parsing BSON element list: unexpected end of input", json::parse_error&);
|
|
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(1);
|
|
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("Incomplete BSON Input 4")
|
|
{
|
|
std::vector<std::uint8_t> const incomplete_bson =
|
|
{
|
|
0x0D, 0x00, // size (incomplete), unexpected EOF
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 3: syntax error while parsing BSON number: unexpected end of input", json::parse_error&);
|
|
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(0);
|
|
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("Incomplete BSON Input 5")
|
|
{
|
|
std::vector<std::uint8_t> const incomplete_bson =
|
|
{
|
|
0x09, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x08, // entry: boolean
|
|
'b', '\x00' // key, unexpected EOF before the value
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 8: syntax error while parsing BSON number: unexpected end of input", json::parse_error&);
|
|
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(0);
|
|
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("Incomplete BSON Input 6")
|
|
{
|
|
std::vector<std::uint8_t> const incomplete_bson =
|
|
{
|
|
0x0F, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x05, // entry: binary
|
|
'b', '\x00', // key
|
|
0x00, 0x00, 0x00, 0x00 // length, unexpected EOF before the subtype
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 12: syntax error while parsing BSON number: unexpected end of input", json::parse_error&);
|
|
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(0);
|
|
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("Improve coverage")
|
|
{
|
|
SECTION("key")
|
|
{
|
|
json const j = {{"key", "value"}};
|
|
auto bson_vec = json::to_bson(j);
|
|
SaxCountdown scp(2);
|
|
CHECK(!json::sax_parse(bson_vec, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("array")
|
|
{
|
|
json const j =
|
|
{
|
|
{ "entry", json::array() }
|
|
};
|
|
auto bson_vec = json::to_bson(j);
|
|
SaxCountdown scp(2);
|
|
CHECK(!json::sax_parse(bson_vec, &scp, json::input_format_t::bson));
|
|
}
|
|
}
|
|
}
|
|
|
|
// the test catches the exceptions of invalid input
|
|
#if !defined(JSON_NOEXCEPTION)
|
|
TEST_CASE("BSON keys from contiguous and stream input")
|
|
{
|
|
// contiguous input reads a key up to its \x00-byte in one step, a stream
|
|
// reads it byte by byte; both must give the same value or error for the
|
|
// complete document and for every truncation of it
|
|
const json j = {{"", true}, {"k", {1, 2, 3}}, {std::string(40, 'x'), {{"nested key", "value"}}}};
|
|
const std::vector<std::uint8_t> bson = json::to_bson(j);
|
|
CHECK(json::from_bson(bson) == j);
|
|
|
|
for (std::size_t length = 0; length <= bson.size(); ++length)
|
|
{
|
|
CAPTURE(length)
|
|
const std::vector<std::uint8_t> input(bson.begin(), bson.begin() + static_cast<std::ptrdiff_t>(length));
|
|
std::string from_vector;
|
|
std::string from_stream;
|
|
try
|
|
{
|
|
from_vector = json::from_bson(input).dump();
|
|
}
|
|
catch (const json::parse_error& e)
|
|
{
|
|
from_vector = e.what();
|
|
}
|
|
try
|
|
{
|
|
std::istringstream stream(std::string(input.begin(), input.end()));
|
|
from_stream = json::from_bson(stream).dump();
|
|
}
|
|
catch (const json::parse_error& e)
|
|
{
|
|
from_stream = e.what();
|
|
}
|
|
CHECK(from_vector == from_stream);
|
|
}
|
|
}
|
|
#endif
|
|
|
|
TEST_CASE("Negative size of binary value")
|
|
{
|
|
// invalid BSON: the size of the binary value is -1
|
|
std::vector<std::uint8_t> const input =
|
|
{
|
|
0x21, 0x00, 0x00, 0x00, // size (little endian)
|
|
0x05, // entry: binary
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
|
|
0xFF, 0xFF, 0xFF, 0xFF, // size of binary (little endian)
|
|
0x05, // MD5 binary subtype
|
|
0xd7, 0x7e, 0x27, 0x54, 0xbe, 0x12, 0x37, 0xfe, 0xd6, 0x0c, 0x33, 0x98, 0x30, 0x3b, 0x8d, 0xc4,
|
|
|
|
0x00 // end marker
|
|
};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1", json::parse_error);
|
|
}
|
|
|
|
TEST_CASE("Unsupported BSON input")
|
|
{
|
|
std::vector<std::uint8_t> const bson =
|
|
{
|
|
0x0C, 0x00, 0x00, 0x00, // size (little endian)
|
|
0xFF, // entry type: Min key (not supported yet)
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
0x00 // end marker
|
|
};
|
|
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(bson), "[json.exception.parse_error.114] parse error at byte 5: Unsupported BSON record type 0xFF", json::parse_error&);
|
|
CHECK(json::from_bson(bson, true, false).is_discarded());
|
|
|
|
SaxCountdown scp(0);
|
|
CHECK(!json::sax_parse(bson, &scp, json::input_format_t::bson));
|
|
}
|
|
|
|
TEST_CASE("BSON document size mismatch")
|
|
{
|
|
json _;
|
|
|
|
SECTION("top-level document declaring more bytes than it contains")
|
|
{
|
|
// empty object, but the length prefix claims 6 bytes instead of 5
|
|
std::vector<std::uint8_t> const input = {0x06, 0x00, 0x00, 0x00, 0x00};
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)", json::parse_error&);
|
|
CHECK(json::from_bson(input, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("top-level document with a negative size")
|
|
{
|
|
std::vector<std::uint8_t> const input = {0xFF, 0xFF, 0xFF, 0xFF, 0x00};
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size -1 does not match the number of bytes read (5)", json::parse_error&);
|
|
CHECK(json::from_bson(input, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("embedded document whose size disagrees with its terminator")
|
|
{
|
|
// the embedded document "d" declares 0x7FFFFFFF bytes but its 0x00
|
|
// terminator falls right after {"a":null}; the length prefix would
|
|
// otherwise let the following "h" element be read as a member of the
|
|
// enclosing document instead of "d"
|
|
std::vector<std::uint8_t> const input =
|
|
{
|
|
0x00, 0x00, 0x00, 0x00, // outer size
|
|
0x03, 'd', 0x00, // entry: embedded document "d"
|
|
0xFF, 0xFF, 0xFF, 0x7F, // embedded size 0x7FFFFFFF
|
|
0x0A, 'a', 0x00, // entry: null "a"
|
|
0x00, // embedded end marker
|
|
0x08, 'h', 0x00, 0x01, // entry: bool "h" = true
|
|
0x00 // outer end marker
|
|
};
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON document: document size 2147483647 does not match the number of bytes read (8)", json::parse_error&);
|
|
CHECK(json::from_bson(input, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("embedded array whose size disagrees with its terminator")
|
|
{
|
|
// array [42] is 12 bytes, but the length prefix claims 13
|
|
std::vector<std::uint8_t> const input =
|
|
{
|
|
0x00, 0x00, 0x00, 0x00, // outer size
|
|
0x04, 'a', 0x00, // entry: array "a"
|
|
0x0D, 0x00, 0x00, 0x00, // array size 13 (real is 12)
|
|
0x10, '0', 0x00, 0x2A, 0x00, 0x00, 0x00, // entry: int32 "0" = 42
|
|
0x00, // array end marker
|
|
0x00 // outer end marker
|
|
};
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 19: syntax error while parsing BSON document: document size 13 does not match the number of bytes read (12)", json::parse_error&);
|
|
CHECK(json::from_bson(input, true, false).is_discarded());
|
|
}
|
|
}
|
|
|
|
TEST_CASE("BSON nesting does not consume the call stack")
|
|
{
|
|
// An embedded document or array used to be read by calling back into the
|
|
// document reader, so the native call stack grew with the nesting depth of
|
|
// the input (#5104). The open documents are kept on a heap stack now.
|
|
//
|
|
// Deeply nested values must not be compared, copied or dumped here: those
|
|
// operations are still recursive and would reintroduce the crash.
|
|
|
|
// A document nested deeply enough to have crashed. The bytes are built
|
|
// here rather than with to_bson(), because the writer still recurses once
|
|
// per level and would overflow the stack before the reader is ever
|
|
// reached. Every level is
|
|
// <int32 size> 0x03 'a' 0x00 <inner document> 0x00
|
|
// so a level is eight bytes larger than the one it holds, and the sizes
|
|
// can be filled in from the outside in.
|
|
const std::size_t depth = 30000;
|
|
std::vector<uint8_t> input;
|
|
input.reserve(5 + (8 * depth));
|
|
for (std::size_t i = 0; i < depth; ++i)
|
|
{
|
|
const auto size = static_cast<std::uint32_t>(5 + (8 * (depth - i)));
|
|
input.push_back(static_cast<uint8_t>(size & 0xFF));
|
|
input.push_back(static_cast<uint8_t>((size >> 8) & 0xFF));
|
|
input.push_back(static_cast<uint8_t>((size >> 16) & 0xFF));
|
|
input.push_back(static_cast<uint8_t>((size >> 24) & 0xFF));
|
|
input.push_back(0x03); // embedded document
|
|
input.push_back('a');
|
|
input.push_back(0x00);
|
|
}
|
|
// the innermost document is empty, then one terminator closes each level
|
|
input.insert(input.end(), {0x05, 0x00, 0x00, 0x00, 0x00});
|
|
input.insert(input.end(), depth, 0x00);
|
|
|
|
SECTION("a well-formed deep document is read through the SAX interface")
|
|
{
|
|
SaxCountdown accept_all(1000000);
|
|
CHECK(json::sax_parse(input, &accept_all, json::input_format_t::bson));
|
|
}
|
|
|
|
SECTION("a well-formed deep document is read into a value")
|
|
{
|
|
json j = json::from_bson(input);
|
|
|
|
// walked rather than compared: comparing, copying or dumping a value
|
|
// this deep is still recursive
|
|
std::size_t measured = 0;
|
|
const json* q = &j;
|
|
while (q->is_object() && !q->empty())
|
|
{
|
|
q = &q->begin().value();
|
|
++measured;
|
|
}
|
|
CHECK(measured == depth);
|
|
}
|
|
|
|
SECTION("embedded documents and arrays are still read the same way")
|
|
{
|
|
const json values = {{"a", {{"b", {{"c", 1}}}}}};
|
|
CHECK(json::from_bson(json::to_bson(values)) == values);
|
|
|
|
const json array = {{"a", {1, 2, 3}}};
|
|
CHECK(json::from_bson(json::to_bson(array)) == array);
|
|
|
|
const json mixed = {{"a", {json{{"x", 1}}, json{{"y", 2}}}}};
|
|
CHECK(json::from_bson(json::to_bson(mixed)) == mixed);
|
|
|
|
CHECK(json::from_bson(json::to_bson(json::object())) == json::object());
|
|
}
|
|
|
|
SECTION("a size that does not match is still reported per document")
|
|
{
|
|
// the embedded document claims one byte too many
|
|
std::vector<uint8_t> const bad =
|
|
{
|
|
0x15, 0x00, 0x00, 0x00, 0x03, 'a', 0x00,
|
|
0x0D, 0x00, 0x00, 0x00, 0x08, 'b', 0x00, 0x01, 0x00,
|
|
0x00
|
|
};
|
|
json _;
|
|
CHECK_THROWS_AS(_ = json::from_bson(bad), json::parse_error&);
|
|
CHECK(json::from_bson(bad, true, false).is_discarded());
|
|
}
|
|
}
|
|
|
|
TEST_CASE("BSON input that cannot be read is discarded by every overload")
|
|
{
|
|
std::vector<std::uint8_t> input = json::to_bson(json({{"a", {1, 2}}}));
|
|
input.pop_back();
|
|
|
|
json _;
|
|
CHECK_THROWS_AS(_ = json::from_bson(input.begin(), input.end()), json::parse_error&);
|
|
CHECK(json::from_bson(input, true, false).is_discarded());
|
|
CHECK(json::from_bson(input.begin(), input.end(), true, false).is_discarded());
|
|
CHECK(json::from_bson(input.data(), input.size(), true, false).is_discarded());
|
|
CHECK(json::from_bson({input.data(), input.size()}, true, false).is_discarded());
|
|
}
|
|
|
|
TEST_CASE("BSON SAX parsing stops at every event")
|
|
{
|
|
// Containers are opened and closed by the loop that reads them; a SAX
|
|
// handler that rejects any event - including the end of a nested
|
|
// container - must stop the parse right there.
|
|
const auto count_events = [](const std::vector<std::uint8_t>& input)
|
|
{
|
|
int events = 0;
|
|
while (true)
|
|
{
|
|
SaxCountdown scp(events);
|
|
if (json::sax_parse(input, &scp, json::input_format_t::bson))
|
|
{
|
|
return events;
|
|
}
|
|
++events;
|
|
REQUIRE(events < 1000);
|
|
}
|
|
};
|
|
|
|
// 20 events: every container kind closes inside another one
|
|
const json j = json::parse(R"({"a": [1, {"b": []}], "c": {"d": [[2]]}})");
|
|
CHECK(count_events(json::to_bson(j)) == 20);
|
|
}
|
|
|
|
TEST_CASE("BSON numerical data")
|
|
{
|
|
SECTION("number")
|
|
{
|
|
SECTION("signed")
|
|
{
|
|
SECTION("std::int64_t: INT64_MIN .. INT32_MIN-1")
|
|
{
|
|
std::vector<int64_t> const numbers
|
|
{
|
|
(std::numeric_limits<int64_t>::min)(),
|
|
-1000000000000000000LL,
|
|
-100000000000000000LL,
|
|
-10000000000000000LL,
|
|
-1000000000000000LL,
|
|
-100000000000000LL,
|
|
-10000000000000LL,
|
|
-1000000000000LL,
|
|
-100000000000LL,
|
|
-10000000000LL,
|
|
static_cast<std::int64_t>((std::numeric_limits<std::int32_t>::min)()) - 1,
|
|
};
|
|
|
|
for (const auto i : numbers)
|
|
{
|
|
|
|
CAPTURE(i)
|
|
|
|
json const j =
|
|
{
|
|
{ "entry", i }
|
|
};
|
|
CHECK(j.at("entry").is_number_integer());
|
|
|
|
std::uint64_t const iu = *reinterpret_cast<const std::uint64_t*>(&i);
|
|
std::vector<std::uint8_t> const expected_bson =
|
|
{
|
|
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
|
|
0x12u, /// entry: int64
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
|
|
0x00u // end marker
|
|
};
|
|
|
|
const auto bson = json::to_bson(j);
|
|
CHECK(bson == expected_bson);
|
|
|
|
auto j_roundtrip = json::from_bson(bson);
|
|
|
|
CHECK(j_roundtrip.at("entry").is_number_integer());
|
|
CHECK(j_roundtrip == j);
|
|
CHECK(json::from_bson(bson, true, false) == j);
|
|
|
|
}
|
|
}
|
|
|
|
SECTION("signed std::int32_t: INT32_MIN .. INT32_MAX")
|
|
{
|
|
std::vector<int32_t> const numbers
|
|
{
|
|
(std::numeric_limits<int32_t>::min)(),
|
|
-2147483647L,
|
|
-1000000000L,
|
|
-100000000L,
|
|
-10000000L,
|
|
-1000000L,
|
|
-100000L,
|
|
-10000L,
|
|
-1000L,
|
|
-100L,
|
|
-10L,
|
|
-1L,
|
|
0L,
|
|
1L,
|
|
10L,
|
|
100L,
|
|
1000L,
|
|
10000L,
|
|
100000L,
|
|
1000000L,
|
|
10000000L,
|
|
100000000L,
|
|
1000000000L,
|
|
2147483646L,
|
|
(std::numeric_limits<int32_t>::max)()
|
|
};
|
|
|
|
for (const auto i : numbers)
|
|
{
|
|
|
|
CAPTURE(i)
|
|
|
|
json const j =
|
|
{
|
|
{ "entry", i }
|
|
};
|
|
CHECK(j.at("entry").is_number_integer());
|
|
|
|
std::uint32_t const iu = *reinterpret_cast<const std::uint32_t*>(&i);
|
|
std::vector<std::uint8_t> const expected_bson =
|
|
{
|
|
0x10u, 0x00u, 0x00u, 0x00u, // size (little endian)
|
|
0x10u, /// entry: int32
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
|
|
0x00u // end marker
|
|
};
|
|
|
|
const auto bson = json::to_bson(j);
|
|
CHECK(bson == expected_bson);
|
|
|
|
auto j_roundtrip = json::from_bson(bson);
|
|
|
|
CHECK(j_roundtrip.at("entry").is_number_integer());
|
|
CHECK(j_roundtrip == j);
|
|
CHECK(json::from_bson(bson, true, false) == j);
|
|
|
|
}
|
|
}
|
|
|
|
SECTION("signed std::int64_t: INT32_MAX+1 .. INT64_MAX")
|
|
{
|
|
std::vector<int64_t> const numbers
|
|
{
|
|
(std::numeric_limits<int64_t>::max)(),
|
|
1000000000000000000LL,
|
|
100000000000000000LL,
|
|
10000000000000000LL,
|
|
1000000000000000LL,
|
|
100000000000000LL,
|
|
10000000000000LL,
|
|
1000000000000LL,
|
|
100000000000LL,
|
|
10000000000LL,
|
|
static_cast<std::int64_t>((std::numeric_limits<int32_t>::max)()) + 1,
|
|
};
|
|
|
|
for (const auto i : numbers)
|
|
{
|
|
|
|
CAPTURE(i)
|
|
|
|
json const j =
|
|
{
|
|
{ "entry", i }
|
|
};
|
|
CHECK(j.at("entry").is_number_integer());
|
|
|
|
std::uint64_t const iu = *reinterpret_cast<const std::uint64_t*>(&i);
|
|
std::vector<std::uint8_t> const expected_bson =
|
|
{
|
|
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
|
|
0x12u, /// entry: int64
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
|
|
0x00u // end marker
|
|
};
|
|
|
|
const auto bson = json::to_bson(j);
|
|
CHECK(bson == expected_bson);
|
|
|
|
auto j_roundtrip = json::from_bson(bson);
|
|
|
|
CHECK(j_roundtrip.at("entry").is_number_integer());
|
|
CHECK(j_roundtrip == j);
|
|
CHECK(json::from_bson(bson, true, false) == j);
|
|
|
|
}
|
|
}
|
|
}
|
|
|
|
SECTION("unsigned")
|
|
{
|
|
SECTION("unsigned std::uint64_t: 0 .. INT32_MAX")
|
|
{
|
|
std::vector<std::uint64_t> const numbers
|
|
{
|
|
0ULL,
|
|
1ULL,
|
|
10ULL,
|
|
100ULL,
|
|
1000ULL,
|
|
10000ULL,
|
|
100000ULL,
|
|
1000000ULL,
|
|
10000000ULL,
|
|
100000000ULL,
|
|
1000000000ULL,
|
|
2147483646ULL,
|
|
static_cast<std::uint64_t>((std::numeric_limits<int32_t>::max)())
|
|
};
|
|
|
|
for (const auto i : numbers)
|
|
{
|
|
|
|
CAPTURE(i)
|
|
|
|
json const j =
|
|
{
|
|
{ "entry", i }
|
|
};
|
|
|
|
auto iu = i;
|
|
std::vector<std::uint8_t> const expected_bson =
|
|
{
|
|
0x10u, 0x00u, 0x00u, 0x00u, // size (little endian)
|
|
0x10u, /// entry: int32
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
|
|
0x00u // end marker
|
|
};
|
|
|
|
const auto bson = json::to_bson(j);
|
|
CHECK(bson == expected_bson);
|
|
|
|
auto j_roundtrip = json::from_bson(bson);
|
|
|
|
CHECK(j.at("entry").is_number_unsigned());
|
|
CHECK(j_roundtrip.at("entry").is_number_integer());
|
|
CHECK(j_roundtrip == j);
|
|
CHECK(json::from_bson(bson, true, false) == j);
|
|
|
|
}
|
|
}
|
|
|
|
SECTION("unsigned std::uint64_t: INT32_MAX+1 .. INT64_MAX")
|
|
{
|
|
std::vector<std::uint64_t> const numbers
|
|
{
|
|
static_cast<std::uint64_t>((std::numeric_limits<std::int32_t>::max)()) + 1,
|
|
4000000000ULL,
|
|
static_cast<std::uint64_t>((std::numeric_limits<std::uint32_t>::max)()),
|
|
10000000000ULL,
|
|
100000000000ULL,
|
|
1000000000000ULL,
|
|
10000000000000ULL,
|
|
100000000000000ULL,
|
|
1000000000000000ULL,
|
|
10000000000000000ULL,
|
|
100000000000000000ULL,
|
|
1000000000000000000ULL,
|
|
static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()),
|
|
};
|
|
|
|
for (const auto i : numbers)
|
|
{
|
|
|
|
CAPTURE(i)
|
|
|
|
json const j =
|
|
{
|
|
{ "entry", i }
|
|
};
|
|
|
|
auto iu = i;
|
|
std::vector<std::uint8_t> const expected_bson =
|
|
{
|
|
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
|
|
0x12u, /// entry: int64
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
|
|
0x00u // end marker
|
|
};
|
|
|
|
const auto bson = json::to_bson(j);
|
|
CHECK(bson == expected_bson);
|
|
|
|
auto j_roundtrip = json::from_bson(bson);
|
|
|
|
CHECK(j.at("entry").is_number_unsigned());
|
|
CHECK(j_roundtrip.at("entry").is_number_integer());
|
|
CHECK(j_roundtrip == j);
|
|
CHECK(json::from_bson(bson, true, false) == j);
|
|
}
|
|
}
|
|
|
|
SECTION("unsigned std::uint64_t: INT64_MAX+1 .. UINT64_MAX")
|
|
{
|
|
std::vector<std::uint64_t> const numbers
|
|
{
|
|
static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1ULL,
|
|
0xffffffffffffffff,
|
|
};
|
|
|
|
for (const auto i : numbers)
|
|
{
|
|
|
|
CAPTURE(i)
|
|
|
|
json const j =
|
|
{
|
|
{ "entry", i }
|
|
};
|
|
|
|
auto iu = i;
|
|
std::vector<std::uint8_t> const expected_bson =
|
|
{
|
|
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
|
|
0x11u, /// entry: uint64
|
|
'e', 'n', 't', 'r', 'y', '\x00',
|
|
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
|
|
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
|
|
0x00u // end marker
|
|
};
|
|
|
|
const auto bson = json::to_bson(j);
|
|
CHECK(bson == expected_bson);
|
|
|
|
auto j_roundtrip = json::from_bson(bson);
|
|
|
|
CHECK(j.at("entry").is_number_unsigned());
|
|
CHECK(j_roundtrip.at("entry").is_number_unsigned());
|
|
CHECK(j_roundtrip == j);
|
|
CHECK(json::from_bson(bson, true, false) == j);
|
|
}
|
|
}
|
|
|
|
}
|
|
}
|
|
}
|
|
|
|
TEST_CASE("Parse BSON directly from a file using iterator and sentinel")
|
|
{
|
|
std::string const filename = TEST_DATA_DIRECTORY "/json.org/1.json";
|
|
|
|
std::ifstream f_json(filename);
|
|
const json expected = json::parse(f_json);
|
|
|
|
std::ifstream file(filename + ".bson", std::ios::binary);
|
|
const std::istreambuf_iterator<char> first(file);
|
|
const json parsed = json::from_bson(first, utils::istreambuf_sentinel{});
|
|
CHECK(parsed == expected);
|
|
}
|
|
|
|
TEST_CASE("BSON round-trip invariants")
|
|
{
|
|
// This checks what the parse_bson_fuzzer driver checks (see
|
|
// tests/src/fuzzer-parse_bson.cpp), so that a regression shows up in CI
|
|
// rather than as an OSS-Fuzz report: anything from_bson() returns (j1)
|
|
// can be serialized, parsed back (j2), and serialized again to reproduce
|
|
// the exact bytes. BSON only serializes objects, so non-object corpus
|
|
// values are skipped.
|
|
for (const auto& j0 : utils::round_trip_corpus::values())
|
|
{
|
|
if (!j0.is_object())
|
|
{
|
|
continue;
|
|
}
|
|
|
|
json j1;
|
|
try
|
|
{
|
|
// turn the corpus value into a value as from_bson() returns it
|
|
j1 = json::from_bson(json::to_bson(j0));
|
|
}
|
|
catch (const json::exception&)
|
|
{
|
|
// the fuzzer driver only ever sees values from_bson() actually
|
|
// produced, so skip corpus values that do not survive the
|
|
// round trip here, too
|
|
continue;
|
|
}
|
|
|
|
INFO("j1 = " << j1.dump());
|
|
const std::vector<std::uint8_t> vec = json::to_bson(j1);
|
|
json j2;
|
|
// anything the library writes must be parsable by the library
|
|
REQUIRE_NOTHROW(j2 = json::from_bson(vec));
|
|
CHECK(json::to_bson(j2) == vec);
|
|
}
|
|
}
|
|
|
|
TEST_CASE("BSON roundtrips" * doctest::skip())
|
|
{
|
|
SECTION("reference files")
|
|
{
|
|
for (const std::string filename :
|
|
{
|
|
TEST_DATA_DIRECTORY "/json.org/1.json",
|
|
TEST_DATA_DIRECTORY "/json.org/2.json",
|
|
TEST_DATA_DIRECTORY "/json.org/3.json",
|
|
TEST_DATA_DIRECTORY "/json.org/4.json",
|
|
TEST_DATA_DIRECTORY "/json.org/5.json"
|
|
})
|
|
{
|
|
CAPTURE(filename)
|
|
|
|
{
|
|
INFO_WITH_TEMP(filename + ": std::vector<std::uint8_t>");
|
|
// parse JSON file
|
|
std::ifstream f_json(filename);
|
|
const json j1 = json::parse(f_json);
|
|
|
|
// parse BSON file
|
|
auto packed = utils::read_binary_file(filename + ".bson");
|
|
json j2;
|
|
CHECK_NOTHROW(j2 = json::from_bson(packed));
|
|
|
|
// compare parsed JSON values
|
|
CHECK(j1 == j2);
|
|
}
|
|
|
|
{
|
|
INFO_WITH_TEMP(filename + ": std::ifstream");
|
|
// parse JSON file
|
|
std::ifstream f_json(filename);
|
|
const json j1 = json::parse(f_json);
|
|
|
|
// parse BSON file
|
|
std::ifstream f_bson(filename + ".bson", std::ios::binary);
|
|
json j2;
|
|
CHECK_NOTHROW(j2 = json::from_bson(f_bson));
|
|
|
|
// compare parsed JSON values
|
|
CHECK(j1 == j2);
|
|
}
|
|
|
|
{
|
|
INFO_WITH_TEMP(filename + ": uint8_t* and size");
|
|
// parse JSON file
|
|
std::ifstream f_json(filename);
|
|
const json j1 = json::parse(f_json);
|
|
|
|
// parse BSON file
|
|
auto packed = utils::read_binary_file(filename + ".bson");
|
|
json j2;
|
|
CHECK_NOTHROW(j2 = json::from_bson({packed.data(), packed.size()}));
|
|
|
|
// compare parsed JSON values
|
|
CHECK(j1 == j2);
|
|
}
|
|
|
|
{
|
|
INFO_WITH_TEMP(filename + ": output to output adapters");
|
|
// parse JSON file
|
|
std::ifstream f_json(filename);
|
|
json const j1 = json::parse(f_json);
|
|
|
|
// parse BSON file
|
|
auto packed = utils::read_binary_file(filename + ".bson");
|
|
|
|
{
|
|
INFO_WITH_TEMP(filename + ": output adapters: std::vector<std::uint8_t>");
|
|
std::vector<std::uint8_t> vec;
|
|
json::to_bson(j1, vec);
|
|
|
|
if (vec != packed)
|
|
{
|
|
// the exact serializations may differ due to the order of
|
|
// object keys; in these cases, just compare whether both
|
|
// serializations create the same JSON value
|
|
CHECK(json::from_bson(vec) == json::from_bson(packed));
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
TEST_CASE("BSON: deeply nested values")
|
|
{
|
|
SECTION("documents and arrays round-trip at every depth")
|
|
{
|
|
// nested documents and arrays, with siblings on every level, so
|
|
// every length prefix covers entries of both kinds
|
|
json value = "leaf";
|
|
for (std::size_t depth = 0; depth <= 300; ++depth)
|
|
{
|
|
CAPTURE(depth);
|
|
const json document = {{"value", value}, {"n", depth}};
|
|
CHECK(json::from_bson(json::to_bson(document)) == document);
|
|
|
|
value = depth % 2 == 0 ? json{{"a", std::move(value)}, {"b", {1, "x"}}} :
|
|
json::array({std::move(value), depth, json::object()});
|
|
}
|
|
}
|
|
|
|
SECTION("a key containing U+0000 is rejected before anything is written")
|
|
{
|
|
json value = json::object({{std::string("bad\0key", 7), 1}});
|
|
for (std::size_t depth = 0; depth < 200; ++depth)
|
|
{
|
|
value = json{{"a", {{"b", 1}}}, {"z", std::move(value)}};
|
|
}
|
|
std::vector<std::uint8_t> output;
|
|
CHECK_THROWS_AS(json::to_bson(value, output), json::out_of_range&);
|
|
CHECK(output.empty());
|
|
}
|
|
|
|
SECTION("a binary subtype that doesn't fit a byte is rejected before anything is written (#5675)")
|
|
{
|
|
// the offending value is nested, so this also covers that the check
|
|
// is not limited to a directly written value's own document
|
|
json const j = {{"a", {{"b", json::binary({1, 2}, 300)}}}};
|
|
|
|
std::vector<std::uint8_t> vector_output;
|
|
CHECK_THROWS_AS(json::to_bson(j, vector_output), json::out_of_range&);
|
|
CHECK(vector_output.empty());
|
|
|
|
std::string string_output;
|
|
CHECK_THROWS_AS(json::to_bson(j, string_output), json::out_of_range&);
|
|
CHECK(string_output.empty());
|
|
}
|
|
|
|
SECTION("values nested too deeply for the call stack (#5392)")
|
|
{
|
|
// serializing recursed once per nesting level, and computed every
|
|
// nested document's length by walking everything below it again.
|
|
// The values are only parsed, serialized and walked, never copied or
|
|
// compared, since those recurse too.
|
|
const std::size_t depth = 100000;
|
|
for (const bool objects :
|
|
{
|
|
false, true
|
|
})
|
|
{
|
|
CAPTURE(objects);
|
|
std::string text = "{\"a\":";
|
|
for (std::size_t i = 0; i < depth; ++i)
|
|
{
|
|
text += objects ? "{\"a\":" : "[";
|
|
}
|
|
text += "1";
|
|
text.append(depth, objects ? '}' : ']');
|
|
text += "}";
|
|
|
|
const auto bson = json::to_bson(json::parse(text));
|
|
const auto result = json::from_bson(bson);
|
|
const json* p = &result.at("a");
|
|
for (std::size_t i = 0; i < depth; ++i)
|
|
{
|
|
p = objects ? &p->at("a") : &p->at(0);
|
|
}
|
|
CHECK(*p == 1);
|
|
}
|
|
}
|
|
}
|
|
|
|
TEST_CASE("Invalid document size handling")
|
|
{
|
|
SECTION("document size must be at least 5")
|
|
{
|
|
std::vector<std::uint8_t> const v = {0x04, 0x00, 0x00, 0x00, 0x00};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 4 does not match the number of bytes read (5)", json::parse_error&);
|
|
CHECK(json::from_bson(v, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("declared document size must match consumed bytes (extra trailing element)")
|
|
{
|
|
// Declares 5-byte empty document but appends an int32 element after the declared end.
|
|
std::vector<std::uint8_t> const v =
|
|
{
|
|
0x05, 0x00, 0x00, 0x00,
|
|
0x10, 'a', 'd', 'm', 'i', 'n', 0x00,
|
|
0x01, 0x00, 0x00, 0x00,
|
|
0x00
|
|
};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 16: syntax error while parsing BSON document: document size 5 does not match the number of bytes read (16)", json::parse_error&);
|
|
CHECK(json::from_bson(v, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("declared document size must match consumed bytes (premature terminator)")
|
|
{
|
|
// Declares 32-byte document but only contains the size field followed by an immediate terminator.
|
|
std::vector<std::uint8_t> const v =
|
|
{
|
|
0x20, 0x00, 0x00, 0x00,
|
|
0x00
|
|
};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 32 does not match the number of bytes read (5)", json::parse_error&);
|
|
CHECK(json::from_bson(v, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("array declared size must match consumed bytes")
|
|
{
|
|
// Outer object contains an array "a" that declares 5 bytes (empty) but
|
|
// actually contains an int32 element before its terminator.
|
|
std::vector<std::uint8_t> const v =
|
|
{
|
|
0x14, 0x00, 0x00, 0x00, // object size = 20
|
|
0x04, 'a', 0x00, // key "a", array type
|
|
0x05, 0x00, 0x00, 0x00, // array declared size = 5 (empty)
|
|
0x10, '0', 0x00, 0x01, 0x00, 0x00, 0x00, // extra int32 element "0" = 1
|
|
0x00, // array terminator
|
|
0x00 // object terminator
|
|
};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 19: syntax error while parsing BSON document: document size 5 does not match the number of bytes read (12)", json::parse_error&);
|
|
CHECK(json::from_bson(v, true, false).is_discarded());
|
|
}
|
|
|
|
SECTION("BSON string must end with 0x00")
|
|
{
|
|
// Length-prefixed string whose terminator byte is 'X' (0x58), not 0x00.
|
|
std::vector<std::uint8_t> const v =
|
|
{
|
|
0x0F, 0x00, 0x00, 0x00,
|
|
0x02, 's', 0x00,
|
|
0x02, 0x00, 0x00, 0x00,
|
|
'A', 'X',
|
|
0x00
|
|
};
|
|
json _;
|
|
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 13: syntax error while parsing BSON string: BSON string is not null-terminated", json::parse_error&);
|
|
CHECK(json::from_bson(v, true, false).is_discarded());
|
|
}
|
|
}
|