Files
json/tests/src/unit-bson.cpp
T
Niels Lohmann e5a89d671f Fix lint debt: enum-macro NOLINTs, doctest as SYSTEM, no-op analyzer (#5737)
* Drop stale LCOV_EXCL_LINE from the json_pointer out_of_range.410 throw

The comment said the size_type overflow check in array_index() is only
triggered on special platforms like 32-bit, and the throw was excluded
from coverage. On 64-bit platforms the check is true for SIZE_MAX
itself, and unit-json_pointer.cpp has asserted that case four times
since #5395, so the line is executed in the coverage job. Reword the
comment and remove the exclusion marker so the coverage report notices
if the tests stop reaching it.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name all three C-array check aliases in the enum-macro NOLINTs

NLOHMANN_JSON_SERIALIZE_ENUM(_STRICT) suppressed the c-array warning
under modernize-avoid-c-arrays only, but clang-tidy emits the same
diagnostic under the aliases cppcoreguidelines-avoid-c-arrays and
hicpp-avoid-c-arrays too. Any user running those checks got a false
positive at every macro expansion, and our own tests needed a local
NOLINT at each call site to work around it.

Name all three aliases in the four macro comments instead, and drop
the now-redundant c-array names from the five test call-site NOLINTs.
Comment-only change; behavior, the public API, and the ABI do not
change. Ran make amalgamate.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Include doctest as a SYSTEM directory instead of disabling warnings for all tests

test_main added -Wno-deprecated and -Wno-float-equal as PUBLIC compile
options for every non-MSVC compiler, so they were applied to every
translation unit, library headers included, and silenced the CI
warnings meant to check the library's own -Wfloat-equal pragmas. The
only code that actually needed the suppression was the vendored
doctest.h, which was included as a normal (non-SYSTEM) directory.

Include thirdparty/doctest as SYSTEM for test_main, matching what
tests/abi/CMakeLists.txt already does, and drop the two suppressions
from both targets. Verified locally that unit-comparison,
unit-conversions and unit-constructor1 compile clean with
-Werror -Weverything and doctest as -isystem, and that CMake still
configures with JSON_BuildTests=ON.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the no-op ci_clang_analyze target

ci_clang_analyze configured the build with the real compiler and only
then wrapped ninja with scan-build. scan-build intercepts compiles by
overriding CC/CXX, but build.ninja already had the compiler path baked
in from the configure step, so every run bypassed the analyzer: CI
logs show "No bugs found" after a normal build, never an analysis.
The job also used Debian's frozen clang-tools-14 rather than the
image's own clang, and CLANG_ANALYZER_CHECKS still named three
valist.* checkers that current clang merged into security.VAList.

ci_clang_tidy already runs every clang-analyzer-* check (via
.clang-tidy's "Checks: '*'") with warnings as errors, so nothing is
lost by removing the dead job. Delete ci_clang_analyze,
CLANG_ANALYZER_CHECKS and the SCAN_BUILD_TOOL lookup from
cmake/ci.cmake, drop it from the ubuntu.yml ci_static_analysis_clang
matrix, and drop the now-unused clang-tools apt package (iwyu stays
for ci_single_binaries). Reword quality_assurance.md and
assurance_case.md, which described the dead job as a working control,
to say the Clang Static Analyzer checks run through clang-tidy.

Verified that `cmake -DJSON_CI=ON` still configures cleanly and that
ci_clang_analyze no longer appears in the generated build or in any
CMake/workflow file.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable portability-template-virtual-member-function; remove redundant forwards

.clang-tidy disabled three checks "to get the CI going" (#4489,
2024-11-13): portability-template-virtual-member-function,
bugprone-use-after-move and its alias hicpp-invalid-access-moved.

portability-template-virtual-member-function only flagged
output_stream_adapter::write_character/write_characters; annotate
both with NOLINT and re-enable the check.

bugprone-use-after-move flagged several double forwards that have no
effect at runtime:
- from_json.hpp calls std::forward<BasicJsonType>(j).at(Idx) inside
  pack expansions; at() has no ref-qualified overloads and always
  returns an lvalue reference, so the forward is a no-op. Replace with
  plain j.at(Idx) in all four places.
- the move constructor forwards the whole object to its base class
  and then reads other's members. That is item 9 of #5724 (together
  with its cppcheck suppressions) and is left to that change.
- input_adapters.hpp forwards the container twice on purpose, so the
  begin/end iterator types match adapter_type; annotate with NOLINT
  and a comment instead of changing behavior.

The check still flags the move constructor (see above) and two sites
in at(KeyType&&) (both overloads, json.hpp, in the throw's
string_t(std::forward<KeyType>(key)) after
find(std::forward<KeyType>(key))). Open PR #5689 rewrites that hunk,
so bugprone-use-after-move (and hicpp-invalid-access-moved)
stay disabled for now, with a comment explaining why; re-enable them
once #5689 and the #5724 move-constructor change have landed.

Also resolve the portability-avoid-pragma-once TODO: single_include
never has #pragma once (amalgamate.py strips it) and every supported
compiler accepts it in include/, so keep it disabled with an
explanatory comment instead of a TODO. Fix the stale "json.hpp,
around line 1265" comment in unit-class_parser.cpp, which now points
at the move constructor's actual line.

Behavior, the public API and the ABI do not change. Verified with
clang-tidy 22.1.8 that portability-template-virtual-member-function
now reports nothing, that bugprone-use-after-move/
hicpp-invalid-access-moved report only the known at(KeyType&&) and
move-constructor sites, and that unit-custom-base-class, unit-constructor1,
unit-conversions, unit-element_access2, unit-class_parser and
unit-diagnostic-positions (JSON_DIAGNOSTIC_POSITIONS=1) compile
under ASan/UBSan and pass with the same assertion counts as before.
Ran make amalgamate.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale and malformed NOLINT comments

json_sax.hpp named "-warnings-as-errors" in the NOLINT list on the two
JSON_ASSERT(false) lines; that is the suffix clang-tidy appends to a
diagnostic tag under WarningsAsErrors, not a check name, and every
other JSON_ASSERT(false) omits it.

unit-capacity.cpp carried 30 "// NOLINT(misc-const-correctness)"
comments on "json j = ...;" declarations that are all used with
non-const members afterwards, so the check has nothing to report
there.

unit-constructor2.cpp used a blanket "// NOLINT: access after move is
OK here" on a use-after-move that hides every check on the line;
naming bugprone-use-after-move and hicpp-invalid-access-moved keeps
the intent once those checks are re-enabled (#5724).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 10

* Remove stale .clang-tidy entries

-google-runtime-references disabled a check that neither clang-tidy
22.1.8 nor 23.1.2 lists under --list-checks -checks='*'; it was
removed upstream. The commented-out HeaderFilterRegex line has been
unused since the active HeaderFilterRegex was introduced in #2561
(2021).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 11

* Remove the GCC C++20 -Wignored-attributes pragma in json.hpp

The pragma (added in #5164) claimed to work around the C++ modules
redefinition errors of #5103, but #5103 is about hard errors (e.g.
"redefinition of std::__is_constant_evaluated()", conflicting
std::integral_constant) that ignoring a warning cannot suppress; they
are traced to GCC PR 124430 and reproduce with <map> or <string>
instead of json.hpp too. A GCC 16.2 -std=gnu++20 -fmodules build
following #5103's repro steps still fails with the pragma in place,
and a build of all test TUs with GCC_CXXFLAGS (which enable
-Wignored-attributes) and the pragma removed produces no such
warning. The block only hid a warning class from GCC C++20 users
while suggesting #5103 was handled.

Overlaps #5610, whose hunks touch the closing half of this pragma to
insert the json_literals.hpp include.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 9

* Fix stale doxygen comments hidden by the -Wdocumentation pragma

macro_scope.hpp ignores -Wdocumentation and -Wdocumentation-unknown-command
for the whole library, which also hides genuine documentation mistakes:

- detail::unescape() documented "@return unescaped string" but returns
  void and unescapes its argument in place; reworded to
  "@param[in,out] s string to unescape in place" and dropped the
  bogus @return.
- basic_json::get()'s copy-conversion overload wrote "converted to
  @tparam ValueType" inside @return, which Doxygen and Clang parse as
  a second, malformed @tparam; changed to "@a ValueType", matching the
  two other get() overloads a few lines above that already use it.

This narrows the gap the -Wdocumentation pragma needs to cover; fully
replacing the Doxygen-only commands it also hides (item 2c) is left
for after #5267.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 2

* Fix -Wextra-semi-stmt at its actual source, not assert()

clang_flags.cmake blamed the global -Wno-extra-semi-stmt on assert(),
but assert() expands to an expression under glibc and libc++ and does
not trigger this warning. unit-assert_macro.cpp overrides JSON_ASSERT
with "{if (!(x)) ++assert_counter; }", a bare block followed by a
semicolon at every JSON_ASSERT(...) call site in the library; that
was the actual source of 151 of the 208 -Wextra-semi-stmt sites found
in a Clang 22 -Weverything sweep of the test suite with the flag
removed. Switched to the standard do/while(false) macro idiom, which
does not expand to a statement-plus-semicolon, and corrected the
comment to name the remaining source instead: vendored Doctest's
CAPTURE(x) shim, which already ends in a semicolon.

Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that unit-assert_macro.cpp now compiles without any
-Wextra-semi-stmt diagnostic.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 8 (step 1 of 2; step 2 covers the CAPTURE() call sites)

* Drop the redundant semicolon from CAPTURE() call sites; remove -Wno-extra-semi-stmt

doctest_compatibility.h defines CAPTURE(x) as DOCTEST_CAPTURE(x); (with
a trailing semicolon baked into the macro), specifically so call sites
do not need to add one themselves; most of the ~267 call sites already
follow that convention. The remaining 64 call sites across 20 files
wrote "CAPTURE(x);" anyway, turning into a statement plus an empty
statement and triggering -Wextra-semi-stmt. Dropped the redundant
semicolon at each of those sites.

With item 6 having already made vendored Doctest a SYSTEM include, and
this the last known source of -Wextra-semi-stmt findings, removed the
flag from clang_flags.cmake entirely.

Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that all 20 touched files, plus a file with no CAPTURE() use
(unit-json_pointer.cpp), compile without any -Wextra-semi-stmt
diagnostic.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 8 (step 2 of 2)

* Switch ci_static_analysis_clang off the frozen LLVM 22 dev image

ubuntu.yml pinned the clang-tidy/clang-tidy-sanitizer/single-binaries
job to silkeh/clang:dev, a tag last pushed 2026-02-18 that reports
"clang version 22.0.0 (...+20251015...)", a pre-release snapshot from
before the LLVM 22 release; the maintainer now updates dev-unstable,
22, and latest instead. Switched to silkeh/clang:22, matching the
other clang jobs on :latest.

Verified with clang-tidy 22.1.8 (the image's actual version) against
this repository's .clang-tidy and library headers what the release
image newly reports compared to :dev:

- readability-redundant-typename fires at ~250 sites across the
  _cpp20-relevant conversion/to_chars headers; the library targets
  C++11 and keeps the typenames, so the check is disabled in
  .clang-tidy, matching how the file already handles checks that
  don't fit a C++11 codebase.
- misc-anonymous-namespace-in-header fires on the two anonymous
  namespaces in from_json.hpp and to_json.hpp; added the alias to
  their existing NOLINT (cert-dcl59-cpp, fuchsia-header-anon-namespaces,
  google-build-namespaces).
- bugprone-std-namespace-modification fires on every addition to
  namespace std: the std::hash, std::formatter and std::swap
  overloads in json.hpp, and the std::tuple_size/std::tuple_element
  specializations in iteration_proxy.hpp (this last file is not named
  in #5725's item 5, found by actually running clang-tidy 22.1.8
  against the current tree). All six are legal, deliberate additions
  to namespace std (explicit/partial specializations of std types, or
  the pre-C++20 std::swap overload); annotated each with the check
  name next to its existing cert-dcl58-cpp NOLINT.
- modernize-avoid-c-style-cast reported nothing new.

Also added clang++-22/21, clang-tidy-22/21, g++-16 and gcov-16 to the
find_program search lists in ci.cmake so a local "maximal warnings"
configure prefers the current toolchain version over an older one on
PATH.

#5725 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Regenerate cmake/gcc_flags.cmake for GCC 16.2.0

GCC_CXXFLAGS was generated for GCC 15.1.0, but ci_test_gcc and
ci_test_gcc_cxx{11..26} now run in gcc:latest, currently GCC 16.2.0,
so the "maximal warnings" job was missing warnings introduced since
15.1.0 while carrying entries GCC 16 treats as duplicates or no-ops.

Regenerated with https://github.com/nlohmann/gcc_flags (patched
locally to not crash on an option whose "-x c++ <opt> -" probe fails
before it reads stdin, e.g. -Wabi=; the tool otherwise raises
BrokenPipeError instead of recording the option as an error) run
against g++ 16.2.0 in the official gcc:16 Docker image, keeping the
documented -Wno-* exclusions and the same alphabetical placement
scheme as before.

Also added three GCC 16 warnings the generator cannot discover on its
own because it only probes value ranges/lists it finds in the -Q
option name itself, not in the enum choices --help=warnings documents
separately:

- -Wbidi-chars=any, -Wleading-whitespace=spaces: manually verified
  these compile cleanly with g++ 16.2.0.
- -Wstrict-flex-arrays: deliberately NOT added, unlike the other two.
  Without -fstrict-flex-arrays (which the library does not enable, as
  it would change codegen for flexible array members), GCC prints
  "'-Wstrict-flex-arrays' is ignored when '-fstrict-flex-arrays' is
  not present" on every translation unit, and under our -Werror that
  note itself aborts the build. This differs from the harmless
  no-op warnings already kept in the file (-Whsa, -Wsynth,
  -Wunreachable-code, -Wunsafe-loop-optimizations), which emit
  nothing; #5725 item 7 named -Wstrict-flex-arrays as one of the
  flags GCC 16 adds, but did not anticipate this failure mode.

Verified: compiled the library header and a representative set of
test translation units (including ones touched by items 1, 3, 8, 9,
10 of this issue) with the regenerated GCC_CXXFLAGS plus -Werror
under g++ 16.2.0 at -std=c++11 through -std=c++26, with zero warnings;
ran the full local test suite (129/129 passing, unrelated to this
compiler) as a regression check. CI must still confirm the actual
ci_test_gcc / ci_test_standards_gcc targets end to end, since this was
verified with direct g++ invocations rather than through the CMake/
CXXFLAGS environment-variable plumbing in ci.cmake.

#5725 item 7

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid std::basic_string<CharType> for non-character output_adapter CharType

output_adapter<CharType, StringType> defaulted StringType to
std::basic_string<CharType>, and (with JSON_NO_IO undefined) always
declared a std::basic_ostream<CharType>&-taking constructor. For
CharType with no non-deprecated std::char_traits specialization (only
std::uint8_t is ever used this way, by the binary writers), simply
naming either type - as an unused default template argument, or as an
unused, never-called constructor's parameter type - instantiates
std::char_traits<CharType> merely to name it, which some standard
libraries mark deprecated: with the library-wide -Wdocumentation
pragma (item 2's other half, left for a later commit) temporarily
removed, an Apple clang 21 / libc++ TU calling json::to_cbor(j, vec)
with std::vector<std::uint8_t>& got one -Wdeprecated-declarations
warning per binary writer at the old output_adapters.hpp:193.

Replaced the eager std::basic_string<CharType> / std::basic_ostream
<CharType> defaults with a bool-tagged partial specialization (not
std::conditional, which requires naming both branches' types up
front regardless of which is selected, reproducing the same warning)
that only ever names std::basic_string<CharType> / std::basic_ostream
<CharType> when CharType is actually one of char, wchar_t, char16_t,
char32_t, or (with __cpp_lib_char8_t) char8_t. For any other
CharType, output_adapter's StringType and ostream-constructor
parameter fall back to two distinct empty placeholder types, kept
distinct so the two constructor overloads do not collide into a
single redeclaration.

Public API / behavior: passing a std::basic_string<std::uint8_t>& or
std::basic_ostream<std::uint8_t>& directly to a binary writer's
output_adapter now fails to compile instead of compiling with a
deprecation warning; this was neither documented nor tested. All
documented uses (std::vector<CharType>, std::basic_ostream<CharType>
and StringType for character CharType) are unaffected.

Verified with Apple clang 21 / libc++, with the two -Wdocumentation*
"ignored" pragma lines in macro_scope.hpp temporarily removed and
-std=c++11/c++20 plus the project's -Weverything flag set: calling
to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson/to_bon8 on a
std::vector<std::uint8_t> now produces no char_traits<unsigned char>
(or any other) deprecation warning, while the char-based string- and
ostream-adapter paths, and a to_cbor/from_cbor round trip, still
compile and run correctly; also verified with GCC 16.2.0. Ran the
full local test suite, including the binary-format unit tests
(unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bson,
unit-bon8, unit-binary_writer_sinks, unit-binary_formats,
unit-custom-binary-type): 129/129 passing.

#5725 item 2 (step a)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the library-wide -Wdocumentation pragma; fix what it hid

macro_scope.hpp / macro_unscope.hpp pushed and popped a Clang
diagnostic region over the entire library that ignored -Wdocumentation
and -Wdocumentation-unknown-command. Removed both pragmas and fixed
every finding a full -Wdocumentation (which implies
-Wdocumentation-unknown-command and -Wdocumentation-deprecated-sync)
build reports, so the library now compiles clean under Clang's
documentation checks without a blanket suppression. Overlaps #5267,
which is still open and edits a nearby doc block (json.hpp's
get()/get_impl() @return, already fixed in the item 2 step (b) commit
of this branch); this commit does not touch that block again.

Unknown Doxygen alias commands (Doxyfile removed in #3071, so these
were never rendered by anything) rewritten as plain prose, keeping the
same information:
- @requirement REQ-JSON-01 / REQ-JSON-02 (iter_impl.hpp,
  json_reverse_iterator.hpp): now "This class satisfies the following
  concept requirements (REQ-JSON-0N):".
- @liveexample{prose,example-id} (three sites in json.hpp): kept the
  prose, dropped the command wrapper and the trailing example-id
  (docs/mkdocs/docs/examples/*.cpp still exist and are used directly
  by the rendered docs, not through this in-header alias) and
  unescaped the "\," commas that were only needed for the old alias's
  comma-separated argument syntax.
- @complexity X (json.hpp x4, json_pointer.hpp x2, serializer.hpp x1):
  now "Complexity: X".

Backslash sequences Clang's comment lexer tried to parse as commands,
escaped to render as literal backslashes:
- lexer.hpp get_codepoint(): two `\u` occurrences.
- binary_reader.hpp get_bson_cstr() / get_bson_cstr_bulk(): two
  `\x00` occurrences.
- serializer.hpp: three `\uXXXX` occurrences (constructor @param,
  append_codepoint_to_string_buffer() @brief, and the ensure_ascii
  member comment).

One finding remained after all of the above: Clang reports
"declaration is marked with '@deprecated' command but does not have a
deprecation attribute" on the deprecated sax_parse(span_input_adapter&&, ...)
overload, even though JSON_HEDLEY_DEPRECATED_FOR does expand to
__attribute__((deprecated(...))) for Clang. Several isolated
reproductions of this exact declaration shape - doc comment,
template<>, two stacked __attribute__ macros, an overload set sharing
the name - did not reproduce the warning, so this looks like a
Clang comment/declaration-association quirk specific to this overload
inside the much larger basic_json class template, not an actual
documentation defect. Rather than keep the pragma library-wide for one
Clang false positive, added a tightly scoped
-Wdocumentation-deprecated-sync push/pop around just that overload.

Verified with Apple clang 21 and the project's actual -Weverything
flag set (cmake/clang_flags.cmake) on the full header at -std=c++11
and -std=c++20: zero -Wdocumentation* diagnostics. Also compiled
clean with GCC 16.2.0 (the pragmas are already __clang__-gated, so
this only confirms no unrelated breakage). Ran make check-amalgamation
and the full local test suite: 129/129 passing.

#5725 item 2 (step c)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Take the JSON value by const reference in the array and tuple from_json paths

Review feedback on #5737 (gregmarr): once the no-op std::forward calls
are gone, the forwarding references have no purpose. from_json_fn
passes the value as const BasicJsonType&, so these functions were only
ever instantiated with a const lvalue anyway.

The std::array, std::pair and std::tuple overloads of from_json and
their helpers now take const BasicJsonType& and pass j on unchanged.
Because the deduced BasicJsonType is now the plain type, tuple_type and
the static_assert name const BasicJsonType& explicitly, so the
reference checks are unchanged: get<std::tuple<const std::string&>>()
still works, and get<std::tuple<std::string&>>() still fails the same
static_assert. from_json_tuple_get_impl keeps its forwarding reference,
since tuple_type calls it through std::declval.

Behavior, the public API and the ABI do not change. unit-conversions,
unit-constructor1, unit-udt, unit-udt_macro, unit-regression1/2/3,
unit-deserialization, unit-noexcept, unit-items, unit-allocator,
unit-custom-object-type, unit-ordered_json2 and
unit-brace-init-copy-semantics pass at C++11, C++17 and C++20 with
unchanged assertion counts. Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:37:47 +02:00

1897 lines
71 KiB
C++

// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <cstdint>
#include <fstream>
#include <limits>
#include <sstream>
#include <vector>
#include "make_test_data_available.hpp"
#include "round_trip_corpus.hpp"
#include "test_utils.hpp"
#include "sax_countdown.hpp"
using utils::SaxCountdown;
namespace
{
// a binary container that reports a size beyond INT32_MAX without allocating
// that much memory, so the BSON length overflow can be tested cheaply
class huge_binary_t : public std::vector<std::uint8_t>
{
public:
using std::vector<std::uint8_t>::vector;
size_type size() const noexcept // NOLINT(readability-convert-member-functions-to-static)
{
// one byte more than the BSON length field can represent
return static_cast<size_type>((std::numeric_limits<std::int32_t>::max)()) + 1;
}
};
using huge_binary_json = nlohmann::basic_json <
std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t,
double, std::allocator, nlohmann::adl_serializer, huge_binary_t, void >;
// a string type that can be made to report a size beyond INT32_MAX without
// allocating that much memory, so BSON length overflow can be tested for
// strings and (embedded) documents as well, following the same idea as
// huge_binary_t.
//
// Unlike huge_binary_t (which is only ever used as the BSON *value* type),
// this type doubles as basic_json's StringType and is therefore also used
// for *object keys* (e.g. "s" or "nested" below). Only the designated test
// value is meant to lie about its size - if every huge_string_t (including
// keys) reported a huge size, the running totals computed while walking the
// BSON document (see calc_bson_sizes in binary_writer.hpp)
// would need more than 32 bits, and on platforms where std::size_t is only
// 32 bits wide that arithmetic would silently wrap around, producing wrong
// (or even unguarded) lengths. The fake size is therefore opt-in via
// as_huge(), and plain strings - in particular object keys - keep reporting
// their real, small size.
class huge_string_t : public std::string
{
public:
using std::string::string;
huge_string_t(const std::string& s) : std::string(s) {} // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
// returns a copy of @a s whose size() pretends to be huge
static huge_string_t as_huge(const std::string& s)
{
huge_string_t result(s);
result.pretend_huge = true;
return result;
}
size_type size() const noexcept
{
if (pretend_huge)
{
// one byte more than the BSON length field can represent
return static_cast<size_type>((std::numeric_limits<std::int32_t>::max)()) + 1;
}
return std::string::size();
}
private:
bool pretend_huge = false;
};
using huge_string_json = nlohmann::basic_json <
std::map, std::vector, huge_string_t, bool, std::int64_t, std::uint64_t,
double, std::allocator, nlohmann::adl_serializer, std::vector<std::uint8_t>, void >;
} // namespace
TEST_CASE("BSON")
{
SECTION("individual values not supported")
{
SECTION("null")
{
json const j = nullptr;
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is null", json::type_error&);
}
SECTION("boolean")
{
SECTION("true")
{
json const j = true;
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is boolean", json::type_error&);
}
SECTION("false")
{
json const j = false;
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is boolean", json::type_error&);
}
}
SECTION("number")
{
json const j = 42;
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is number", json::type_error&);
}
SECTION("float")
{
json const j = 4.2;
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is number", json::type_error&);
}
SECTION("string")
{
json const j = "not supported";
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is string", json::type_error&);
}
SECTION("array")
{
json const j = std::vector<int> {1, 2, 3, 4, 5, 6, 7};
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.type_error.317] to serialize to BSON, top-level type must be object, but is array", json::type_error&);
}
}
SECTION("keys containing code-point U+0000 cannot be serialized to BSON")
{
json const j =
{
{ std::string("en\0try", 6), true }
};
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.out_of_range.409] (/en) BSON key cannot contain code point U+0000 (at byte 2)", json::out_of_range&);
#else
CHECK_THROWS_WITH_AS(json::to_bson(j), "[json.exception.out_of_range.409] BSON key cannot contain code point U+0000 (at byte 2)", json::out_of_range&);
#endif
}
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
{
// out_of_range.412 is thrown from a single shared helper
// (to_bson_length) that guards the BSON length fields of binary
// values, strings, and (embedded) documents alike
SECTION("binary")
{
huge_binary_json j;
j["b"] = huge_binary_json::binary(huge_binary_t{});
CHECK_THROWS_WITH_AS(huge_binary_json::to_bson(j), "[json.exception.out_of_range.412] BSON length 2147483661 exceeds maximum of 2147483647", huge_binary_json::out_of_range&);
}
SECTION("string")
{
huge_string_json j;
j["s"] = huge_string_t::as_huge("value");
CHECK_THROWS_WITH_AS(huge_string_json::to_bson(j), "[json.exception.out_of_range.412] BSON length 2147483661 exceeds maximum of 2147483647", huge_string_json::out_of_range&);
}
SECTION("document")
{
// an oversized string nested one level deep makes the
// *embedded* document's own length exceed INT32_MAX as well
huge_string_json nested;
nested["s"] = huge_string_t::as_huge("value");
huge_string_json j;
j["nested"] = nested;
CHECK_THROWS_WITH_AS(huge_string_json::to_bson(j), "[json.exception.out_of_range.412] BSON length 2147483674 exceeds maximum of 2147483647", huge_string_json::out_of_range&);
}
}
SECTION("string length must be at least 1")
{
// from https://bugs.chromium.org/p/oss-fuzz/issues/detail?id=11175
std::vector<std::uint8_t> const v =
{
0x20, 0x20, 0x20, 0x20,
0x02,
0x00,
0x00, 0x00, 0x00, 0x80
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 10: syntax error while parsing BSON string: string length must be at least 1, is -2147483648", json::parse_error&);
}
SECTION("objects")
{
SECTION("empty object")
{
json const j = json::object();
std::vector<std::uint8_t> const expected =
{
0x05, 0x00, 0x00, 0x00, // size (little endian)
// no entries
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with bool")
{
json const j =
{
{ "entry", true }
};
std::vector<std::uint8_t> const expected =
{
0x0D, 0x00, 0x00, 0x00, // size (little endian)
0x08, // entry: boolean
'e', 'n', 't', 'r', 'y', '\x00',
0x01, // value = true
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with bool")
{
json const j =
{
{ "entry", false }
};
std::vector<std::uint8_t> const expected =
{
0x0D, 0x00, 0x00, 0x00, // size (little endian)
0x08, // entry: boolean
'e', 'n', 't', 'r', 'y', '\x00',
0x00, // value = false
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with bool from a non-0/1 byte (lenient parsing)")
{
// documented lenient behavior (see gh-5333): any non-zero byte
// is accepted as `true`, not just 0x01
std::vector<std::uint8_t> const input =
{
0x0D, 0x00, 0x00, 0x00, // size (little endian)
0x08, // entry: boolean
'e', 'n', 't', 'r', 'y', '\x00',
0x02, // value = 0x02 (neither 0x00 nor 0x01)
0x00 // end marker
};
const json expected = { { "entry", true } };
CHECK(json::from_bson(input) == expected);
}
SECTION("non-empty object with double")
{
json const j =
{
{ "entry", 4.2 }
};
std::vector<std::uint8_t> const expected =
{
0x14, 0x00, 0x00, 0x00, // size (little endian)
0x01, /// entry: double
'e', 'n', 't', 'r', 'y', '\x00',
0xcd, 0xcc, 0xcc, 0xcc, 0xcc, 0xcc, 0x10, 0x40,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with string")
{
json const j =
{
{ "entry", "bsonstr" }
};
std::vector<std::uint8_t> const expected =
{
0x18, 0x00, 0x00, 0x00, // size (little endian)
0x02, /// entry: string (UTF-8)
'e', 'n', 't', 'r', 'y', '\x00',
0x08, 0x00, 0x00, 0x00, 'b', 's', 'o', 'n', 's', 't', 'r', '\x00',
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with null member")
{
json const j =
{
{ "entry", nullptr }
};
std::vector<std::uint8_t> const expected =
{
0x0C, 0x00, 0x00, 0x00, // size (little endian)
0x0A, /// entry: null
'e', 'n', 't', 'r', 'y', '\x00',
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with integer (32-bit) member")
{
json const j =
{
{ "entry", std::int32_t{0x12345678} }
};
std::vector<std::uint8_t> const expected =
{
0x10, 0x00, 0x00, 0x00, // size (little endian)
0x10, /// entry: int32
'e', 'n', 't', 'r', 'y', '\x00',
0x78, 0x56, 0x34, 0x12,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with integer (64-bit) member")
{
json const j =
{
{ "entry", std::int64_t{0x1234567804030201} }
};
std::vector<std::uint8_t> const expected =
{
0x14, 0x00, 0x00, 0x00, // size (little endian)
0x12, /// entry: int64
'e', 'n', 't', 'r', 'y', '\x00',
0x01, 0x02, 0x03, 0x04, 0x78, 0x56, 0x34, 0x12,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with negative integer (32-bit) member")
{
json const j =
{
{ "entry", std::int32_t{-1} }
};
std::vector<std::uint8_t> const expected =
{
0x10, 0x00, 0x00, 0x00, // size (little endian)
0x10, /// entry: int32
'e', 'n', 't', 'r', 'y', '\x00',
0xFF, 0xFF, 0xFF, 0xFF,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with negative integer (64-bit) member")
{
json const j =
{
{ "entry", std::int64_t{-1} }
};
std::vector<std::uint8_t> const expected =
{
0x10, 0x00, 0x00, 0x00, // size (little endian)
0x10, /// entry: int32
'e', 'n', 't', 'r', 'y', '\x00',
0xFF, 0xFF, 0xFF, 0xFF,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with unsigned integer (64-bit) member")
{
json const j =
{
{ "entry", std::uint64_t{0x1234567804030201} }
};
std::vector<std::uint8_t> const expected =
{
0x14, 0x00, 0x00, 0x00, // size (little endian)
0x12, /// entry: int64
'e', 'n', 't', 'r', 'y', '\x00',
0x01, 0x02, 0x03, 0x04, 0x78, 0x56, 0x34, 0x12,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with small unsigned integer member")
{
json const j =
{
{ "entry", std::uint64_t{0x42} }
};
std::vector<std::uint8_t> const expected =
{
0x10, 0x00, 0x00, 0x00, // size (little endian)
0x10, /// entry: int32
'e', 'n', 't', 'r', 'y', '\x00',
0x42, 0x00, 0x00, 0x00,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with object member")
{
json const j =
{
{ "entry", json::object() }
};
std::vector<std::uint8_t> const expected =
{
0x11, 0x00, 0x00, 0x00, // size (little endian)
0x03, /// entry: embedded document
'e', 'n', 't', 'r', 'y', '\x00',
0x05, 0x00, 0x00, 0x00, // size (little endian)
// no entries
0x00, // end marker (embedded document)
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with array member")
{
json const j =
{
{ "entry", json::array() }
};
std::vector<std::uint8_t> const expected =
{
0x11, 0x00, 0x00, 0x00, // size (little endian)
0x04, /// entry: embedded document
'e', 'n', 't', 'r', 'y', '\x00',
0x05, 0x00, 0x00, 0x00, // size (little endian)
// no entries
0x00, // end marker (embedded document)
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with non-empty array member")
{
json const j =
{
{ "entry", json::array({1, 2, 3, 4, 5, 6, 7, 8}) }
};
std::vector<std::uint8_t> const expected =
{
0x49, 0x00, 0x00, 0x00, // size (little endian)
0x04, /// entry: embedded document
'e', 'n', 't', 'r', 'y', '\x00',
0x3D, 0x00, 0x00, 0x00, // size (little endian)
0x10, '0', 0x00, 0x01, 0x00, 0x00, 0x00,
0x10, '1', 0x00, 0x02, 0x00, 0x00, 0x00,
0x10, '2', 0x00, 0x03, 0x00, 0x00, 0x00,
0x10, '3', 0x00, 0x04, 0x00, 0x00, 0x00,
0x10, '4', 0x00, 0x05, 0x00, 0x00, 0x00,
0x10, '5', 0x00, 0x06, 0x00, 0x00, 0x00,
0x10, '6', 0x00, 0x07, 0x00, 0x00, 0x00,
0x10, '7', 0x00, 0x08, 0x00, 0x00, 0x00,
0x00, // end marker (embedded document)
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("array elements with non-conforming keys (lenient parsing)")
{
// documented lenient behavior (see gh-5333): BSON array element
// keys are not checked against the required decimal sequence
// "0", "1", "2", ... - elements are taken in encoded order
std::vector<std::uint8_t> const input =
{
0x26, 0x00, 0x00, 0x00, // size (little endian)
0x04, 'e', 'n', 't', 'r', 'y', '\x00', // entry: embedded array
0x1A, 0x00, 0x00, 0x00, // size (little endian)
0x10, '5', 0x00, 0x0A, 0x00, 0x00, 0x00, // key "5" (bogus) -> 10
0x10, 'x', 0x00, 0x14, 0x00, 0x00, 0x00, // key "x" (non-numeric) -> 20
0x10, '1', 0x00, 0x1E, 0x00, 0x00, 0x00, // key "1" (out of order) -> 30
0x00, // end marker (embedded array)
0x00 // end marker
};
const json expected = { { "entry", json::array({10, 20, 30}) } };
CHECK(json::from_bson(input) == expected);
}
SECTION("non-empty object with binary member")
{
const size_t N = 10;
const auto s = std::vector<std::uint8_t>(N, 'x');
json const j =
{
{ "entry", json::binary(s, 0) }
};
std::vector<std::uint8_t> const expected =
{
0x1B, 0x00, 0x00, 0x00, // size (little endian)
0x05, // entry: binary
'e', 'n', 't', 'r', 'y', '\x00',
0x0A, 0x00, 0x00, 0x00, // size of binary (little endian)
0x00, // Generic binary subtype
0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("non-empty object with binary member without subtype")
{
const size_t N = 10;
const auto s = std::vector<std::uint8_t>(N, 'x');
json const j =
{
{ "entry", json::binary(s) }
};
CHECK(!j.at("entry").get_binary().has_subtype());
std::vector<std::uint8_t> const expected =
{
0x1B, 0x00, 0x00, 0x00, // size (little endian)
0x05, // entry: binary
'e', 'n', 't', 'r', 'y', '\x00',
0x0A, 0x00, 0x00, 0x00, // size of binary (little endian)
0x00, // Generic binary subtype
0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78, 0x78,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip adds the generic binary subtype
const auto roundtrip = json::from_bson(result);
CHECK(roundtrip != j);
CHECK(roundtrip.at("entry").get_binary().has_subtype());
CHECK(roundtrip.at("entry").get_binary().subtype() == 0);
CHECK(json::from_bson(result, true, false) == roundtrip);
}
SECTION("non-empty object with binary member with subtype")
{
// an MD5 hash
const std::vector<std::uint8_t> md5hash = {0xd7, 0x7e, 0x27, 0x54, 0xbe, 0x12, 0x37, 0xfe, 0xd6, 0x0c, 0x33, 0x98, 0x30, 0x3b, 0x8d, 0xc4};
json const j =
{
{ "entry", json::binary(md5hash, 5) }
};
std::vector<std::uint8_t> const expected =
{
0x21, 0x00, 0x00, 0x00, // size (little endian)
0x05, // entry: binary
'e', 'n', 't', 'r', 'y', '\x00',
0x10, 0x00, 0x00, 0x00, // size of binary (little endian)
0x05, // MD5 binary subtype
0xd7, 0x7e, 0x27, 0x54, 0xbe, 0x12, 0x37, 0xfe, 0xd6, 0x0c, 0x33, 0x98, 0x30, 0x3b, 0x8d, 0xc4,
0x00 // end marker
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
SECTION("binary member with subtype 0x02 (old binary) keeps its inner length prefix (lenient parsing)")
{
// documented lenient behavior (see gh-5333): the payload for
// binary subtype 0x02 ("old binary") is returned as-is,
// including its own inner 4-byte length prefix; it is not
// stripped or reinterpreted
std::vector<std::uint8_t> const input =
{
0x17, 0x00, 0x00, 0x00, // size (little endian)
0x05, 'e', 'n', 't', 'r', 'y', '\x00', // entry: binary
0x06, 0x00, 0x00, 0x00, // size of binary (little endian)
0x02, // "old binary" subtype
0x02, 0x00, 0x00, 0x00, // inner length prefix (part of the old-binary payload)
0x68, 0x69, // payload ('h', 'i')
0x00 // end marker
};
// the inner length prefix is part of the (unmodified) payload
const std::vector<std::uint8_t> expected_payload = {0x02, 0x00, 0x00, 0x00, 0x68, 0x69};
const json expected = { { "entry", json::binary(expected_payload, 0x02) } };
CHECK(json::from_bson(input) == expected);
}
SECTION("Some more complex document")
{
json const j =
{
{"double", 42.5},
{"entry", 4.2},
{"number", 12345},
{"object", {{ "string", "value" }}}
};
std::vector<std::uint8_t> const expected =
{
/*size */ 0x4f, 0x00, 0x00, 0x00,
/*entry*/ 0x01, 'd', 'o', 'u', 'b', 'l', 'e', 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x40, 0x45, 0x40,
/*entry*/ 0x01, 'e', 'n', 't', 'r', 'y', 0x00, 0xcd, 0xcc, 0xcc, 0xcc, 0xcc, 0xcc, 0x10, 0x40,
/*entry*/ 0x10, 'n', 'u', 'm', 'b', 'e', 'r', 0x00, 0x39, 0x30, 0x00, 0x00,
/*entry*/ 0x03, 'o', 'b', 'j', 'e', 'c', 't', 0x00,
/*entry: obj-size */ 0x17, 0x00, 0x00, 0x00,
/*entry: obj-entry*/0x02, 's', 't', 'r', 'i', 'n', 'g', 0x00, 0x06, 0x00, 0x00, 0x00, 'v', 'a', 'l', 'u', 'e', 0,
/*entry: obj-term.*/0x00,
/*obj-term*/ 0x00
};
const auto result = json::to_bson(j);
CHECK(result == expected);
// roundtrip
CHECK(json::from_bson(result) == j);
CHECK(json::from_bson(result, true, false) == j);
}
}
SECTION("Examples from https://bsonspec.org/faq.html")
{
SECTION("Example 1")
{
std::vector<std::uint8_t> input = {0x16, 0x00, 0x00, 0x00, 0x02, 'h', 'e', 'l', 'l', 'o', 0x00, 0x06, 0x00, 0x00, 0x00, 'w', 'o', 'r', 'l', 'd', 0x00, 0x00};
const json parsed = json::from_bson(input);
const json expected = {{"hello", "world"}};
CHECK(parsed == expected);
const auto dumped = json::to_bson(parsed);
CHECK(dumped == input);
CHECK(json::from_bson(dumped) == expected);
}
SECTION("Example 2")
{
std::vector<std::uint8_t> input = {0x31, 0x00, 0x00, 0x00, 0x04, 'B', 'S', 'O', 'N', 0x00, 0x26, 0x00, 0x00, 0x00, 0x02, 0x30, 0x00, 0x08, 0x00, 0x00, 0x00, 'a', 'w', 'e', 's', 'o', 'm', 'e', 0x00, 0x01, 0x31, 0x00, 0x33, 0x33, 0x33, 0x33, 0x33, 0x33, 0x14, 0x40, 0x10, 0x32, 0x00, 0xc2, 0x07, 0x00, 0x00, 0x00, 0x00};
const json parsed = json::from_bson(input);
const json expected = {{"BSON", {"awesome", 5.05, 1986}}};
CHECK(parsed == expected);
const auto dumped = json::to_bson(parsed);
CHECK(dumped == input);
CHECK(json::from_bson(dumped) == expected);
}
}
}
TEST_CASE("regression test - BSON binary subtype rejects a value that doesn't fit a single byte")
{
json const doc255 = {{"b", json::binary({1, 2}, 255)}};
CHECK(json::from_bson(json::to_bson(doc255))["b"].get_binary().subtype() == 255);
CHECK_THROWS_AS(json::to_bson(json{{"b", json::binary({1, 2}, 256)}}), json::out_of_range);
#if JSON_DIAGNOSTICS
CHECK_THROWS_WITH_AS(json::to_bson(json {{"b", json::binary({1, 2}, 300)}}), "[json.exception.out_of_range.415] (/b) subtype 300 is too large for the BSON binary subtype (max 255)", json::out_of_range);
#else
CHECK_THROWS_WITH_AS(json::to_bson(json {{"b", json::binary({1, 2}, 300)}}), "[json.exception.out_of_range.415] subtype 300 is too large for the BSON binary subtype (max 255)", json::out_of_range);
#endif
}
TEST_CASE("BSON input/output_adapters")
{
const json json_representation =
{
{"double", 42.5},
{"entry", 4.2},
{"number", 12345},
{"object", {{ "string", "value" }}}
};
const std::vector<std::uint8_t> bson_representation =
{
/*size */ 0x4f, 0x00, 0x00, 0x00,
/*entry*/ 0x01, 'd', 'o', 'u', 'b', 'l', 'e', 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x40, 0x45, 0x40,
/*entry*/ 0x01, 'e', 'n', 't', 'r', 'y', 0x00, 0xcd, 0xcc, 0xcc, 0xcc, 0xcc, 0xcc, 0x10, 0x40,
/*entry*/ 0x10, 'n', 'u', 'm', 'b', 'e', 'r', 0x00, 0x39, 0x30, 0x00, 0x00,
/*entry*/ 0x03, 'o', 'b', 'j', 'e', 'c', 't', 0x00,
/*entry: obj-size */ 0x17, 0x00, 0x00, 0x00,
/*entry: obj-entry*/0x02, 's', 't', 'r', 'i', 'n', 'g', 0x00, 0x06, 0x00, 0x00, 0x00, 'v', 'a', 'l', 'u', 'e', 0,
/*entry: obj-term.*/0x00,
/*obj-term*/ 0x00
};
json j2;
CHECK_NOTHROW(j2 = json::from_bson(bson_representation));
// compare parsed JSON values
CHECK(json_representation == j2);
SECTION("roundtrips")
{
SECTION("std::ostringstream")
{
std::basic_ostringstream<char> ss;
json::to_bson(json_representation, ss);
const json j3 = json::from_bson(ss.str());
CHECK(json_representation == j3);
}
SECTION("std::string")
{
std::string s;
json::to_bson(json_representation, s);
const json j3 = json::from_bson(s);
CHECK(json_representation == j3);
}
SECTION("std::vector")
{
std::vector<std::uint8_t> v;
json::to_bson(json_representation, v);
const json j3 = json::from_bson(v);
CHECK(json_representation == j3);
}
}
}
TEST_CASE("Incomplete BSON Input")
{
SECTION("Incomplete BSON Input 1")
{
std::vector<std::uint8_t> const incomplete_bson =
{
0x0D, 0x00, 0x00, 0x00, // size (little endian)
0x08, // entry: boolean
'e', 'n', 't' // unexpected EOF
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 9: syntax error while parsing BSON cstring: unexpected end of input", json::parse_error&);
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
SaxCountdown scp(0);
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
}
SECTION("Incomplete BSON Input 2")
{
std::vector<std::uint8_t> const incomplete_bson =
{
0x0D, 0x00, 0x00, 0x00, // size (little endian)
0x08, // entry: boolean, unexpected EOF
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 6: syntax error while parsing BSON cstring: unexpected end of input", json::parse_error&);
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
SaxCountdown scp(0);
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
}
SECTION("Incomplete BSON Input 3")
{
std::vector<std::uint8_t> const incomplete_bson =
{
0x41, 0x00, 0x00, 0x00, // size (little endian)
0x04, /// entry: embedded document
'e', 'n', 't', 'r', 'y', '\x00',
0x35, 0x00, 0x00, 0x00, // size (little endian)
0x10, 0x00, 0x01, 0x00, 0x00, 0x00,
0x10, 0x00, 0x02, 0x00, 0x00, 0x00
// missing input data...
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 28: syntax error while parsing BSON element list: unexpected end of input", json::parse_error&);
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
SaxCountdown scp(1);
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
}
SECTION("Incomplete BSON Input 4")
{
std::vector<std::uint8_t> const incomplete_bson =
{
0x0D, 0x00, // size (incomplete), unexpected EOF
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 3: syntax error while parsing BSON number: unexpected end of input", json::parse_error&);
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
SaxCountdown scp(0);
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
}
SECTION("Incomplete BSON Input 5")
{
std::vector<std::uint8_t> const incomplete_bson =
{
0x09, 0x00, 0x00, 0x00, // size (little endian)
0x08, // entry: boolean
'b', '\x00' // key, unexpected EOF before the value
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 8: syntax error while parsing BSON number: unexpected end of input", json::parse_error&);
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
SaxCountdown scp(0);
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
}
SECTION("Incomplete BSON Input 6")
{
std::vector<std::uint8_t> const incomplete_bson =
{
0x0F, 0x00, 0x00, 0x00, // size (little endian)
0x05, // entry: binary
'b', '\x00', // key
0x00, 0x00, 0x00, 0x00 // length, unexpected EOF before the subtype
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(incomplete_bson), "[json.exception.parse_error.110] parse error at byte 12: syntax error while parsing BSON number: unexpected end of input", json::parse_error&);
CHECK(json::from_bson(incomplete_bson, true, false).is_discarded());
SaxCountdown scp(0);
CHECK(!json::sax_parse(incomplete_bson, &scp, json::input_format_t::bson));
}
SECTION("Improve coverage")
{
SECTION("key")
{
json const j = {{"key", "value"}};
auto bson_vec = json::to_bson(j);
SaxCountdown scp(2);
CHECK(!json::sax_parse(bson_vec, &scp, json::input_format_t::bson));
}
SECTION("array")
{
json const j =
{
{ "entry", json::array() }
};
auto bson_vec = json::to_bson(j);
SaxCountdown scp(2);
CHECK(!json::sax_parse(bson_vec, &scp, json::input_format_t::bson));
}
}
}
// the test catches the exceptions of invalid input
#if !defined(JSON_NOEXCEPTION)
TEST_CASE("BSON keys from contiguous and stream input")
{
// contiguous input reads a key up to its \x00-byte in one step, a stream
// reads it byte by byte; both must give the same value or error for the
// complete document and for every truncation of it
const json j = {{"", true}, {"k", {1, 2, 3}}, {std::string(40, 'x'), {{"nested key", "value"}}}};
const std::vector<std::uint8_t> bson = json::to_bson(j);
CHECK(json::from_bson(bson) == j);
for (std::size_t length = 0; length <= bson.size(); ++length)
{
CAPTURE(length)
const std::vector<std::uint8_t> input(bson.begin(), bson.begin() + static_cast<std::ptrdiff_t>(length));
std::string from_vector;
std::string from_stream;
try
{
from_vector = json::from_bson(input).dump();
}
catch (const json::parse_error& e)
{
from_vector = e.what();
}
try
{
std::istringstream stream(std::string(input.begin(), input.end()));
from_stream = json::from_bson(stream).dump();
}
catch (const json::parse_error& e)
{
from_stream = e.what();
}
CHECK(from_vector == from_stream);
}
}
#endif
TEST_CASE("Negative size of binary value")
{
// invalid BSON: the size of the binary value is -1
std::vector<std::uint8_t> const input =
{
0x21, 0x00, 0x00, 0x00, // size (little endian)
0x05, // entry: binary
'e', 'n', 't', 'r', 'y', '\x00',
0xFF, 0xFF, 0xFF, 0xFF, // size of binary (little endian)
0x05, // MD5 binary subtype
0xd7, 0x7e, 0x27, 0x54, 0xbe, 0x12, 0x37, 0xfe, 0xd6, 0x0c, 0x33, 0x98, 0x30, 0x3b, 0x8d, 0xc4,
0x00 // end marker
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1", json::parse_error);
}
TEST_CASE("Unsupported BSON input")
{
std::vector<std::uint8_t> const bson =
{
0x0C, 0x00, 0x00, 0x00, // size (little endian)
0xFF, // entry type: Min key (not supported yet)
'e', 'n', 't', 'r', 'y', '\x00',
0x00 // end marker
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(bson), "[json.exception.parse_error.114] parse error at byte 5: Unsupported BSON record type 0xFF", json::parse_error&);
CHECK(json::from_bson(bson, true, false).is_discarded());
SaxCountdown scp(0);
CHECK(!json::sax_parse(bson, &scp, json::input_format_t::bson));
}
TEST_CASE("BSON document size mismatch")
{
json _;
SECTION("top-level document declaring more bytes than it contains")
{
// empty object, but the length prefix claims 6 bytes instead of 5
std::vector<std::uint8_t> const input = {0x06, 0x00, 0x00, 0x00, 0x00};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("top-level document with a negative size")
{
std::vector<std::uint8_t> const input = {0xFF, 0xFF, 0xFF, 0xFF, 0x00};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size -1 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("embedded document whose size disagrees with its terminator")
{
// the embedded document "d" declares 0x7FFFFFFF bytes but its 0x00
// terminator falls right after {"a":null}; the length prefix would
// otherwise let the following "h" element be read as a member of the
// enclosing document instead of "d"
std::vector<std::uint8_t> const input =
{
0x00, 0x00, 0x00, 0x00, // outer size
0x03, 'd', 0x00, // entry: embedded document "d"
0xFF, 0xFF, 0xFF, 0x7F, // embedded size 0x7FFFFFFF
0x0A, 'a', 0x00, // entry: null "a"
0x00, // embedded end marker
0x08, 'h', 0x00, 0x01, // entry: bool "h" = true
0x00 // outer end marker
};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON document: document size 2147483647 does not match the number of bytes read (8)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("embedded array whose size disagrees with its terminator")
{
// array [42] is 12 bytes, but the length prefix claims 13
std::vector<std::uint8_t> const input =
{
0x00, 0x00, 0x00, 0x00, // outer size
0x04, 'a', 0x00, // entry: array "a"
0x0D, 0x00, 0x00, 0x00, // array size 13 (real is 12)
0x10, '0', 0x00, 0x2A, 0x00, 0x00, 0x00, // entry: int32 "0" = 42
0x00, // array end marker
0x00 // outer end marker
};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 19: syntax error while parsing BSON document: document size 13 does not match the number of bytes read (12)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
}
TEST_CASE("BSON nesting does not consume the call stack")
{
// An embedded document or array used to be read by calling back into the
// document reader, so the native call stack grew with the nesting depth of
// the input (#5104). The open documents are kept on a heap stack now.
//
// Deeply nested values must not be compared, copied or dumped here: those
// operations are still recursive and would reintroduce the crash.
// A document nested deeply enough to have crashed. The bytes are built
// here rather than with to_bson(), because the writer still recurses once
// per level and would overflow the stack before the reader is ever
// reached. Every level is
// <int32 size> 0x03 'a' 0x00 <inner document> 0x00
// so a level is eight bytes larger than the one it holds, and the sizes
// can be filled in from the outside in.
const std::size_t depth = 30000;
std::vector<uint8_t> input;
input.reserve(5 + (8 * depth));
for (std::size_t i = 0; i < depth; ++i)
{
const auto size = static_cast<std::uint32_t>(5 + (8 * (depth - i)));
input.push_back(static_cast<uint8_t>(size & 0xFF));
input.push_back(static_cast<uint8_t>((size >> 8) & 0xFF));
input.push_back(static_cast<uint8_t>((size >> 16) & 0xFF));
input.push_back(static_cast<uint8_t>((size >> 24) & 0xFF));
input.push_back(0x03); // embedded document
input.push_back('a');
input.push_back(0x00);
}
// the innermost document is empty, then one terminator closes each level
input.insert(input.end(), {0x05, 0x00, 0x00, 0x00, 0x00});
input.insert(input.end(), depth, 0x00);
SECTION("a well-formed deep document is read through the SAX interface")
{
SaxCountdown accept_all(1000000);
CHECK(json::sax_parse(input, &accept_all, json::input_format_t::bson));
}
SECTION("a well-formed deep document is read into a value")
{
json j = json::from_bson(input);
// walked rather than compared: comparing, copying or dumping a value
// this deep is still recursive
std::size_t measured = 0;
const json* q = &j;
while (q->is_object() && !q->empty())
{
q = &q->begin().value();
++measured;
}
CHECK(measured == depth);
}
SECTION("embedded documents and arrays are still read the same way")
{
const json values = {{"a", {{"b", {{"c", 1}}}}}};
CHECK(json::from_bson(json::to_bson(values)) == values);
const json array = {{"a", {1, 2, 3}}};
CHECK(json::from_bson(json::to_bson(array)) == array);
const json mixed = {{"a", {json{{"x", 1}}, json{{"y", 2}}}}};
CHECK(json::from_bson(json::to_bson(mixed)) == mixed);
CHECK(json::from_bson(json::to_bson(json::object())) == json::object());
}
SECTION("a size that does not match is still reported per document")
{
// the embedded document claims one byte too many
std::vector<uint8_t> const bad =
{
0x15, 0x00, 0x00, 0x00, 0x03, 'a', 0x00,
0x0D, 0x00, 0x00, 0x00, 0x08, 'b', 0x00, 0x01, 0x00,
0x00
};
json _;
CHECK_THROWS_AS(_ = json::from_bson(bad), json::parse_error&);
CHECK(json::from_bson(bad, true, false).is_discarded());
}
}
TEST_CASE("BSON input that cannot be read is discarded by every overload")
{
std::vector<std::uint8_t> input = json::to_bson(json({{"a", {1, 2}}}));
input.pop_back();
json _;
CHECK_THROWS_AS(_ = json::from_bson(input.begin(), input.end()), json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
CHECK(json::from_bson(input.begin(), input.end(), true, false).is_discarded());
CHECK(json::from_bson(input.data(), input.size(), true, false).is_discarded());
CHECK(json::from_bson({input.data(), input.size()}, true, false).is_discarded());
}
TEST_CASE("BSON SAX parsing stops at every event")
{
// Containers are opened and closed by the loop that reads them; a SAX
// handler that rejects any event - including the end of a nested
// container - must stop the parse right there.
const auto count_events = [](const std::vector<std::uint8_t>& input)
{
int events = 0;
while (true)
{
SaxCountdown scp(events);
if (json::sax_parse(input, &scp, json::input_format_t::bson))
{
return events;
}
++events;
REQUIRE(events < 1000);
}
};
// 20 events: every container kind closes inside another one
const json j = json::parse(R"({"a": [1, {"b": []}], "c": {"d": [[2]]}})");
CHECK(count_events(json::to_bson(j)) == 20);
}
TEST_CASE("BSON numerical data")
{
SECTION("number")
{
SECTION("signed")
{
SECTION("std::int64_t: INT64_MIN .. INT32_MIN-1")
{
std::vector<int64_t> const numbers
{
(std::numeric_limits<int64_t>::min)(),
-1000000000000000000LL,
-100000000000000000LL,
-10000000000000000LL,
-1000000000000000LL,
-100000000000000LL,
-10000000000000LL,
-1000000000000LL,
-100000000000LL,
-10000000000LL,
static_cast<std::int64_t>((std::numeric_limits<std::int32_t>::min)()) - 1,
};
for (const auto i : numbers)
{
CAPTURE(i)
json const j =
{
{ "entry", i }
};
CHECK(j.at("entry").is_number_integer());
std::uint64_t const iu = *reinterpret_cast<const std::uint64_t*>(&i);
std::vector<std::uint8_t> const expected_bson =
{
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
0x12u, /// entry: int64
'e', 'n', 't', 'r', 'y', '\x00',
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
0x00u // end marker
};
const auto bson = json::to_bson(j);
CHECK(bson == expected_bson);
auto j_roundtrip = json::from_bson(bson);
CHECK(j_roundtrip.at("entry").is_number_integer());
CHECK(j_roundtrip == j);
CHECK(json::from_bson(bson, true, false) == j);
}
}
SECTION("signed std::int32_t: INT32_MIN .. INT32_MAX")
{
std::vector<int32_t> const numbers
{
(std::numeric_limits<int32_t>::min)(),
-2147483647L,
-1000000000L,
-100000000L,
-10000000L,
-1000000L,
-100000L,
-10000L,
-1000L,
-100L,
-10L,
-1L,
0L,
1L,
10L,
100L,
1000L,
10000L,
100000L,
1000000L,
10000000L,
100000000L,
1000000000L,
2147483646L,
(std::numeric_limits<int32_t>::max)()
};
for (const auto i : numbers)
{
CAPTURE(i)
json const j =
{
{ "entry", i }
};
CHECK(j.at("entry").is_number_integer());
std::uint32_t const iu = *reinterpret_cast<const std::uint32_t*>(&i);
std::vector<std::uint8_t> const expected_bson =
{
0x10u, 0x00u, 0x00u, 0x00u, // size (little endian)
0x10u, /// entry: int32
'e', 'n', 't', 'r', 'y', '\x00',
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
0x00u // end marker
};
const auto bson = json::to_bson(j);
CHECK(bson == expected_bson);
auto j_roundtrip = json::from_bson(bson);
CHECK(j_roundtrip.at("entry").is_number_integer());
CHECK(j_roundtrip == j);
CHECK(json::from_bson(bson, true, false) == j);
}
}
SECTION("signed std::int64_t: INT32_MAX+1 .. INT64_MAX")
{
std::vector<int64_t> const numbers
{
(std::numeric_limits<int64_t>::max)(),
1000000000000000000LL,
100000000000000000LL,
10000000000000000LL,
1000000000000000LL,
100000000000000LL,
10000000000000LL,
1000000000000LL,
100000000000LL,
10000000000LL,
static_cast<std::int64_t>((std::numeric_limits<int32_t>::max)()) + 1,
};
for (const auto i : numbers)
{
CAPTURE(i)
json const j =
{
{ "entry", i }
};
CHECK(j.at("entry").is_number_integer());
std::uint64_t const iu = *reinterpret_cast<const std::uint64_t*>(&i);
std::vector<std::uint8_t> const expected_bson =
{
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
0x12u, /// entry: int64
'e', 'n', 't', 'r', 'y', '\x00',
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
0x00u // end marker
};
const auto bson = json::to_bson(j);
CHECK(bson == expected_bson);
auto j_roundtrip = json::from_bson(bson);
CHECK(j_roundtrip.at("entry").is_number_integer());
CHECK(j_roundtrip == j);
CHECK(json::from_bson(bson, true, false) == j);
}
}
}
SECTION("unsigned")
{
SECTION("unsigned std::uint64_t: 0 .. INT32_MAX")
{
std::vector<std::uint64_t> const numbers
{
0ULL,
1ULL,
10ULL,
100ULL,
1000ULL,
10000ULL,
100000ULL,
1000000ULL,
10000000ULL,
100000000ULL,
1000000000ULL,
2147483646ULL,
static_cast<std::uint64_t>((std::numeric_limits<int32_t>::max)())
};
for (const auto i : numbers)
{
CAPTURE(i)
json const j =
{
{ "entry", i }
};
auto iu = i;
std::vector<std::uint8_t> const expected_bson =
{
0x10u, 0x00u, 0x00u, 0x00u, // size (little endian)
0x10u, /// entry: int32
'e', 'n', 't', 'r', 'y', '\x00',
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
0x00u // end marker
};
const auto bson = json::to_bson(j);
CHECK(bson == expected_bson);
auto j_roundtrip = json::from_bson(bson);
CHECK(j.at("entry").is_number_unsigned());
CHECK(j_roundtrip.at("entry").is_number_integer());
CHECK(j_roundtrip == j);
CHECK(json::from_bson(bson, true, false) == j);
}
}
SECTION("unsigned std::uint64_t: INT32_MAX+1 .. INT64_MAX")
{
std::vector<std::uint64_t> const numbers
{
static_cast<std::uint64_t>((std::numeric_limits<std::int32_t>::max)()) + 1,
4000000000ULL,
static_cast<std::uint64_t>((std::numeric_limits<std::uint32_t>::max)()),
10000000000ULL,
100000000000ULL,
1000000000000ULL,
10000000000000ULL,
100000000000000ULL,
1000000000000000ULL,
10000000000000000ULL,
100000000000000000ULL,
1000000000000000000ULL,
static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()),
};
for (const auto i : numbers)
{
CAPTURE(i)
json const j =
{
{ "entry", i }
};
auto iu = i;
std::vector<std::uint8_t> const expected_bson =
{
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
0x12u, /// entry: int64
'e', 'n', 't', 'r', 'y', '\x00',
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
0x00u // end marker
};
const auto bson = json::to_bson(j);
CHECK(bson == expected_bson);
auto j_roundtrip = json::from_bson(bson);
CHECK(j.at("entry").is_number_unsigned());
CHECK(j_roundtrip.at("entry").is_number_integer());
CHECK(j_roundtrip == j);
CHECK(json::from_bson(bson, true, false) == j);
}
}
SECTION("unsigned std::uint64_t: INT64_MAX+1 .. UINT64_MAX")
{
std::vector<std::uint64_t> const numbers
{
static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1ULL,
0xffffffffffffffff,
};
for (const auto i : numbers)
{
CAPTURE(i)
json const j =
{
{ "entry", i }
};
auto iu = i;
std::vector<std::uint8_t> const expected_bson =
{
0x14u, 0x00u, 0x00u, 0x00u, // size (little endian)
0x11u, /// entry: uint64
'e', 'n', 't', 'r', 'y', '\x00',
static_cast<std::uint8_t>((iu >> (8u * 0u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 1u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 2u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 3u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 4u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 5u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 6u)) & 0xffu),
static_cast<std::uint8_t>((iu >> (8u * 7u)) & 0xffu),
0x00u // end marker
};
const auto bson = json::to_bson(j);
CHECK(bson == expected_bson);
auto j_roundtrip = json::from_bson(bson);
CHECK(j.at("entry").is_number_unsigned());
CHECK(j_roundtrip.at("entry").is_number_unsigned());
CHECK(j_roundtrip == j);
CHECK(json::from_bson(bson, true, false) == j);
}
}
}
}
}
TEST_CASE("Parse BSON directly from a file using iterator and sentinel")
{
std::string const filename = TEST_DATA_DIRECTORY "/json.org/1.json";
std::ifstream f_json(filename);
const json expected = json::parse(f_json);
std::ifstream file(filename + ".bson", std::ios::binary);
const std::istreambuf_iterator<char> first(file);
const json parsed = json::from_bson(first, utils::istreambuf_sentinel{});
CHECK(parsed == expected);
}
TEST_CASE("BSON round-trip invariants")
{
// This checks what the parse_bson_fuzzer driver checks (see
// tests/src/fuzzer-parse_bson.cpp), so that a regression shows up in CI
// rather than as an OSS-Fuzz report: anything from_bson() returns (j1)
// can be serialized, parsed back (j2), and serialized again to reproduce
// the exact bytes. BSON only serializes objects, so non-object corpus
// values are skipped.
for (const auto& j0 : utils::round_trip_corpus::values())
{
if (!j0.is_object())
{
continue;
}
json j1;
try
{
// turn the corpus value into a value as from_bson() returns it
j1 = json::from_bson(json::to_bson(j0));
}
catch (const json::exception&)
{
// the fuzzer driver only ever sees values from_bson() actually
// produced, so skip corpus values that do not survive the
// round trip here, too
continue;
}
INFO("j1 = " << j1.dump());
const std::vector<std::uint8_t> vec = json::to_bson(j1);
json j2;
// anything the library writes must be parsable by the library
REQUIRE_NOTHROW(j2 = json::from_bson(vec));
CHECK(json::to_bson(j2) == vec);
}
}
TEST_CASE("BSON roundtrips" * doctest::skip())
{
SECTION("reference files")
{
for (const std::string filename :
{
TEST_DATA_DIRECTORY "/json.org/1.json",
TEST_DATA_DIRECTORY "/json.org/2.json",
TEST_DATA_DIRECTORY "/json.org/3.json",
TEST_DATA_DIRECTORY "/json.org/4.json",
TEST_DATA_DIRECTORY "/json.org/5.json"
})
{
CAPTURE(filename)
{
INFO_WITH_TEMP(filename + ": std::vector<std::uint8_t>");
// parse JSON file
std::ifstream f_json(filename);
const json j1 = json::parse(f_json);
// parse BSON file
auto packed = utils::read_binary_file(filename + ".bson");
json j2;
CHECK_NOTHROW(j2 = json::from_bson(packed));
// compare parsed JSON values
CHECK(j1 == j2);
}
{
INFO_WITH_TEMP(filename + ": std::ifstream");
// parse JSON file
std::ifstream f_json(filename);
const json j1 = json::parse(f_json);
// parse BSON file
std::ifstream f_bson(filename + ".bson", std::ios::binary);
json j2;
CHECK_NOTHROW(j2 = json::from_bson(f_bson));
// compare parsed JSON values
CHECK(j1 == j2);
}
{
INFO_WITH_TEMP(filename + ": uint8_t* and size");
// parse JSON file
std::ifstream f_json(filename);
const json j1 = json::parse(f_json);
// parse BSON file
auto packed = utils::read_binary_file(filename + ".bson");
json j2;
CHECK_NOTHROW(j2 = json::from_bson({packed.data(), packed.size()}));
// compare parsed JSON values
CHECK(j1 == j2);
}
{
INFO_WITH_TEMP(filename + ": output to output adapters");
// parse JSON file
std::ifstream f_json(filename);
json const j1 = json::parse(f_json);
// parse BSON file
auto packed = utils::read_binary_file(filename + ".bson");
{
INFO_WITH_TEMP(filename + ": output adapters: std::vector<std::uint8_t>");
std::vector<std::uint8_t> vec;
json::to_bson(j1, vec);
if (vec != packed)
{
// the exact serializations may differ due to the order of
// object keys; in these cases, just compare whether both
// serializations create the same JSON value
CHECK(json::from_bson(vec) == json::from_bson(packed));
}
}
}
}
}
}
TEST_CASE("BSON: deeply nested values")
{
SECTION("documents and arrays round-trip at every depth")
{
// nested documents and arrays, with siblings on every level, so
// every length prefix covers entries of both kinds
json value = "leaf";
for (std::size_t depth = 0; depth <= 300; ++depth)
{
CAPTURE(depth)
const json document = {{"value", value}, {"n", depth}};
CHECK(json::from_bson(json::to_bson(document)) == document);
value = depth % 2 == 0 ? json{{"a", std::move(value)}, {"b", {1, "x"}}} :
json::array({std::move(value), depth, json::object()});
}
}
SECTION("a key containing U+0000 is rejected before anything is written")
{
json value = json::object({{std::string("bad\0key", 7), 1}});
for (std::size_t depth = 0; depth < 200; ++depth)
{
value = json{{"a", {{"b", 1}}}, {"z", std::move(value)}};
}
std::vector<std::uint8_t> output;
CHECK_THROWS_AS(json::to_bson(value, output), json::out_of_range&);
CHECK(output.empty());
}
SECTION("a binary subtype that doesn't fit a byte is rejected before anything is written (#5675)")
{
// the offending value is nested, so this also covers that the check
// is not limited to a directly written value's own document
json const j = {{"a", {{"b", json::binary({1, 2}, 300)}}}};
std::vector<std::uint8_t> vector_output;
CHECK_THROWS_AS(json::to_bson(j, vector_output), json::out_of_range&);
CHECK(vector_output.empty());
std::string string_output;
CHECK_THROWS_AS(json::to_bson(j, string_output), json::out_of_range&);
CHECK(string_output.empty());
}
SECTION("values nested too deeply for the call stack (#5392)")
{
// serializing recursed once per nesting level, and computed every
// nested document's length by walking everything below it again.
// The values are only parsed, serialized and walked, never copied or
// compared, since those recurse too.
const std::size_t depth = 100000;
for (const bool objects :
{
false, true
})
{
CAPTURE(objects)
std::string text = "{\"a\":";
for (std::size_t i = 0; i < depth; ++i)
{
text += objects ? "{\"a\":" : "[";
}
text += "1";
text.append(depth, objects ? '}' : ']');
text += "}";
const auto bson = json::to_bson(json::parse(text));
const auto result = json::from_bson(bson);
const json* p = &result.at("a");
for (std::size_t i = 0; i < depth; ++i)
{
p = objects ? &p->at("a") : &p->at(0);
}
CHECK(*p == 1);
}
}
}
TEST_CASE("Invalid document size handling")
{
SECTION("document size must be at least 5")
{
std::vector<std::uint8_t> const v = {0x04, 0x00, 0x00, 0x00, 0x00};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 4 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(v, true, false).is_discarded());
}
SECTION("declared document size must match consumed bytes (extra trailing element)")
{
// Declares 5-byte empty document but appends an int32 element after the declared end.
std::vector<std::uint8_t> const v =
{
0x05, 0x00, 0x00, 0x00,
0x10, 'a', 'd', 'm', 'i', 'n', 0x00,
0x01, 0x00, 0x00, 0x00,
0x00
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 16: syntax error while parsing BSON document: document size 5 does not match the number of bytes read (16)", json::parse_error&);
CHECK(json::from_bson(v, true, false).is_discarded());
}
SECTION("declared document size must match consumed bytes (premature terminator)")
{
// Declares 32-byte document but only contains the size field followed by an immediate terminator.
std::vector<std::uint8_t> const v =
{
0x20, 0x00, 0x00, 0x00,
0x00
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 32 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(v, true, false).is_discarded());
}
SECTION("array declared size must match consumed bytes")
{
// Outer object contains an array "a" that declares 5 bytes (empty) but
// actually contains an int32 element before its terminator.
std::vector<std::uint8_t> const v =
{
0x14, 0x00, 0x00, 0x00, // object size = 20
0x04, 'a', 0x00, // key "a", array type
0x05, 0x00, 0x00, 0x00, // array declared size = 5 (empty)
0x10, '0', 0x00, 0x01, 0x00, 0x00, 0x00, // extra int32 element "0" = 1
0x00, // array terminator
0x00 // object terminator
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 19: syntax error while parsing BSON document: document size 5 does not match the number of bytes read (12)", json::parse_error&);
CHECK(json::from_bson(v, true, false).is_discarded());
}
SECTION("BSON string must end with 0x00")
{
// Length-prefixed string whose terminator byte is 'X' (0x58), not 0x00.
std::vector<std::uint8_t> const v =
{
0x0F, 0x00, 0x00, 0x00,
0x02, 's', 0x00,
0x02, 0x00, 0x00, 0x00,
'A', 'X',
0x00
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.112] parse error at byte 13: syntax error while parsing BSON string: BSON string is not null-terminated", json::parse_error&);
CHECK(json::from_bson(v, true, false).is_discarded());
}
}