mirror of
https://github.com/nlohmann/json.git
synced 2026-09-29 22:10:47 +01:00
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
592 lines
26 KiB
C++
592 lines
26 KiB
C++
// __ _____ _____ _____
|
|
// __| | __| | | | JSON for Modern C++ (supporting code)
|
|
// | | |__ | | | | | | version 3.12.0
|
|
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
//
|
|
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
// SPDX-License-Identifier: MIT
|
|
|
|
#include <benchmark/benchmark.h>
|
|
#include <nlohmann/json.hpp>
|
|
#include <fstream>
|
|
#include <numeric>
|
|
#include <vector>
|
|
#include <test_data.hpp>
|
|
|
|
using json = nlohmann::json;
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// parse JSON from file
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
static void ParseFile(benchmark::State& state, const char* filename)
|
|
{
|
|
while (state.KeepRunning())
|
|
{
|
|
state.PauseTiming();
|
|
auto* f = new std::ifstream(filename);
|
|
auto* j = new json();
|
|
state.ResumeTiming();
|
|
|
|
*j = json::parse(*f);
|
|
|
|
state.PauseTiming();
|
|
delete f;
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
std::ifstream file(filename, std::ios::binary | std::ios::ate);
|
|
state.SetBytesProcessed(state.iterations() * file.tellg());
|
|
}
|
|
BENCHMARK_CAPTURE(ParseFile, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json");
|
|
BENCHMARK_CAPTURE(ParseFile, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json");
|
|
BENCHMARK_CAPTURE(ParseFile, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json");
|
|
BENCHMARK_CAPTURE(ParseFile, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json");
|
|
BENCHMARK_CAPTURE(ParseFile, floats, TEST_DATA_DIRECTORY "/regression/floats.json");
|
|
BENCHMARK_CAPTURE(ParseFile, signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json");
|
|
BENCHMARK_CAPTURE(ParseFile, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
|
|
BENCHMARK_CAPTURE(ParseFile, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// parse JSON from string
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
static void ParseString(benchmark::State& state, const char* filename)
|
|
{
|
|
std::ifstream f(filename);
|
|
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
|
|
|
while (state.KeepRunning())
|
|
{
|
|
state.PauseTiming();
|
|
auto* j = new json();
|
|
state.ResumeTiming();
|
|
|
|
*j = json::parse(str);
|
|
|
|
state.PauseTiming();
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * str.size());
|
|
}
|
|
BENCHMARK_CAPTURE(ParseString, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json");
|
|
BENCHMARK_CAPTURE(ParseString, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json");
|
|
BENCHMARK_CAPTURE(ParseString, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json");
|
|
BENCHMARK_CAPTURE(ParseString, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json");
|
|
BENCHMARK_CAPTURE(ParseString, floats, TEST_DATA_DIRECTORY "/regression/floats.json");
|
|
BENCHMARK_CAPTURE(ParseString, signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json");
|
|
BENCHMARK_CAPTURE(ParseString, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
|
|
BENCHMARK_CAPTURE(ParseString, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// parse pretty-printed JSON from string
|
|
//
|
|
// Every file in the corpus above is minified or only lightly spaced, so none of
|
|
// them exercise the lexer's whitespace handling. Real-world JSON is frequently
|
|
// indented - configuration files, pretty-printed API responses, anything kept
|
|
// under version control - where insignificant whitespace can outweigh the data.
|
|
// Re-serializing a document with an indentation and parsing that keeps the
|
|
// content identical to the ParseString row above, so the pair isolates the cost
|
|
// of the whitespace alone.
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
static void ParseIndented(benchmark::State& state, const char* filename, int indent)
|
|
{
|
|
std::ifstream f(filename);
|
|
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
|
const std::string indented = json::parse(str).dump(indent);
|
|
|
|
while (state.KeepRunning())
|
|
{
|
|
state.PauseTiming();
|
|
auto* j = new json();
|
|
state.ResumeTiming();
|
|
|
|
*j = json::parse(indented);
|
|
|
|
state.PauseTiming();
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * indented.size());
|
|
}
|
|
BENCHMARK_CAPTURE(ParseIndented, jeopardy / 4, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", 4);
|
|
BENCHMARK_CAPTURE(ParseIndented, canada / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", 4);
|
|
BENCHMARK_CAPTURE(ParseIndented, citm_catalog / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", 4);
|
|
BENCHMARK_CAPTURE(ParseIndented, twitter / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", 4);
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// serialize JSON
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
static void Dump(benchmark::State& state, const char* filename, int indent)
|
|
{
|
|
std::ifstream f(filename);
|
|
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
|
json j = json::parse(str);
|
|
|
|
while (state.KeepRunning())
|
|
{
|
|
std::string output = j.dump(indent);
|
|
benchmark::DoNotOptimize(output);
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * j.dump(indent).size());
|
|
}
|
|
BENCHMARK_CAPTURE(Dump, jeopardy / -, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, jeopardy / 4, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, canada / -, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, canada / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, citm_catalog / -, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, citm_catalog / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, twitter / -, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, twitter / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, floats / -, TEST_DATA_DIRECTORY "/regression/floats.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, floats / 4, TEST_DATA_DIRECTORY "/regression/floats.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, signed_ints / -, TEST_DATA_DIRECTORY "/regression/signed_ints.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, signed_ints / 4, TEST_DATA_DIRECTORY "/regression/signed_ints.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, unsigned_ints / -, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, unsigned_ints / 4, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json", 4);
|
|
BENCHMARK_CAPTURE(Dump, small_signed_ints / -, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json", -1);
|
|
BENCHMARK_CAPTURE(Dump, small_signed_ints / 4, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json", 4);
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// serialize CBOR
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
static void ToCbor(benchmark::State& state, const char* filename)
|
|
{
|
|
std::ifstream f(filename);
|
|
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
|
json j = json::parse(str);
|
|
|
|
while (state.KeepRunning())
|
|
{
|
|
json::to_cbor(j);
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * json::to_cbor(j).size());
|
|
}
|
|
BENCHMARK_CAPTURE(ToCbor, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json");
|
|
BENCHMARK_CAPTURE(ToCbor, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json");
|
|
BENCHMARK_CAPTURE(ToCbor, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json");
|
|
BENCHMARK_CAPTURE(ToCbor, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json");
|
|
BENCHMARK_CAPTURE(ToCbor, floats, TEST_DATA_DIRECTORY "/regression/floats.json");
|
|
BENCHMARK_CAPTURE(ToCbor, signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json");
|
|
BENCHMARK_CAPTURE(ToCbor, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
|
|
BENCHMARK_CAPTURE(ToCbor, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// Parse Msgpack
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
static void FromMsgpack(benchmark::State& state, const char* filename)
|
|
{
|
|
std::ifstream f(filename);
|
|
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
|
auto bytes = json::to_msgpack(json::parse(str));
|
|
std::ofstream o("test.msgpack");
|
|
o.write((char*)bytes.data(), bytes.size());
|
|
o.flush();
|
|
o.close();
|
|
for (auto _ : state)
|
|
{
|
|
state.PauseTiming();
|
|
auto* j = new json();
|
|
auto file = fopen("test.msgpack", "rb");
|
|
state.ResumeTiming();
|
|
|
|
*j = json::from_msgpack(file);
|
|
|
|
state.PauseTiming();
|
|
fclose(file);
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * bytes.size());
|
|
}
|
|
|
|
BENCHMARK_CAPTURE(FromMsgpack, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, floats, TEST_DATA_DIRECTORY "/regression/floats.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
|
|
BENCHMARK_CAPTURE(FromMsgpack, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// serialize binary CBOR
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
static void BinaryToCbor(benchmark::State& state)
|
|
{
|
|
std::vector<uint8_t> data(256);
|
|
std::iota(data.begin(), data.end(), 0);
|
|
|
|
auto it = data.begin();
|
|
std::vector<uint8_t> in;
|
|
in.reserve(state.range(0));
|
|
for (int i = 0; i < state.range(0); ++i)
|
|
{
|
|
if (it == data.end())
|
|
{
|
|
it = data.begin();
|
|
}
|
|
|
|
in.push_back(*it);
|
|
++it;
|
|
}
|
|
|
|
json::binary_t bin{in};
|
|
json j{{"type", "binary"}, {"data", bin}};
|
|
|
|
while (state.KeepRunning())
|
|
{
|
|
json::to_cbor(j);
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * json::to_cbor(j).size());
|
|
}
|
|
BENCHMARK(BinaryToCbor)->RangeMultiplier(2)->Range(8, 8 << 12);
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// parse binary formats
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
// Only MessagePack had a read benchmark (FromMsgpack above, left untouched so
|
|
// its numbers stay comparable across releases). The benchmarks below cover the
|
|
// other formats, and read from a contiguous buffer as well as from a FILE*:
|
|
// most callers pass a container, and the two adapters compile to different
|
|
// code. The test data repository ships JSON only, so the input for each is
|
|
// derived at setup time by serializing a parsed test file.
|
|
|
|
/// binary format to benchmark; the _optimized variants add UBJSON/BJData size
|
|
/// and type annotations, which the readers handle in a separate code path
|
|
enum class binary_format
|
|
{
|
|
cbor,
|
|
msgpack,
|
|
ubjson,
|
|
ubjson_optimized,
|
|
bjdata,
|
|
bjdata_optimized,
|
|
bson,
|
|
bon8
|
|
};
|
|
|
|
static std::vector<std::uint8_t> to_binary(const json& j, const binary_format format)
|
|
{
|
|
switch (format)
|
|
{
|
|
case binary_format::cbor:
|
|
return json::to_cbor(j);
|
|
case binary_format::msgpack:
|
|
return json::to_msgpack(j);
|
|
case binary_format::ubjson:
|
|
return json::to_ubjson(j);
|
|
case binary_format::ubjson_optimized:
|
|
return json::to_ubjson(j, true, true);
|
|
case binary_format::bjdata:
|
|
return json::to_bjdata(j);
|
|
case binary_format::bjdata_optimized:
|
|
return json::to_bjdata(j, true, true);
|
|
case binary_format::bon8:
|
|
return json::to_bon8(j);
|
|
case binary_format::bson:
|
|
default:
|
|
return json::to_bson(j);
|
|
}
|
|
}
|
|
|
|
static json from_binary(const std::vector<std::uint8_t>& bytes, const binary_format format)
|
|
{
|
|
switch (format)
|
|
{
|
|
case binary_format::cbor:
|
|
return json::from_cbor(bytes);
|
|
case binary_format::msgpack:
|
|
return json::from_msgpack(bytes);
|
|
case binary_format::ubjson:
|
|
case binary_format::ubjson_optimized:
|
|
return json::from_ubjson(bytes);
|
|
case binary_format::bjdata:
|
|
case binary_format::bjdata_optimized:
|
|
return json::from_bjdata(bytes);
|
|
case binary_format::bon8:
|
|
return json::from_bon8(bytes);
|
|
case binary_format::bson:
|
|
default:
|
|
return json::from_bson(bytes);
|
|
}
|
|
}
|
|
|
|
static json from_binary(std::FILE* file, const binary_format format)
|
|
{
|
|
switch (format)
|
|
{
|
|
case binary_format::cbor:
|
|
return json::from_cbor(file);
|
|
case binary_format::msgpack:
|
|
return json::from_msgpack(file);
|
|
case binary_format::ubjson:
|
|
case binary_format::ubjson_optimized:
|
|
return json::from_ubjson(file);
|
|
case binary_format::bjdata:
|
|
case binary_format::bjdata_optimized:
|
|
return json::from_bjdata(file);
|
|
case binary_format::bon8:
|
|
return json::from_bon8(file);
|
|
case binary_format::bson:
|
|
default:
|
|
return json::from_bson(file);
|
|
}
|
|
}
|
|
|
|
/*!
|
|
@brief serialize a parsed test file to @a format
|
|
|
|
Returns an empty vector and marks the benchmark as skipped if the file cannot
|
|
be represented in the format, rather than letting the exception escape: BSON
|
|
requires an object at the top level, and several test files are arrays.
|
|
*/
|
|
static std::vector<std::uint8_t> binary_input(benchmark::State& state, const char* filename, const binary_format format)
|
|
{
|
|
std::ifstream f(filename);
|
|
std::string const str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
|
|
const json j = json::parse(str);
|
|
|
|
if (format == binary_format::bson && !j.is_object())
|
|
{
|
|
state.SkipWithError("BSON requires an object at the top level");
|
|
return {};
|
|
}
|
|
|
|
return to_binary(j, format);
|
|
}
|
|
|
|
static void FromBinaryBuffer(benchmark::State& state, const char* filename, const binary_format format)
|
|
{
|
|
const std::vector<std::uint8_t> bytes = binary_input(state, filename, format);
|
|
if (bytes.empty())
|
|
{
|
|
return;
|
|
}
|
|
|
|
for (auto _ : state)
|
|
{
|
|
// the value is destroyed outside the timed section, because destroying
|
|
// a large DOM is not what this benchmark measures
|
|
state.PauseTiming();
|
|
auto* j = new json();
|
|
state.ResumeTiming();
|
|
|
|
*j = from_binary(bytes, format);
|
|
|
|
state.PauseTiming();
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * bytes.size());
|
|
}
|
|
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / floats, TEST_DATA_DIRECTORY "/regression/floats.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson_optimized / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson_optimized);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson_optimized / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson_optimized);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bjdata);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata_optimized / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bjdata_optimized);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata_optimized / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata_optimized);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bon8 / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::bon8);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bon8 / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bon8);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bon8 / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::bon8);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bon8 / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bon8);
|
|
// BSON requires an object at the top level, so the array-rooted test files
|
|
// (jeopardy and the regression files) cannot be captured here
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bson);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::bson);
|
|
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bson);
|
|
|
|
static void FromBinaryFile(benchmark::State& state, const char* filename, const binary_format format)
|
|
{
|
|
const std::vector<std::uint8_t> bytes = binary_input(state, filename, format);
|
|
if (bytes.empty())
|
|
{
|
|
return;
|
|
}
|
|
|
|
const char* tmp = "benchmark_input.bin";
|
|
std::ofstream o(tmp, std::ios::binary);
|
|
o.write(reinterpret_cast<const char*>(bytes.data()), static_cast<std::streamsize>(bytes.size()));
|
|
o.flush();
|
|
o.close();
|
|
|
|
for (auto _ : state)
|
|
{
|
|
state.PauseTiming();
|
|
auto* j = new json();
|
|
auto* file = std::fopen(tmp, "rb");
|
|
state.ResumeTiming();
|
|
|
|
*j = from_binary(file, format);
|
|
|
|
state.PauseTiming();
|
|
std::fclose(file);
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * bytes.size());
|
|
}
|
|
|
|
BENCHMARK_CAPTURE(FromBinaryFile, cbor / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, cbor / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, ubjson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, ubjson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, bjdata / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, bon8 / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bon8);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, bon8 / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bon8);
|
|
BENCHMARK_CAPTURE(FromBinaryFile, bson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bson);
|
|
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
// parse binary formats: value shapes
|
|
//////////////////////////////////////////////////////////////////////////////
|
|
|
|
// The test files above are wide and shallow, but the readers' cost is per
|
|
// container, so these cover the shapes that stress the container handling
|
|
// itself. Every shape is wrapped in an object so that BSON, which requires an
|
|
// object at the top level, measures the same value as the other formats.
|
|
|
|
/// deeply nested arrays: one container per level, no other work
|
|
static json make_nested()
|
|
{
|
|
json nested = json::array();
|
|
json* p = &nested;
|
|
for (std::size_t i = 1; i < 1000; ++i)
|
|
{
|
|
p->push_back(json::array());
|
|
p = &p->operator[](0);
|
|
}
|
|
|
|
json j = json::object();
|
|
j["data"] = std::move(nested);
|
|
return j;
|
|
}
|
|
|
|
/// many sibling containers: maximum container churn, minimum nesting
|
|
static json make_containers()
|
|
{
|
|
json data = json::array();
|
|
for (std::size_t i = 0; i < 100000; ++i)
|
|
{
|
|
data.push_back(json::array({1, 2}));
|
|
}
|
|
|
|
json j = json::object();
|
|
j["data"] = std::move(data);
|
|
return j;
|
|
}
|
|
|
|
/// one flat array of numbers: the scalar decoding path, which must not move
|
|
static json make_scalars()
|
|
{
|
|
json data = json::array();
|
|
for (std::size_t i = 0; i < 1000000; ++i)
|
|
{
|
|
data.push_back(i);
|
|
}
|
|
|
|
json j = json::object();
|
|
j["data"] = std::move(data);
|
|
return j;
|
|
}
|
|
|
|
static void FromBinaryShape(benchmark::State& state, json (*build)(), const binary_format format)
|
|
{
|
|
const std::vector<std::uint8_t> bytes = to_binary(build(), format);
|
|
|
|
for (auto _ : state)
|
|
{
|
|
state.PauseTiming();
|
|
auto* j = new json();
|
|
state.ResumeTiming();
|
|
|
|
*j = from_binary(bytes, format);
|
|
|
|
state.PauseTiming();
|
|
delete j;
|
|
state.ResumeTiming();
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * bytes.size());
|
|
}
|
|
|
|
BENCHMARK_CAPTURE(FromBinaryShape, nested / cbor, make_nested, binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, nested / msgpack, make_nested, binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, nested / ubjson, make_nested, binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, nested / bjdata, make_nested, binary_format::bjdata);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, nested / bson, make_nested, binary_format::bson);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, nested / bon8, make_nested, binary_format::bon8);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / cbor, make_containers, binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / msgpack, make_containers, binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / ubjson, make_containers, binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / ubjson_optimized, make_containers, binary_format::ubjson_optimized);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / bjdata, make_containers, binary_format::bjdata);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / bson, make_containers, binary_format::bson);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, containers / bon8, make_containers, binary_format::bon8);
|
|
// BSON names every array element, so a large array measures key generation
|
|
// rather than scalar decoding and is left out here
|
|
BENCHMARK_CAPTURE(FromBinaryShape, scalars / cbor, make_scalars, binary_format::cbor);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, scalars / msgpack, make_scalars, binary_format::msgpack);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, scalars / ubjson, make_scalars, binary_format::ubjson);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, scalars / bjdata, make_scalars, binary_format::bjdata);
|
|
BENCHMARK_CAPTURE(FromBinaryShape, scalars / bon8, make_scalars, binary_format::bon8);
|
|
|
|
/*!
|
|
@brief parse an indefinite-length CBOR string
|
|
|
|
The writer never emits this form, so the input is assembled by hand: 0x7F
|
|
opens the string, each chunk is a one-character string, and 0xFF closes it.
|
|
*/
|
|
static void FromCborChunkedString(benchmark::State& state, const std::size_t chunks)
|
|
{
|
|
std::vector<std::uint8_t> bytes;
|
|
bytes.reserve(2 * chunks + 2);
|
|
bytes.push_back(0x7F);
|
|
for (std::size_t i = 0; i < chunks; ++i)
|
|
{
|
|
bytes.push_back(0x61); // string of length 1
|
|
bytes.push_back(0x61); // 'a'
|
|
}
|
|
bytes.push_back(0xFF);
|
|
|
|
for (auto _ : state)
|
|
{
|
|
json j = json::from_cbor(bytes);
|
|
benchmark::DoNotOptimize(j);
|
|
}
|
|
|
|
state.SetBytesProcessed(state.iterations() * bytes.size());
|
|
}
|
|
|
|
BENCHMARK_CAPTURE(FromCborChunkedString, 10000 chunks, 10000);
|
|
|
|
BENCHMARK_MAIN();
|