* Document the benchmarks, and make them build and compare versions again
The benchmark project hasn't configured since #4793: download_test_data.cmake
compiles cmake/detect_libcpp_version.cpp relative to CMAKE_SOURCE_DIR, which
is tests/benchmarks when that is the top-level project, so try_run fails and
so does `make run_benchmarks`. The path is now relative to the module itself,
which is the same file for the main build.
The Dump benchmark discarded dump()'s result, which is [[nodiscard]] by now;
it warned, and left the optimizer free to shorten the loop. The result is
now kept with benchmark::DoNotOptimize.
A new cache variable, JSON_BENCHMARK_INCLUDE_DIR, names the directory holding
the nlohmann/json.hpp to benchmark (single_include by default, as before),
so the same benchmarks can be built against two versions and compared.
tests/benchmarks/README.md documents what is measured, how to build and run
the benchmarks, how to read the output, and how to compare two versions with
Google Benchmark's compare.py; it recommends doing so by hand before a
release rather than in CI.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point ci_benchmarks at tests/benchmarks
The target has configured ${PROJECT_SOURCE_DIR}/benchmarks since it was
added in #2561, but the benchmarks live in tests/benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Pin Google Benchmark to release 1.9.5
The benchmarks fetched Google Benchmark's main branch, so two builds on
different days could measure with different library code, and CMake 3.30
and later warn that the single-argument FetchContent_Populate() is
deprecated. Fetch the 1.9.5 release archive, verified by its SHA-256,
with FetchContent_MakeAvailable() instead. That needs CMake 3.14; Google
Benchmark itself already needed 3.13.
Its -Werror is switched off, so a newer compiler's new warnings cannot
break the pinned release, and its install rules are no longer added.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.
The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.
The round-trip tests need the .bon8 files of json_test_data 3.2.0.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The benchmark suite covered parsing JSON text, dumping, and serializing to
CBOR, but only one binary read: FromMsgpack. Nothing measured from_cbor,
from_ubjson, from_bjdata or from_bson, so a change to binary_reader.hpp had
no baseline to be compared against.
Add read benchmarks for every format, in the two shapes that matter: from a
contiguous buffer, which is what most callers pass, and from a FILE*, which
is what FromMsgpack already measures and which compiles to different code.
FromMsgpack itself is left untouched so its numbers stay comparable across
releases. The input is derived at setup time by serializing a parsed test
file, because the test data repository ships JSON only.
The test files are wide and shallow, but the readers' cost is per container,
so add three value shapes they do not cover -- deeply nested containers, many
sibling containers, and one flat array of numbers -- plus an indefinite-length
CBOR string, a form the writer never emits and which therefore has to be
assembled by hand. UBJSON and BJData are also captured in their size- and
type-annotated form, which the readers handle in a separate code path.
BSON requires an object at the top level, so it cannot reuse the array-rooted
test files; it is captured on the object-rooted ones, and the shapes are
wrapped in an object so every format measures the same value. The setup marks
the benchmark as skipped rather than letting the exception escape if that
requirement is ever violated.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Every input in the benchmark corpus is minified or only lightly spaced, so
none of them exercise the lexer's whitespace handling. Real-world JSON is
frequently indented - configuration files, pretty-printed API responses,
anything kept under version control - where the insignificant whitespace can
outweigh the data itself.
Add a ParseIndented family that re-serializes each document with an
indentation and parses that. The content is identical to the matching
ParseString row, so the pair isolates the cost of the whitespace alone.
Measured here, parsing the indented form costs 16-34% more than the minified
form of the same document, which nothing in the suite currently reports.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* multibyte binary reader
* wide_string_input_adapter fallback to get_character
Update input_adapters.hpp
* Update json.hpp
* Add from msgpack test
* Test for broken msgpack with stream, address some warnings
* Reading binary number from wchar as an error, address warnings
* Not casting float to int, it violates strict aliasing rule
* Add benchmark for cbor serialization and cbor binary data serialization
* Use std::vector::insert method instead of std::copy in output_vector_adapter::write_characters
This change increases a lot the performance when writing lots of binary data.