mirror of
https://github.com/nlohmann/json.git
synced 2026-09-28 22:06:02 +01:00
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
378 lines
11 KiB
C++
378 lines
11 KiB
C++
// __ _____ _____ _____
|
|
// __| | __| | | | JSON for Modern C++ (supporting code)
|
|
// | | |__ | | | | | | version 3.12.0
|
|
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
//
|
|
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
// SPDX-FileCopyrightText: 2018 Vitaliy Manushkin <agri@akamo.info>
|
|
// SPDX-License-Identifier: MIT
|
|
|
|
#include "doctest_compatibility.h"
|
|
|
|
#include <nlohmann/json.hpp>
|
|
|
|
#include <cstdint>
|
|
#include <string>
|
|
#include <utility>
|
|
#include <vector>
|
|
|
|
/* forward declarations */
|
|
class alt_string;
|
|
bool operator<(const char* op1, const alt_string& op2) noexcept; // NOLINT(misc-use-internal-linkage)
|
|
void int_to_string(alt_string& target, std::size_t value); // NOLINT(misc-use-internal-linkage)
|
|
|
|
/*
|
|
* This is virtually a string class.
|
|
* It covers std::string under the hood.
|
|
*
|
|
* It deliberately does not provide c_str(), back(), find(str, pos), replace(),
|
|
* or substr(): the library must not rely on them. Do not add members here
|
|
* without checking that the library actually needs them.
|
|
*/
|
|
class alt_string
|
|
{
|
|
public:
|
|
using value_type = std::string::value_type;
|
|
|
|
static constexpr auto npos = (std::numeric_limits<std::size_t>::max)();
|
|
|
|
alt_string(const char* str): str_impl(str) {}
|
|
alt_string(const char* str, std::size_t count): str_impl(str, count) {}
|
|
alt_string(size_t count, char chr): str_impl(count, chr) {}
|
|
alt_string() = default;
|
|
|
|
alt_string& append(char ch)
|
|
{
|
|
str_impl.push_back(ch);
|
|
return *this;
|
|
}
|
|
|
|
alt_string& append(const alt_string& str)
|
|
{
|
|
str_impl.append(str.str_impl);
|
|
return *this;
|
|
}
|
|
|
|
alt_string& append(const char* s, std::size_t length)
|
|
{
|
|
str_impl.append(s, length);
|
|
return *this;
|
|
}
|
|
|
|
void push_back(char c)
|
|
{
|
|
str_impl.push_back(c);
|
|
}
|
|
|
|
template <typename op_type>
|
|
bool operator==(const op_type& op) const
|
|
{
|
|
return str_impl == op;
|
|
}
|
|
|
|
bool operator==(const alt_string& op) const
|
|
{
|
|
return str_impl == op.str_impl;
|
|
}
|
|
|
|
template <typename op_type>
|
|
bool operator!=(const op_type& op) const
|
|
{
|
|
return str_impl != op;
|
|
}
|
|
|
|
bool operator!=(const alt_string& op) const
|
|
{
|
|
return str_impl != op.str_impl;
|
|
}
|
|
|
|
std::size_t size() const noexcept
|
|
{
|
|
return str_impl.size();
|
|
}
|
|
|
|
void resize (std::size_t n)
|
|
{
|
|
str_impl.resize(n);
|
|
}
|
|
|
|
void resize (std::size_t n, char c)
|
|
{
|
|
str_impl.resize(n, c);
|
|
}
|
|
|
|
template <typename op_type>
|
|
bool operator<(const op_type& op) const noexcept
|
|
{
|
|
return str_impl < op;
|
|
}
|
|
|
|
bool operator<(const alt_string& op) const noexcept
|
|
{
|
|
return str_impl < op.str_impl;
|
|
}
|
|
|
|
char& operator[](std::size_t index)
|
|
{
|
|
return str_impl[index];
|
|
}
|
|
|
|
const char& operator[](std::size_t index) const
|
|
{
|
|
return str_impl[index];
|
|
}
|
|
|
|
void clear()
|
|
{
|
|
str_impl.clear();
|
|
}
|
|
|
|
const value_type* data() const
|
|
{
|
|
return str_impl.data();
|
|
}
|
|
|
|
bool empty() const
|
|
{
|
|
return str_impl.empty();
|
|
}
|
|
|
|
std::size_t find_first_of(char c, std::size_t pos = 0) const
|
|
{
|
|
return str_impl.find_first_of(c, pos);
|
|
}
|
|
|
|
void reserve( std::size_t new_cap = 0 )
|
|
{
|
|
str_impl.reserve(new_cap);
|
|
}
|
|
|
|
private:
|
|
std::string str_impl {}; // NOLINT(readability-redundant-member-init)
|
|
|
|
friend bool operator<(const char* /*op1*/, const alt_string& /*op2*/) noexcept;
|
|
};
|
|
|
|
void int_to_string(alt_string& target, std::size_t value)
|
|
{
|
|
target = std::to_string(value).c_str();
|
|
}
|
|
|
|
using alt_json = nlohmann::basic_json <
|
|
std::map,
|
|
std::vector,
|
|
alt_string,
|
|
bool,
|
|
std::int64_t,
|
|
std::uint64_t,
|
|
double,
|
|
std::allocator,
|
|
nlohmann::adl_serializer >;
|
|
|
|
bool operator<(const char* op1, const alt_string& op2) noexcept
|
|
{
|
|
return op1 < op2.str_impl;
|
|
}
|
|
|
|
TEST_CASE("alternative string type")
|
|
{
|
|
SECTION("binary formats")
|
|
{
|
|
alt_json doc;
|
|
doc["pi"] = 3.141;
|
|
doc["happy"] = true;
|
|
doc["list"] = {1, 2, 3};
|
|
|
|
CHECK(alt_json::from_cbor(alt_json::to_cbor(doc)) == doc);
|
|
CHECK(alt_json::from_msgpack(alt_json::to_msgpack(doc)) == doc);
|
|
CHECK(alt_json::from_bon8(alt_json::to_bon8(doc)) == doc);
|
|
// BSON is not covered: it additionally needs string_t::find(value_type),
|
|
// which alt_string does not provide
|
|
CHECK(alt_json::from_ubjson(alt_json::to_ubjson(doc)) == doc);
|
|
|
|
// a UBJSON high-precision number is parsed into a std::string that the
|
|
// reader has to hand to the SAX interface as an alt_string
|
|
const std::vector<uint8_t> high_precision =
|
|
{
|
|
'H', 'i', 0x16, '3', '.', '1', '4', '1', '5', '9', '2', '6', '5', '3',
|
|
'5', '8', '9', '7', '9', '3', '2', '3', '8', '4', '6'
|
|
};
|
|
const auto number = alt_json::from_ubjson(high_precision);
|
|
CHECK(number.is_number_float());
|
|
CHECK(number.get<double>() == doctest::Approx(3.14159265358979323846));
|
|
}
|
|
|
|
SECTION("dump")
|
|
{
|
|
{
|
|
alt_json doc;
|
|
doc["pi"] = 3.141;
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"pi":3.141})");
|
|
}
|
|
|
|
{
|
|
alt_json doc;
|
|
doc["happy"] = true;
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"happy":true})");
|
|
}
|
|
|
|
{
|
|
alt_json doc;
|
|
doc["name"] = "I'm Batman";
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"name":"I'm Batman"})");
|
|
}
|
|
|
|
{
|
|
alt_json doc;
|
|
doc["nothing"] = nullptr;
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"nothing":null})");
|
|
}
|
|
|
|
{
|
|
alt_json doc;
|
|
doc["answer"]["everything"] = 42;
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"answer":{"everything":42}})");
|
|
}
|
|
|
|
{
|
|
alt_json doc;
|
|
doc["list"] = { 1, 0, 2 };
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"list":[1,0,2]})");
|
|
}
|
|
|
|
{
|
|
alt_json doc;
|
|
doc["object"] = { {"currency", "USD"}, {"value", 42.99} };
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"object":{"currency":"USD","value":42.99}})");
|
|
}
|
|
}
|
|
|
|
SECTION("parse")
|
|
{
|
|
auto doc = alt_json::parse(R"({"foo": "bar"})");
|
|
const alt_string dump = doc.dump();
|
|
CHECK(dump == R"({"foo":"bar"})");
|
|
}
|
|
|
|
SECTION("items")
|
|
{
|
|
auto doc = alt_json::parse(R"({"foo": "bar"})");
|
|
|
|
for (const auto& item : doc.items())
|
|
{
|
|
CHECK(item.key() == "foo");
|
|
CHECK(item.value() == "bar");
|
|
}
|
|
|
|
auto doc_array = alt_json::parse(R"(["foo", "bar"])");
|
|
|
|
for (const auto& item : doc_array.items())
|
|
{
|
|
if (item.key() == "0" )
|
|
{
|
|
CHECK( item.value() == "foo" );
|
|
}
|
|
else if (item.key() == "1" )
|
|
{
|
|
CHECK(item.value() == "bar");
|
|
}
|
|
else
|
|
{
|
|
CHECK(false);
|
|
}
|
|
}
|
|
}
|
|
|
|
SECTION("equality")
|
|
{
|
|
alt_json doc;
|
|
doc["Who are you?"] = "I'm Batman";
|
|
|
|
CHECK("I'm Batman" == doc["Who are you?"]);
|
|
CHECK(doc["Who are you?"] == "I'm Batman");
|
|
CHECK_FALSE("I'm Batman" != doc["Who are you?"]);
|
|
CHECK_FALSE(doc["Who are you?"] != "I'm Batman");
|
|
|
|
CHECK("I'm Bruce Wayne" != doc["Who are you?"]);
|
|
CHECK(doc["Who are you?"] != "I'm Bruce Wayne");
|
|
CHECK_FALSE("I'm Bruce Wayne" == doc["Who are you?"]);
|
|
CHECK_FALSE(doc["Who are you?"] == "I'm Bruce Wayne");
|
|
|
|
{
|
|
const alt_json& const_doc = doc;
|
|
|
|
CHECK("I'm Batman" == const_doc["Who are you?"]);
|
|
CHECK(const_doc["Who are you?"] == "I'm Batman");
|
|
CHECK_FALSE("I'm Batman" != const_doc["Who are you?"]);
|
|
CHECK_FALSE(const_doc["Who are you?"] != "I'm Batman");
|
|
|
|
CHECK("I'm Bruce Wayne" != const_doc["Who are you?"]);
|
|
CHECK(const_doc["Who are you?"] != "I'm Bruce Wayne");
|
|
CHECK_FALSE("I'm Bruce Wayne" == const_doc["Who are you?"]);
|
|
CHECK_FALSE(const_doc["Who are you?"] == "I'm Bruce Wayne");
|
|
}
|
|
}
|
|
|
|
SECTION("JSON pointer")
|
|
{
|
|
// Direct conversion from a json literal to alt_json is not supported due to issue #3425:
|
|
// alt_json's string_t (alt_string) is not directly constructible from std::string, so the
|
|
// cross-basic_json conversion falls back to the array-conversion path, incorrectly representing
|
|
// objects as arrays of [key, value] pairs and strings as arrays of character codes.
|
|
// See https://github.com/nlohmann/json/issues/3425 for details.
|
|
// Workaround: use alt_json::parse() instead of implicit conversion.
|
|
auto j = alt_json::parse(R"({"foo": ["bar", "baz"]})");
|
|
|
|
CHECK(j.at(alt_json::json_pointer("/foo/0")) == j["foo"][0]);
|
|
CHECK(j.at(alt_json::json_pointer("/foo/1")) == j["foo"][1]);
|
|
|
|
// RFC 6901 escaping works without string_t::find(str, pos), replace(),
|
|
// and substr()
|
|
auto j2 = alt_json::parse(R"({"a/b": 1, "m~n": 2, "~/~~//": 3})");
|
|
CHECK(j2.at(alt_json::json_pointer("/a~1b")) == 1);
|
|
CHECK(j2.at(alt_json::json_pointer("/m~0n")) == 2);
|
|
CHECK(j2.at(alt_json::json_pointer("/~0~1~0~0~1~1")) == 3);
|
|
CHECK(alt_json::json_pointer("/~0~1~0~0~1~1").to_string() == alt_string("/~0~1~0~0~1~1"));
|
|
CHECK(j2.flatten().unflatten() == j2);
|
|
}
|
|
|
|
SECTION("patch")
|
|
{
|
|
alt_json const patch1 = alt_json::parse(R"([{ "op": "add", "path": "/a/b", "value": [ "foo", "bar" ] }])");
|
|
alt_json const doc1 = alt_json::parse(R"({ "a": { "foo": 1 } })");
|
|
|
|
CHECK_NOTHROW(doc1.patch(patch1));
|
|
alt_json doc1_ans = alt_json::parse(R"(
|
|
{
|
|
"a": {
|
|
"foo": 1,
|
|
"b": [ "foo", "bar" ]
|
|
}
|
|
}
|
|
)");
|
|
CHECK(doc1.patch(patch1) == doc1_ans);
|
|
}
|
|
|
|
SECTION("diff")
|
|
{
|
|
alt_json const j1 = {"foo", "bar", "baz"};
|
|
alt_json const j2 = {"foo", "bam"};
|
|
CHECK(alt_json::diff(j1, j2).dump() == "[{\"op\":\"replace\",\"path\":\"/1\",\"value\":\"bam\"},{\"op\":\"remove\",\"path\":\"/2\"}]");
|
|
}
|
|
|
|
SECTION("flatten")
|
|
{
|
|
// a JSON value
|
|
const alt_json j = alt_json::parse(R"({"foo": ["bar", "baz"]})");
|
|
const auto j2 = j.flatten();
|
|
CHECK(j2.dump() == R"({"/foo/0":"bar","/foo/1":"baz"})");
|
|
}
|
|
}
|