Commit Graph
5 Commits
Author SHA1 Message Date
Niels Lohmann 1e101ecac1 Add BON8 support (#2998)
* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 16:56:21 +02:00
Niels Lohmann 2e91641de2 Test JSON_BRACE_INIT_COPY_SEMANTICS for real, and fix one-element tuples under it (#5544)
* Test JSON_BRACE_INIT_COPY_SEMANTICS for real, and fix one-element tuples under it

The opt-in JSON_BRACE_INIT_COPY_SEMANTICS was never exercised by CI:

- Its only test, in unit-regression3.cpp, was guarded by
  `#if defined(JSON_BRACE_INIT_COPY_SEMANTICS)` after the #include. The
  header #undefs the macro unconditionally in macro_unscope.hpp, so the
  guard was always false and the test compiled to nothing, whatever -D
  flag was passed.
- The ci_test_brace_init_copy_semantics target that passes the flag was
  not named by any workflow.

Move the test into its own translation unit that defines the macro before
including the header, as unit-diagnostics.cpp does for JSON_DIAGNOSTICS.
It now runs in every CI job and for every standard. Remove the unused
target: it ran the whole suite with the macro, and that suite deliberately
relies on default brace-init semantics in about 90 places
(e.g. `json({1})` meaning `[1]`), so it could never pass.

Running the whole suite with the macro did find one library bug:
to_json for std::tuple builds `j = { std::get<Idx>(t)... }`, so with copy
semantics a one-element tuple became its element. `json(std::tuple<int>{5})`
was `5` instead of `[5]`, and `get<std::tuple<int>>()` threw type_error.302
on the result. Under the macro, a one-element tuple now builds exactly what
the default deduction builds. Without the macro nothing changes.

The new tests also pin that the library's other conversions produce the
same values with and without the macro. The macro page now says that the
macro affects every single-element list (`json j = {1}` is `1`), and that
all translation units must agree on it, since it has no ABI tag.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make JSON_BRACE_INIT_COPY_SEMANTICS part of the ABI tag

The macro changes the body of the initializer-list constructor and adds a
to_json_tuple_impl overload, both with the same mangled names in either
mode, so mixing translation units silently picked one definition. Encode
it in the inline namespace as `_bics`, as JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
does with `_ldvcmp`. The macro is new in the unreleased 3.13.0, so no
existing namespace name changes.

- Move the macro's default into abi_macros.hpp so json_fwd.hpp computes
  the same namespace, and keep it defined under JSON_TEST_KEEP_MACROS.
- Check the tag in the ABI config tests and in the unit test.
- List `_bics` (and the missing `_dp`) in the namespace docs and in the
  natvis generator; regenerate nlohmann_json.natvis.
- Replace the "define it consistently" warning with an ABI note.

Suggested by @gregmarr in the review of #5544.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the cppcheck, clang-tidy and legacy-comparison CI failures

- to_json_tuple_impl() moved the element in both branches of a ternary;
  only one runs, but cppcheck reported accessMoved. Use if/else.
- The ABI tag test looked for "json_abi_bics", which misses when another
  tag comes first, as in json_abi_ldvcmp_bics; look for "_bics".
- readability-qualified-auto in the items() test.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-24 17:05:37 +02:00
Niels Lohmann 75efd6b1c3 Test suite: cover untested macro configs, std::formatter branches, patch_inplace, and fix a duplicate TEST_CASE name (#5492)
* test: cover JSON_NO_IO, JSON_THROW/TRY/CATCH_USER, JSON_SKIP_LIBRARY_VERSION_CHECK, and JSON_DisableEnumSerialization in CI (#5423)

These four supported configuration macros were never actually compiled
anywhere in the test matrix:

- JSON_NO_IO and the JSON_THROW_USER/JSON_TRY_USER/JSON_CATCH_USER trio
  are exercised together in a new tests/src/unit-no_io_and_user_exceptions.cpp,
  which is automatically picked up by the existing unit-*.cpp test glob and
  thus built across the whole standard test matrix.
- JSON_SKIP_LIBRARY_VERSION_CHECK is exercised by a new, dedicated
  tests/src/skip_library_version_check.cpp, compiled directly by the new
  ci_test_skiplibraryversioncheck target in cmake/ci.cmake: the scenario it
  simulates (mixing two differently-versioned inclusions of the library)
  unavoidably triggers the compiler's own "macro redefined" warning, which
  would fail under the library's own -Weverything/-Werror unit test matrix
  for a reason unrelated to the macro under test.
- JSON_DisableEnumSerialization already had #if-guarded tests in several
  unit-*.cpp files (from #4384), but no CMake target ever actually set the
  JSON_DisableEnumSerialization CMake option, so that guarded code was never
  compiled. Add ci_test_disableenumserialization, mirroring the existing
  ci_test_noimplicitconversions/ci_test_noglobaludls targets. Building the
  full test suite with this option on surfaced one real, narrow gap: get<T>()
  on std::vector<std::byte> (used by unit-regression2.cpp's custom BinaryType
  tests) relies on std::byte being handled via enum serialization, so add the
  same #if-guard convention to the two affected SECTIONs there.

Both new CI targets are added to the ci_cmake_options matrix in
.github/workflows/ubuntu.yml, alongside the existing ci_test_* targets.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: cover multi-digit widths and bare alignment in std::formatter<json> (#5423)

Every existing std::formatter spec with a width used a single digit (e.g.
"{:2}"), so the width-parsing loop's accumulation of a second/third digit was
never exercised; add multi-digit width cases. Likewise, every existing spec
with an alignment character also had an explicit fill character, so the
bare-alignment branch (e.g. "{:<}", with no fill) was never exercised; add
cases asserting it keeps the default space indent character.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: add coverage for patch_inplace() (#5423)

patch_inplace() had no unit test at all. Add a happy-path case mirroring an
existing patch() example, and -- more importantly -- pin its distinguishing
contract versus patch(): when a multi-operation JSON Patch fails partway
through, patch_inplace() (which mutates the document directly, operation by
operation) leaves whatever operations already succeeded applied, whereas
patch() (which applies the patch to an internal copy that is discarded on
exception) leaves the original completely untouched either way. Verified
empirically against the current implementation before writing the assertions.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: fix duplicate TEST_CASE name in unit-no-mem-leak-on-adl-serialize.cpp (#5423)

Two distinct TEST_CASEs were both named "check_for_mem_leak_on_adl_to_json-2".
doctest allows duplicate names, so both still ran, but it makes
--test-case=<name> filtering and reporting ambiguous. Rename the second one
to "-3", continuing the existing "-1"/"-2" sequence.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: add direct coverage for the std::u8string to_json overload (#5423)

The ADL to_json overload for std::basic_string<char8_t, ...> was only ever
reached indirectly, via std::filesystem::path::u8string(). Add a test that
constructs a json value directly from a std::u8string, gated the same way as
the overload itself (include/nlohmann/detail/conversions/to_json.hpp): behind
both the std::filesystem::path feature guard and __cpp_lib_char8_t, since the
overload only exists when both are satisfied.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: verify move semantics of byte_container_with_subtype's rvalue constructors (#5423)

The two rvalue-reference constructors were never distinguished from their
const-lvalue-reference twins by any test. Add a "move semantics" section that
constructs from an rvalue std::vector, checks the resulting container keeps
the exact same buffer address as the source (a stronger check than just
observing the source ended up empty, since a copy-then-clear could do that
too), and confirms the source vector was left empty.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard patch_inplace() partial-application test against JSON_NOEXCEPTION

The "distinguishing contract vs patch(): partial application on
failure" test relies on doc.patch_inplace(patch) actually throwing so
the partially-applied state can be observed right after the throw
point. Under ci_test_noexceptions, JSON_THROW() calls std::abort()
instead of throwing, and doctest's --no-throw test filter (which that
CI job passes) makes CHECK_THROWS_AS() a no-op that never even
evaluates its expression -- so patch_inplace() is never called and the
follow-up assertions fail against the untouched original document.

Guard the whole SECTION with #if !defined(JSON_NOEXCEPTION), following
the same convention already used elsewhere in the test suite (e.g.
unit-class_parser.cpp) for exception-dependent tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix MSVC C2220 in the std::u8string conversion test

MSVC's C5321 ("nonstandard extension used: encoding '\xNN' as a
multi-byte utf-8 character") is promoted to a hard error by our MSVC CI
configs. It fires because the test composed a non-ASCII UTF-8 sequence
inside a u8"" literal using raw \x byte escapes; MSVC treats that as
nonstandard and suggests using \u universal-character-names instead,
which every compiler agrees on and which compiles down to the exact
same encoded bytes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard the JSON_THROW_USER test against JSON_NOEXCEPTION and GCC's -Wunused-result

Two independent CI configurations failed to build/run this new test:

- ci_test_noexceptions runs the whole suite with -DJSON_NOEXCEPTION and
  doctest's "--no-throw" filter, which compiles CHECK_THROWS_AS() down
  to a no-op that never even invokes the guarded expression. Since this
  test's whole point is to observe json_throw_user_call_count after
  json::parse()/at() actually throw, it can't be meaningfully run under
  that filter (our JSON_THROW_USER override still throws real
  exceptions regardless of JSON_NOEXCEPTION, but the assertion never
  gets a chance to run). Guard the TEST_CASE with
  #if !defined(JSON_NOEXCEPTION), mirroring the existing precedent in
  unit-json_patch.cpp.

- ci_test_gcc and ci_test_standards_gcc(11) failed with
  -Werror=unused-result on the discarded json::parse() return value.
  json::parse() is marked warn_unused_result, and unlike a real
  [[nodiscard]] attribute, GCC does not consider that satisfied by
  doctest's (void)-cast around the expression in C++11 mode. Assign the
  result to a discarded local instead, matching the established
  `json _ = json::parse(...)` idiom already used throughout
  unit-class_parser.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Suppress a clang-tidy false positive on an intentional defensive copy

performance-unnecessary-copy-initialization suggests copy_for_patch
could be a reference since it's never modified -- but the copy is the
point: it guards against a hypothetical regression where patch()
mutates its receiver, which a reference could never catch (the
follow-up assertion would just compare `original` to itself).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy findings in the JSON_NO_IO/JSON_THROW_USER test

- bugprone-macro-parentheses: wrap the JSON_THROW_USER macro argument in
  parentheses at the throw site.
- modernize-raw-string-literal: switch two escaped JSON string literals to
  raw string literals.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:15:46 +02:00
Niels Lohmann 1da2f68992 Only reserve array capacity if the array type supports it (#5522) 2026-09-14 06:44:01 +02:00
Niels Lohmann d9c55eb225 Split unit-regression2.cpp so the MinGW linker can relocate it (#5511)
Linking test-regression2 with clang and MinGW fails with

    relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'

once the translation unit grows past a certain size: the code can no longer
reach the read-only data it references within the range of a 32-bit
relocation. The file is one of the largest in the test suite and had been
sitting just under that limit, so an unrelated change elsewhere in the
library is enough to tip it over. It is already the second such file --
unit-regression1.cpp was split for size before -- and windows.yml already
carries a workaround for the same limit hitting the debug sections of this
same target, where -g0 was enough because that relocation was against
`.debug_line'. This one is against `.rdata', which no compiler flag avoids.

Move the second half of the regression tests, and the helper types only they
use, into unit-regression3.cpp. The sections are independent -- every
statement in "regression tests 2" was already inside a SECTION -- so they
move unchanged, and the counts confirm nothing was lost: 168 assertions
before the split, 50 plus 118 after.

The result is that both files are comfortably smaller than the one that used
to link, measured with clang at -O1 for C++20:

                        read-only data        text     object
    before                      58,233   1,287,764  3,158,120
    unit-regression2.cpp        48,161   1,012,988  2,522,296
    unit-regression3.cpp        41,710     772,704  1,878,880

No CMake change is needed: tests/CMakeLists.txt globs src/unit-*.cpp, so the
new file is picked up and built for every standard like its siblings.

CONTRIBUTING.md pointed contributors at unit-regression2.cpp for new bug
tests; it now points at the smaller file and says why the two exist, so the
split does not quietly undo itself.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:42 +02:00