Add BON8 support (#2998)

* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-27 16:56:21 +02:00
committed by GitHub
parent f682cd2ef1
commit 1e101ecac1
58 changed files with 3950 additions and 53 deletions
@@ -0,0 +1,159 @@
# BON8
BON8 (Binary Object Notation 8) is a compact binary serialization format for JSON values. It uses the byte values that
cannot begin a UTF-8 character as type markers, so strings are stored as plain UTF-8 without a length prefix: a string
ends at the first byte that cannot continue it. Integers from -10 to 39, `true`, `false`, `null`, and the floating-point
values -1.0, 0.0, and 1.0 take a single byte, and arrays and objects with up to four elements need no terminator.
!!! abstract "References"
- [BON8 specification](https://github.com/hikoworks/hikogui/blob/main/docs/BON8.md)
- [Reference implementation](https://github.com/hikoworks/hikogui/blob/main/src/hikogui/codec/BON8.hpp) in HikoGUI
## Serialization
The library uses the following mapping from JSON values types to BON8 types according to the BON8 specification:
| JSON value type | value/range | BON8 type | first byte |
|-----------------|----------------------------------------------|-------------------------------|------------|
| null | `null` | null | 0xFA |
| boolean | `true` | true | 0xF9 |
| boolean | `false` | false | 0xF8 |
| number_integer | -9223372036854775808..-2147483649 | int64 | 0x8D |
| number_integer | -2147483648..-33818507 | int32 | 0x8C |
| number_integer | -33818506..-264075 | 4-byte negative integer | 0xF0..0xF7 |
| number_integer | -264074..-1931 | 3-byte negative integer | 0xE0..0xEF |
| number_integer | -1930..-11 | 2-byte negative integer | 0xC2..0xDF |
| number_integer | -10..-1 | 1-byte negative integer | 0xB8..0xC1 |
| number_integer | 0..39 | 1-byte positive integer | 0x90..0xB7 |
| number_integer | 40..3879 | 2-byte positive integer | 0xC2..0xDF |
| number_integer | 3880..528167 | 3-byte positive integer | 0xE0..0xEF |
| number_integer | 528168..67637031 | 4-byte positive integer | 0xF0..0xF7 |
| number_integer | 67637032..2147483647 | int32 | 0x8C |
| number_integer | 2147483648..9223372036854775807 | int64 | 0x8D |
| number_unsigned | 0..39 | 1-byte positive integer | 0x90..0xB7 |
| number_unsigned | 40..3879 | 2-byte positive integer | 0xC2..0xDF |
| number_unsigned | 3880..528167 | 3-byte positive integer | 0xE0..0xEF |
| number_unsigned | 528168..67637031 | 4-byte positive integer | 0xF0..0xF7 |
| number_unsigned | 67637032..2147483647 | int32 | 0x8C |
| number_unsigned | 2147483648..9223372036854775807 | int64 | 0x8D |
| number_float | `-1.0` | -1.0 | 0xFB |
| number_float | `0.0` | 0.0 | 0xFC |
| number_float | `1.0` | 1.0 | 0xFD |
| number_float | *any other value representable by a float* | binary32 | 0x8E |
| number_float | *any value NOT representable by a float* | binary64 | 0x8F |
| string | *empty* | end of string | 0xFF |
| string | *non-empty* | UTF-8 string | 0x00..0x7F, 0xC2..0xF4 |
| array | *size*: 0..4 | array with count | 0x80..0x84 |
| array | *size*: 5 or more | array (terminated by 0xFE) | 0x85 |
| object | *size*: 0..4 | object with count | 0x86..0x8A |
| object | *size*: 5 or more | object (terminated by 0xFE) | 0x8B |
| binary | *size*: 0..4 | array with count | 0x80..0x84 |
| binary | *size*: 5 or more | array (terminated by 0xFE) | 0x85 |
An integer that takes 2 to 4 bytes starts with a UTF-8 lead byte (0xC2..0xF7) that is followed by a byte that cannot
continue a UTF-8 character: 0x00..0x7F for positive and 0xC0..0xFF for negative integers. A string is terminated by
0xFF only if it is empty, if another string follows it, or if it is the last value of the message; otherwise, the first
byte of the next value ends it.
!!! success "Complete mapping"
Except for the values listed below, any JSON value can be converted to a BON8 value.
Any BON8 output created by `to_bon8` can be successfully parsed by `from_bon8`.
!!! warning "Unsupported values"
The following values can **not** be converted to a BON8 value:
- unsigned integers above 9223372036854775807, because BON8 has no unsigned 64-bit integer type
([out_of_range.407](../../home/exceptions.md#jsonexceptionout_of_range407))
- strings that are not valid UTF-8, because the end of a string is determined from its encoding
([type_error.316](../../home/exceptions.md#jsonexceptiontype_error316))
!!! info "NaN/infinity handling"
`-0.0`, `Infinity`, and `-Infinity` are serialized as binary32 (type 0x8E, 5 bytes total). `NaN` is serialized as
the binary32 value 0x7F800001 that the specification recommends. This is in contrast to the
[dump](../../api/basic_json/dump.md) function which serializes NaN or Infinity to `null`.
!!! warning "Binary values"
BON8 has no binary type. Binary values are serialized as arrays of integers (0..255), so they are read back as
arrays. The subtype is not serialized.
!!! info "Canonical representation"
The output follows the specification's canonical representation rules: every value uses the shortest encoding,
floating-point numbers use binary32 whenever that loses no precision, and object keys are sorted by their UTF-8
code units. There are two exceptions:
- Strings are not normalized to Unicode Normalization Form C (NFC).
- Object keys are written in the order of the object type, which is sorted for `json`, but not for
[`ordered_json`](../../api/ordered_json.md).
??? example
```cpp
--8<-- "examples/to_bon8.cpp"
```
Output:
```c
--8<-- "examples/to_bon8.output"
```
## Deserialization
The library maps BON8 types to JSON value types as follows:
| BON8 type | JSON value type | first byte |
|-------------------------------|-----------------|------------------------|
| UTF-8 string | string | 0x00..0x7F |
| array with count | array | 0x80..0x84 |
| array (terminated by 0xFE) | array | 0x85 |
| object with count | object | 0x86..0x8A |
| object (terminated by 0xFE) | object | 0x8B |
| int32 | number_unsigned or number_integer | 0x8C |
| int64 | number_unsigned or number_integer | 0x8D |
| binary32 | number_float | 0x8E |
| binary64 | number_float | 0x8F |
| 1-byte positive integer | number_unsigned | 0x90..0xB7 |
| 1-byte negative integer | number_integer | 0xB8..0xC1 |
| UTF-8 string | string | 0xC2..0xF4, followed by 0x80..0xBF |
| 2- to 4-byte positive integer | number_unsigned | 0xC2..0xF7, followed by 0x00..0x7F |
| 2- to 4-byte negative integer | number_integer | 0xC2..0xF7, followed by 0xC0..0xFF |
| false | `false` | 0xF8 |
| true | `true` | 0xF9 |
| null | `null` | 0xFA |
| -1.0 | number_float | 0xFB |
| 0.0 | number_float | 0xFC |
| 1.0 | number_float | 0xFD |
| empty string | string | 0xFF |
Non-negative integers are read as number_unsigned, negative integers as number_integer.
!!! info
Values that do not use the canonical representation, such as integers with a longer encoding than necessary,
arrays and objects with up to four elements that are terminated by 0xFE, unsorted object keys, or a 0xFF after a
string that would also end without it, are accepted. A second 0xFF is not a terminator but an empty string.
Strings must be valid UTF-8, and the last string of a message must be terminated by 0xFF.
!!! info
Any BON8 output created by `to_bon8` can be successfully parsed by `from_bon8`.
??? example
```cpp
--8<-- "examples/from_bon8.cpp"
```
Output:
```json
--8<-- "examples/from_bon8.output"
```
@@ -4,6 +4,7 @@ Though JSON is a ubiquitous data format, it is not a very compact format suitabl
a network. Hence, the library supports
- [BJData](bjdata.md) (Binary JData),
- [BON8](bon8.md) (Binary Object Notation 8),
- [BSON](bson.md) (Binary JSON),
- [CBOR](cbor.md) (Concise Binary Object Representation),
- [MessagePack](messagepack.md), and
@@ -18,6 +19,7 @@ to efficiently encode JSON values to byte vectors and to decode such vectors.
| Format | Serialization | Deserialization |
|-------------|-----------------------------------------------|----------------------------------------------|
| BJData | complete | complete |
| BON8 | incomplete: no unsigned integers above int64 | complete |
| BSON | incomplete: top-level value must be an object | incomplete, but all JSON types are supported |
| CBOR | complete | incomplete, but all JSON types are supported |
| MessagePack | complete | complete |
@@ -28,6 +30,7 @@ to efficiently encode JSON values to byte vectors and to decode such vectors.
| Format | Binary values | Binary subtypes |
|-------------|---------------|-----------------|
| BJData | not supported | not supported |
| BON8 | not supported | not supported |
| BSON | supported | supported |
| CBOR | supported | supported |
| MessagePack | supported | supported |
@@ -42,6 +45,7 @@ See [binary values](../binary_values.md) for more information.
| BJData | 53.2 % | 91.1 % | 78.1 % | 96.6 % |
| BJData (size) | 58.6 % | 92.1 % | 86.7 % | 97.4 % |
| BJData (size+type) | 58.6 % | 92.1 % | 86.5 % | 97.4 % |
| BON8 | 50.5 % | 83.8 % | 63.5 % | 87.5 % |
| BSON | 85.8 % | 95.2 % | 95.8 % | 106.7 % |
| CBOR | 50.5 % | 86.3 % | 68.4 % | 88.0 % |
| MessagePack | 50.5 % | 86.0 % | 68.5 % | 87.9 % |
@@ -187,6 +187,41 @@ as an array of uint8 values. The library implements this translation.
}
```
### BON8
[BON8](binary_formats/bon8.md) neither supports binary values nor subtypes. The library serializes binary values as an
array of integers.
??? example
Code:
```cpp
// create a binary value of subtype 42 (will be ignored in BON8)
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);
// convert to BON8
auto v = json::to_bon8(j);
```
`v` is a `std::vector<std::uint8_t>` with the following 16 elements:
```c
0x87 // object with 1 member
0x62 0x69 0x6E 0x61 0x72 0x79 // "binary"
0x84 // array with 4 elements
0xC3 0x22 0xC3 0x56 0xC3 0x12 0xC3 0x16 // content (each byte as a 2-byte integer)
```
Note that the subtype is lost, and deserializing `v` would yield the following value:
```json
{
"binary": [202, 254, 186, 190]
}
```
### BSON
[BSON](binary_formats/bson.md) supports binary values and subtypes. If a subtype is given, it is used and added as an
+2 -2
View File
@@ -35,8 +35,8 @@ C++ types, and finally serialize it again.
- [Serialization](serialization.md) — turn a value back into JSON text with [`dump`](../api/basic_json/dump.md),
including pretty-printing and handling of non-ASCII and invalid UTF-8.
- [Binary formats](binary_formats/index.md) — encode values more compactly as
[BJData](binary_formats/bjdata.md), [BSON](binary_formats/bson.md), [CBOR](binary_formats/cbor.md),
[MessagePack](binary_formats/messagepack.md), or [UBJSON](binary_formats/ubjson.md).
[BJData](binary_formats/bjdata.md), [BON8](binary_formats/bon8.md), [BSON](binary_formats/bson.md),
[CBOR](binary_formats/cbor.md), [MessagePack](binary_formats/messagepack.md), or [UBJSON](binary_formats/ubjson.md).
- [Binary values](binary_values.md) — store and exchange raw byte sequences.
## How values are stored and configured
+1 -1
View File
@@ -117,7 +117,7 @@ For the [{fmt}](https://github.com/fmtlib/fmt) library, the library ships a
## Serializing to other formats
Besides JSON text, a value can also be serialized to the more compact [binary formats](binary_formats/index.md)
(BJData, BSON, CBOR, MessagePack, UBJSON).
(BJData, BON8, BSON, CBOR, MessagePack, UBJSON).
## See also
@@ -547,7 +547,7 @@ Grisu2 algorithm, which produces the shortest representation that round-trips. O
### Required for the binary formats
`NumberFloatType` must be `#!cpp float` or `#!cpp double`. The writers for
[CBOR, MessagePack, UBJSON, BJData, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
[CBOR, MessagePack, UBJSON, BJData, BON8, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
binary32 or binary64 field and have no encoding for `#!cpp long double`.
### Compatible types