Files
json/docs/mkdocs/docs/features/binary_values.md
T
Niels Lohmann 1e101ecac1 Add BON8 support (#2998)
* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 16:56:21 +02:00

12 KiB

Binary Values

The library implements several binary formats that encode JSON in an efficient way. Most of these formats support binary values; that is, values that have semantics defined outside the library and only define a sequence of bytes to be stored.

JSON itself does not have a binary value. As such, binary values are an extension that this library implements to store values received by a binary format. Binary values are never created by the JSON parser and are only part of a serialized JSON text if they have been created manually or via a binary format.

API for binary values

classDiagram

class binary_t ["json::binary_t"] {
    +void set_subtype(std::uint64_t subtype)
    +void clear_subtype()
    +std::uint64_t subtype() const
    +bool has_subtype() const
}

class vector ["std::vector<uint8_t>"]

vector <|-- binary_t

By default, binary values are stored as std::vector<std::uint8_t>. This type can be changed by providing a template parameter to the basic_json type. To store binary subtypes, the storage type is extended and exposed as json::binary_t:

auto binary = json::binary_t({0xCA, 0xFE, 0xBA, 0xBE});
auto binary_with_subtype = json::binary_t({0xCA, 0xFE, 0xBA, 0xBE}, 42);

There are several convenience functions to check and set the subtype:

binary.has_subtype();                   // returns false
binary_with_subtype.has_subtype();      // returns true

binary_with_subtype.clear_subtype();
binary_with_subtype.has_subtype();      // returns false

binary_with_subtype.set_subtype(42);
binary.set_subtype(23);

binary.subtype();                       // returns 23

As json::binary_t is subclassing std::vector<std::uint8_t>, all member functions are available:

binary.size();  // returns 4
binary[1];      // returns 0xFE

JSON values can be constructed from json::binary_t:

json j = binary;

Binary values are primitive values just like numbers or strings:

j.is_binary();    // returns true
j.is_primitive(); // returns true

Given a binary JSON value, the binary_t can be accessed by reference as via get_binary():

j.get_binary().has_subtype();  // returns true
j.get_binary().size();         // returns 4

For convenience, binary JSON values can be constructed via json::binary:

auto j2 = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 23);
auto j3 = json::binary({0xCA, 0xFE, 0xBA, 0xBE});

j2 == j;                        // returns true
j3.get_binary().has_subtype();  // returns false
j3.get_binary().subtype();      // returns std::uint64_t(-1) as j3 has no subtype

Serialization

Binary values are serialized differently according to the formats.

JSON

JSON does not have a binary type, and this library does not introduce a new type as this would break conformance. Instead, binary values are serialized as an object with two keys: bytes holds an array of integers, and subtype is an integer or null.

??? example

Code:

```cpp
// create a binary value of subtype 42
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// serialize to standard output
std::cout << j.dump(2) << std::endl;
```

Output:

```json
{
  "binary": {
    "bytes": [202, 254, 186, 190],
    "subtype": 42
  }
}
```

!!! warning "No roundtrip for binary values"

The JSON parser will not parse the objects generated by binary values back to binary values. This is by design to
remain standards compliant. Serializing binary values to JSON is only implemented for debugging purposes.

BJData

BJData neither supports binary values nor subtypes and proposes to serialize binary values as an array of uint8 values. The library implements this translation.

??? example

Code:

```cpp
// create a binary value of subtype 42 (will be ignored in BJData)
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// convert to BJData
auto v = json::to_bjdata(j);      
```
        
`v` is a `std::vector<std::uint8_t>` with the following 20 elements:

```c
0x7B                                             // '{'
    0x69 0x06                                    // i 6 (length of the key)
    0x62 0x69 0x6E 0x61 0x72 0x79                // "binary"
    0x5B                                         // '['
        0x55 0xCA 0x55 0xFE 0x55 0xBA 0x55 0xBE  // content (each byte prefixed with 'U')
    0x5D                                         // ']'
0x7D                                             // '}'
```

The following code uses the type and size optimization for BJData:

```cpp
// convert to BJData using the size and type optimization
auto v = json::to_bjdata(j, true, true);
```

The resulting vector has 22 elements; the optimization is not effective for examples with few values:

```c
0x7B                                // '{'
    0x23 0x69 0x01                  // '#' 'i' type of the array elements: unsigned integers
    0x69 0x06                       // i 6 (length of the key)
    0x62 0x69 0x6E 0x61 0x72 0x79   // "binary"
    0x5B                            // '[' array
        0x24 0x55                   // '$' 'U' type of the array elements: unsigned integers
        0x23 0x69 0x04              // '#' i 4 number of array elements
        0xCA 0xFE 0xBA 0xBE         // content
```

Note that subtype (42) is **not** serialized and that BJData has **no binary type**, and deserializing `v` would
yield the following value:

```json
{
  "binary": [202, 254, 186, 190]
}
```

BON8

BON8 neither supports binary values nor subtypes. The library serializes binary values as an array of integers.

??? example

Code:

```cpp
// create a binary value of subtype 42 (will be ignored in BON8)
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// convert to BON8
auto v = json::to_bon8(j);
```

`v` is a `std::vector<std::uint8_t>` with the following 16 elements:

```c
0x87                                     // object with 1 member
    0x62 0x69 0x6E 0x61 0x72 0x79        // "binary"
    0x84                                 // array with 4 elements
        0xC3 0x22 0xC3 0x56 0xC3 0x12 0xC3 0x16  // content (each byte as a 2-byte integer)
```

Note that the subtype is lost, and deserializing `v` would yield the following value:

```json
{
  "binary": [202, 254, 186, 190]
}
```

BSON

BSON supports binary values and subtypes. If a subtype is given, it is used and added as an unsigned 8-bit integer. If no subtype is given, the generic binary subtype 0x00 is used.

??? example

Code:

```cpp
// create a binary value of subtype 42
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// convert to BSON
auto v = json::to_bson(j);      
```
        
`v` is a `std::vector<std::uint8_t>` with the following 22 elements:

```c
0x16 0x00 0x00 0x00                         // number of bytes in the document
    0x05                                    // binary value
        0x62 0x69 0x6E 0x61 0x72 0x79 0x00  // key "binary" + null byte
        0x04 0x00 0x00 0x00                 // number of bytes
        0x2a                                // subtype
        0xCA 0xFE 0xBA 0xBE                 // content
0x00                                        // end of the document
```

Note that the serialization preserves the subtype, and deserializing `v` would yield the following value:

```json
{
  "binary": {
    "bytes": [202, 254, 186, 190],
    "subtype": 42
  }
}
```

CBOR

CBOR supports binary values, but no subtypes. Subtypes will be serialized as tags. Any binary value will be serialized as byte strings. The library will choose the smallest representation using the length of the byte array.

??? example

Code:

```cpp
// create a binary value of subtype 42
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// convert to CBOR
auto v = json::to_cbor(j);      
```
        
`v` is a `std::vector<std::uint8_t>` with the following 15 elements:

```c
0xA1                                   // map(1)
    0x66                               // text(6)
        0x62 0x69 0x6E 0x61 0x72 0x79  // "binary"
    0xD8 0x2A                          // tag(42)
    0x44                               // bytes(4)
        0xCA 0xFE 0xBA 0xBE            // content
```

Note that the subtype is serialized as tag. However, parsing tagged values yield a parse error unless
`json::cbor_tag_handler_t::ignore` or `json::cbor_tag_handler_t::store` is passed to `json::from_cbor`.

```json
{
  "binary": {
    "bytes": [202, 254, 186, 190],
    "subtype": null
  }
}
```

MessagePack

MessagePack supports binary values and subtypes. If a subtype is given, the ext family is used. The library will choose the smallest representation among fixext1, fixext2, fixext4, fixext8, ext8, ext16, and ext32. The subtype is then added as a signed 8-bit integer.

If no subtype is given, the bin family (bin8, bin16, bin32) is used.

??? example

Code:

```cpp
// create a binary value of subtype 42
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// convert to MessagePack
auto v = json::to_msgpack(j);      
```
        
`v` is a `std::vector<std::uint8_t>` with the following 14 elements:

```c
0x81                                   // fixmap1
    0xA6                               // fixstr6
        0x62 0x69 0x6E 0x61 0x72 0x79  // "binary"
    0xD6                               // fixext4
        0x2A                           // subtype
        0xCA 0xFE 0xBA 0xBE            // content
```

Note that the serialization preserves the subtype, and deserializing `v` would yield the following value:

```json
{
  "binary": {
    "bytes": [202, 254, 186, 190],
    "subtype": 42
  }
}
```

UBJSON

UBJSON neither supports binary values nor subtypes and proposes to serialize binary values as an array of uint8 values. The library implements this translation.

??? example

Code:

```cpp
// create a binary value of subtype 42 (will be ignored in UBJSON)
json j;
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);

// convert to UBJSON
auto v = json::to_ubjson(j);      
```
        
`v` is a `std::vector<std::uint8_t>` with the following 20 elements:

```c
0x7B                                             // '{'
    0x69 0x06                                    // i 6 (length of the key)
    0x62 0x69 0x6E 0x61 0x72 0x79                // "binary"
    0x5B                                         // '['
        0x55 0xCA 0x55 0xFE 0x55 0xBA 0x55 0xBE  // content (each byte prefixed with 'U')
    0x5D                                         // ']'
0x7D                                             // '}'
```

The following code uses the type and size optimization for UBJSON:

```cpp
// convert to UBJSON using the size and type optimization
auto v = json::to_ubjson(j, true, true);
```

The resulting vector has 23 elements; the optimization is not effective for examples with few values:

```c
0x7B                                // '{'
    0x24                            // '$' type of the object elements
    0x5B                            // '[' array
    0x23 0x69 0x01                  // '#' i 1 number of object elements
    0x69 0x06                       // i 6 (length of the key)
    0x62 0x69 0x6E 0x61 0x72 0x79   // "binary"
        0x24 0x55                   // '$' 'U' type of the array elements: unsigned integers
        0x23 0x69 0x04              // '#' i 4 number of array elements
        0xCA 0xFE 0xBA 0xBE         // content
```

Note that subtype (42) is **not** serialized and that UBJSON has **no binary type**, and deserializing `v` would
yield the following value:

```json
{
  "binary": [202, 254, 186, 190]
}
```