mirror of
https://github.com/nlohmann/json.git
synced 2026-09-30 22:06:42 +01:00
Add BON8 support (#2998)
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -0,0 +1,159 @@
|
||||
# BON8
|
||||
|
||||
BON8 (Binary Object Notation 8) is a compact binary serialization format for JSON values. It uses the byte values that
|
||||
cannot begin a UTF-8 character as type markers, so strings are stored as plain UTF-8 without a length prefix: a string
|
||||
ends at the first byte that cannot continue it. Integers from -10 to 39, `true`, `false`, `null`, and the floating-point
|
||||
values -1.0, 0.0, and 1.0 take a single byte, and arrays and objects with up to four elements need no terminator.
|
||||
|
||||
!!! abstract "References"
|
||||
|
||||
- [BON8 specification](https://github.com/hikoworks/hikogui/blob/main/docs/BON8.md)
|
||||
- [Reference implementation](https://github.com/hikoworks/hikogui/blob/main/src/hikogui/codec/BON8.hpp) in HikoGUI
|
||||
|
||||
## Serialization
|
||||
|
||||
The library uses the following mapping from JSON values types to BON8 types according to the BON8 specification:
|
||||
|
||||
| JSON value type | value/range | BON8 type | first byte |
|
||||
|-----------------|----------------------------------------------|-------------------------------|------------|
|
||||
| null | `null` | null | 0xFA |
|
||||
| boolean | `true` | true | 0xF9 |
|
||||
| boolean | `false` | false | 0xF8 |
|
||||
| number_integer | -9223372036854775808..-2147483649 | int64 | 0x8D |
|
||||
| number_integer | -2147483648..-33818507 | int32 | 0x8C |
|
||||
| number_integer | -33818506..-264075 | 4-byte negative integer | 0xF0..0xF7 |
|
||||
| number_integer | -264074..-1931 | 3-byte negative integer | 0xE0..0xEF |
|
||||
| number_integer | -1930..-11 | 2-byte negative integer | 0xC2..0xDF |
|
||||
| number_integer | -10..-1 | 1-byte negative integer | 0xB8..0xC1 |
|
||||
| number_integer | 0..39 | 1-byte positive integer | 0x90..0xB7 |
|
||||
| number_integer | 40..3879 | 2-byte positive integer | 0xC2..0xDF |
|
||||
| number_integer | 3880..528167 | 3-byte positive integer | 0xE0..0xEF |
|
||||
| number_integer | 528168..67637031 | 4-byte positive integer | 0xF0..0xF7 |
|
||||
| number_integer | 67637032..2147483647 | int32 | 0x8C |
|
||||
| number_integer | 2147483648..9223372036854775807 | int64 | 0x8D |
|
||||
| number_unsigned | 0..39 | 1-byte positive integer | 0x90..0xB7 |
|
||||
| number_unsigned | 40..3879 | 2-byte positive integer | 0xC2..0xDF |
|
||||
| number_unsigned | 3880..528167 | 3-byte positive integer | 0xE0..0xEF |
|
||||
| number_unsigned | 528168..67637031 | 4-byte positive integer | 0xF0..0xF7 |
|
||||
| number_unsigned | 67637032..2147483647 | int32 | 0x8C |
|
||||
| number_unsigned | 2147483648..9223372036854775807 | int64 | 0x8D |
|
||||
| number_float | `-1.0` | -1.0 | 0xFB |
|
||||
| number_float | `0.0` | 0.0 | 0xFC |
|
||||
| number_float | `1.0` | 1.0 | 0xFD |
|
||||
| number_float | *any other value representable by a float* | binary32 | 0x8E |
|
||||
| number_float | *any value NOT representable by a float* | binary64 | 0x8F |
|
||||
| string | *empty* | end of string | 0xFF |
|
||||
| string | *non-empty* | UTF-8 string | 0x00..0x7F, 0xC2..0xF4 |
|
||||
| array | *size*: 0..4 | array with count | 0x80..0x84 |
|
||||
| array | *size*: 5 or more | array (terminated by 0xFE) | 0x85 |
|
||||
| object | *size*: 0..4 | object with count | 0x86..0x8A |
|
||||
| object | *size*: 5 or more | object (terminated by 0xFE) | 0x8B |
|
||||
| binary | *size*: 0..4 | array with count | 0x80..0x84 |
|
||||
| binary | *size*: 5 or more | array (terminated by 0xFE) | 0x85 |
|
||||
|
||||
An integer that takes 2 to 4 bytes starts with a UTF-8 lead byte (0xC2..0xF7) that is followed by a byte that cannot
|
||||
continue a UTF-8 character: 0x00..0x7F for positive and 0xC0..0xFF for negative integers. A string is terminated by
|
||||
0xFF only if it is empty, if another string follows it, or if it is the last value of the message; otherwise, the first
|
||||
byte of the next value ends it.
|
||||
|
||||
!!! success "Complete mapping"
|
||||
|
||||
Except for the values listed below, any JSON value can be converted to a BON8 value.
|
||||
|
||||
Any BON8 output created by `to_bon8` can be successfully parsed by `from_bon8`.
|
||||
|
||||
!!! warning "Unsupported values"
|
||||
|
||||
The following values can **not** be converted to a BON8 value:
|
||||
|
||||
- unsigned integers above 9223372036854775807, because BON8 has no unsigned 64-bit integer type
|
||||
([out_of_range.407](../../home/exceptions.md#jsonexceptionout_of_range407))
|
||||
- strings that are not valid UTF-8, because the end of a string is determined from its encoding
|
||||
([type_error.316](../../home/exceptions.md#jsonexceptiontype_error316))
|
||||
|
||||
!!! info "NaN/infinity handling"
|
||||
|
||||
`-0.0`, `Infinity`, and `-Infinity` are serialized as binary32 (type 0x8E, 5 bytes total). `NaN` is serialized as
|
||||
the binary32 value 0x7F800001 that the specification recommends. This is in contrast to the
|
||||
[dump](../../api/basic_json/dump.md) function which serializes NaN or Infinity to `null`.
|
||||
|
||||
!!! warning "Binary values"
|
||||
|
||||
BON8 has no binary type. Binary values are serialized as arrays of integers (0..255), so they are read back as
|
||||
arrays. The subtype is not serialized.
|
||||
|
||||
!!! info "Canonical representation"
|
||||
|
||||
The output follows the specification's canonical representation rules: every value uses the shortest encoding,
|
||||
floating-point numbers use binary32 whenever that loses no precision, and object keys are sorted by their UTF-8
|
||||
code units. There are two exceptions:
|
||||
|
||||
- Strings are not normalized to Unicode Normalization Form C (NFC).
|
||||
- Object keys are written in the order of the object type, which is sorted for `json`, but not for
|
||||
[`ordered_json`](../../api/ordered_json.md).
|
||||
|
||||
??? example
|
||||
|
||||
```cpp
|
||||
--8<-- "examples/to_bon8.cpp"
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
```c
|
||||
--8<-- "examples/to_bon8.output"
|
||||
```
|
||||
|
||||
## Deserialization
|
||||
|
||||
The library maps BON8 types to JSON value types as follows:
|
||||
|
||||
| BON8 type | JSON value type | first byte |
|
||||
|-------------------------------|-----------------|------------------------|
|
||||
| UTF-8 string | string | 0x00..0x7F |
|
||||
| array with count | array | 0x80..0x84 |
|
||||
| array (terminated by 0xFE) | array | 0x85 |
|
||||
| object with count | object | 0x86..0x8A |
|
||||
| object (terminated by 0xFE) | object | 0x8B |
|
||||
| int32 | number_unsigned or number_integer | 0x8C |
|
||||
| int64 | number_unsigned or number_integer | 0x8D |
|
||||
| binary32 | number_float | 0x8E |
|
||||
| binary64 | number_float | 0x8F |
|
||||
| 1-byte positive integer | number_unsigned | 0x90..0xB7 |
|
||||
| 1-byte negative integer | number_integer | 0xB8..0xC1 |
|
||||
| UTF-8 string | string | 0xC2..0xF4, followed by 0x80..0xBF |
|
||||
| 2- to 4-byte positive integer | number_unsigned | 0xC2..0xF7, followed by 0x00..0x7F |
|
||||
| 2- to 4-byte negative integer | number_integer | 0xC2..0xF7, followed by 0xC0..0xFF |
|
||||
| false | `false` | 0xF8 |
|
||||
| true | `true` | 0xF9 |
|
||||
| null | `null` | 0xFA |
|
||||
| -1.0 | number_float | 0xFB |
|
||||
| 0.0 | number_float | 0xFC |
|
||||
| 1.0 | number_float | 0xFD |
|
||||
| empty string | string | 0xFF |
|
||||
|
||||
Non-negative integers are read as number_unsigned, negative integers as number_integer.
|
||||
|
||||
!!! info
|
||||
|
||||
Values that do not use the canonical representation, such as integers with a longer encoding than necessary,
|
||||
arrays and objects with up to four elements that are terminated by 0xFE, unsorted object keys, or a 0xFF after a
|
||||
string that would also end without it, are accepted. A second 0xFF is not a terminator but an empty string.
|
||||
|
||||
Strings must be valid UTF-8, and the last string of a message must be terminated by 0xFF.
|
||||
|
||||
!!! info
|
||||
|
||||
Any BON8 output created by `to_bon8` can be successfully parsed by `from_bon8`.
|
||||
|
||||
??? example
|
||||
|
||||
```cpp
|
||||
--8<-- "examples/from_bon8.cpp"
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
```json
|
||||
--8<-- "examples/from_bon8.output"
|
||||
```
|
||||
@@ -4,6 +4,7 @@ Though JSON is a ubiquitous data format, it is not a very compact format suitabl
|
||||
a network. Hence, the library supports
|
||||
|
||||
- [BJData](bjdata.md) (Binary JData),
|
||||
- [BON8](bon8.md) (Binary Object Notation 8),
|
||||
- [BSON](bson.md) (Binary JSON),
|
||||
- [CBOR](cbor.md) (Concise Binary Object Representation),
|
||||
- [MessagePack](messagepack.md), and
|
||||
@@ -18,6 +19,7 @@ to efficiently encode JSON values to byte vectors and to decode such vectors.
|
||||
| Format | Serialization | Deserialization |
|
||||
|-------------|-----------------------------------------------|----------------------------------------------|
|
||||
| BJData | complete | complete |
|
||||
| BON8 | incomplete: no unsigned integers above int64 | complete |
|
||||
| BSON | incomplete: top-level value must be an object | incomplete, but all JSON types are supported |
|
||||
| CBOR | complete | incomplete, but all JSON types are supported |
|
||||
| MessagePack | complete | complete |
|
||||
@@ -28,6 +30,7 @@ to efficiently encode JSON values to byte vectors and to decode such vectors.
|
||||
| Format | Binary values | Binary subtypes |
|
||||
|-------------|---------------|-----------------|
|
||||
| BJData | not supported | not supported |
|
||||
| BON8 | not supported | not supported |
|
||||
| BSON | supported | supported |
|
||||
| CBOR | supported | supported |
|
||||
| MessagePack | supported | supported |
|
||||
@@ -42,6 +45,7 @@ See [binary values](../binary_values.md) for more information.
|
||||
| BJData | 53.2 % | 91.1 % | 78.1 % | 96.6 % |
|
||||
| BJData (size) | 58.6 % | 92.1 % | 86.7 % | 97.4 % |
|
||||
| BJData (size+type) | 58.6 % | 92.1 % | 86.5 % | 97.4 % |
|
||||
| BON8 | 50.5 % | 83.8 % | 63.5 % | 87.5 % |
|
||||
| BSON | 85.8 % | 95.2 % | 95.8 % | 106.7 % |
|
||||
| CBOR | 50.5 % | 86.3 % | 68.4 % | 88.0 % |
|
||||
| MessagePack | 50.5 % | 86.0 % | 68.5 % | 87.9 % |
|
||||
|
||||
@@ -187,6 +187,41 @@ as an array of uint8 values. The library implements this translation.
|
||||
}
|
||||
```
|
||||
|
||||
### BON8
|
||||
|
||||
[BON8](binary_formats/bon8.md) neither supports binary values nor subtypes. The library serializes binary values as an
|
||||
array of integers.
|
||||
|
||||
??? example
|
||||
|
||||
Code:
|
||||
|
||||
```cpp
|
||||
// create a binary value of subtype 42 (will be ignored in BON8)
|
||||
json j;
|
||||
j["binary"] = json::binary({0xCA, 0xFE, 0xBA, 0xBE}, 42);
|
||||
|
||||
// convert to BON8
|
||||
auto v = json::to_bon8(j);
|
||||
```
|
||||
|
||||
`v` is a `std::vector<std::uint8_t>` with the following 16 elements:
|
||||
|
||||
```c
|
||||
0x87 // object with 1 member
|
||||
0x62 0x69 0x6E 0x61 0x72 0x79 // "binary"
|
||||
0x84 // array with 4 elements
|
||||
0xC3 0x22 0xC3 0x56 0xC3 0x12 0xC3 0x16 // content (each byte as a 2-byte integer)
|
||||
```
|
||||
|
||||
Note that the subtype is lost, and deserializing `v` would yield the following value:
|
||||
|
||||
```json
|
||||
{
|
||||
"binary": [202, 254, 186, 190]
|
||||
}
|
||||
```
|
||||
|
||||
### BSON
|
||||
|
||||
[BSON](binary_formats/bson.md) supports binary values and subtypes. If a subtype is given, it is used and added as an
|
||||
|
||||
@@ -35,8 +35,8 @@ C++ types, and finally serialize it again.
|
||||
- [Serialization](serialization.md) — turn a value back into JSON text with [`dump`](../api/basic_json/dump.md),
|
||||
including pretty-printing and handling of non-ASCII and invalid UTF-8.
|
||||
- [Binary formats](binary_formats/index.md) — encode values more compactly as
|
||||
[BJData](binary_formats/bjdata.md), [BSON](binary_formats/bson.md), [CBOR](binary_formats/cbor.md),
|
||||
[MessagePack](binary_formats/messagepack.md), or [UBJSON](binary_formats/ubjson.md).
|
||||
[BJData](binary_formats/bjdata.md), [BON8](binary_formats/bon8.md), [BSON](binary_formats/bson.md),
|
||||
[CBOR](binary_formats/cbor.md), [MessagePack](binary_formats/messagepack.md), or [UBJSON](binary_formats/ubjson.md).
|
||||
- [Binary values](binary_values.md) — store and exchange raw byte sequences.
|
||||
|
||||
## How values are stored and configured
|
||||
|
||||
@@ -117,7 +117,7 @@ For the [{fmt}](https://github.com/fmtlib/fmt) library, the library ships a
|
||||
## Serializing to other formats
|
||||
|
||||
Besides JSON text, a value can also be serialized to the more compact [binary formats](binary_formats/index.md)
|
||||
(BJData, BSON, CBOR, MessagePack, UBJSON).
|
||||
(BJData, BON8, BSON, CBOR, MessagePack, UBJSON).
|
||||
|
||||
## See also
|
||||
|
||||
|
||||
@@ -547,7 +547,7 @@ Grisu2 algorithm, which produces the shortest representation that round-trips. O
|
||||
### Required for the binary formats
|
||||
|
||||
`NumberFloatType` must be `#!cpp float` or `#!cpp double`. The writers for
|
||||
[CBOR, MessagePack, UBJSON, BJData, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
|
||||
[CBOR, MessagePack, UBJSON, BJData, BON8, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
|
||||
binary32 or binary64 field and have no encoding for `#!cpp long double`.
|
||||
|
||||
### Compatible types
|
||||
|
||||
Reference in New Issue
Block a user