Validate non-ASCII strings with SSSE3 on every x86-64 CPU that has it

The vector UTF-8 check of json_view needed SSSE3 at compile time
(JSON_VIEW_USE_SSSE3 with -mssse3), so default x86-64 builds validated
non-ASCII text one sequence at a time. The check is now compiled for
SSSE3 with a function attribute (GCC 4.9 and later, Clang; MSVC compiles
the intrinsics anyway) and used where CPUID reports SSSE3. The answer is
kept in an atomic that is initialized at compile time, so neither a
guard nor a global constructor is needed. The definitions do not depend
on compiler flags, so there is no ODR issue. JSON_VIEW_USE_SSSE3 now only
skips the CPU check.

On x86-64 Linux (Haswell), twitter.json parses 23% faster with Clang 18
and 25-34% faster with GCC 13, now ahead of yyjson.

Also always inline read_eight_bytes() and parse_eight_digits(): GCC
called both in the number loops of the lexer and of json_view (52 call
sites), which cost about 10% on citm_catalog.json at -O2.

Document that reusing a document with read() avoids the page faults of
a fresh node index (about 40% of a 55 MB parse on x86-64 Linux).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-10-05 20:16:24 +02:00
parent fbacf1f1a8
commit 606c1e710b
9 changed files with 173 additions and 36 deletions
@@ -49,6 +49,11 @@ whether or not the new parse succeeds; take fresh views from [`root()`](root.md)
`input` is borrowed or owned by the same rules as [`parse()`](parse.md#notes); a document can borrow on one call and
own on the next, since ownership is decided freshly each time.
Reusing a document matters most for large inputs: the operating system provides the memory of a fresh node index one
page at a time, and every page costs a page fault the first time it is written. On x86-64 Linux (4 KiB pages), parsing
a 55 MB document into a reused document took about 40 % less time than parsing it into a fresh one. Programs that parse
many documents of similar size should therefore keep one document and call `read()`.
## Examples
??? example
@@ -5,13 +5,17 @@
```
When defined on x86-64, the parser of [`basic_json_document`](../basic_json_document/index.md)
(`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3, 16 bytes at a time, using the "lookup4"
algorithm of [simdjson](https://github.com/simdjson/simdjson). Without it, non-ASCII text is validated one UTF-8
sequence at a time on x86-64; on AArch64, the vector check uses NEON and is always on.
(`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3 without asking the CPU first.
SSSE3 is not part of the x86-64 baseline, so the code must be compiled for it: define the macro only together with a
compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that has it), and only for programs that
run on such CPUs. The same input is accepted or rejected either way; only the speed of non-ASCII text differs.
By default, the parser checks once at run time whether the CPU has SSSE3 (all x86-64 CPUs since about 2011 have it)
and then validates non-ASCII text 16 bytes at a time, using the "lookup4" algorithm of
[simdjson](https://github.com/simdjson/simdjson); on CPUs without SSSE3, it validates one UTF-8 sequence at a time.
The vector check is compiled for SSSE3 with a function attribute (GCC 4.9 and later, Clang), so this needs no compiler
option. With MSVC, the check uses `__cpuid`. On AArch64, the vector check uses NEON and is always on.
Define the macro only together with a compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that
has it), and only for programs that run on such CPUs. It saves the check of the CPU, which costs little. The same
input is accepted or rejected either way; only the speed of non-ASCII text differs.
!!! warning "Define consistently"
+4 -3
View File
@@ -86,8 +86,9 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc
}
/// eight bytes as a little-endian word (compilers fold this into one load on
/// little-endian targets)
inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
/// little-endian targets; always inlined, as GCC otherwise calls it in the
/// number loops)
JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
{
return static_cast<std::uint64_t>(b[0]) | (static_cast<std::uint64_t>(b[1]) << 8u)
| (static_cast<std::uint64_t>(b[2]) << 16u) | (static_cast<std::uint64_t>(b[3]) << 24u)
@@ -96,7 +97,7 @@ inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
}
/// eight bytes as a little-endian word
inline std::uint64_t read_eight_bytes(const char* p) noexcept
JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const char* p) noexcept
{
return read_eight_bytes(reinterpret_cast<const unsigned char*>(p)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
@@ -249,8 +249,9 @@ template<typename FloatType>
using native_float_t = typename std::conditional<std::numeric_limits<FloatType>::digits == 24, float, double>::type;
/// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three
/// multiplications instead of eight (after simdjson and fast_float)
inline std::uint32_t parse_eight_digits(std::uint64_t v) noexcept
/// multiplications instead of eight (after simdjson and fast_float); always
/// inlined, as GCC otherwise calls it in the number loops
JSON_HEDLEY_ALWAYS_INLINE std::uint32_t parse_eight_digits(std::uint64_t v) noexcept
{
v = ((v & 0x0F0F0F0F0F0F0F0Fu) * 2561u) >> 8u;
v = ((v & 0x00FF00FF00FF00FFu) * 6553601u) >> 16u;
@@ -22,5 +22,7 @@
#undef NLOHMANN_VIEW_NEON
#undef NLOHMANN_VIEW_SSE2
#undef NLOHMANN_VIEW_SSSE3
#undef NLOHMANN_VIEW_SSSE3_DISPATCH
#undef NLOHMANN_VIEW_SSSE3_TARGET
#undef NLOHMANN_VIEW_VECTOR
#undef NLOHMANN_VIEW_VECTOR_UTF8
+9 -3
View File
@@ -121,9 +121,15 @@ stop:
return p; // quote, backslash, or control character
}
#if NLOHMANN_VIEW_VECTOR_UTF8
// non-ASCII: the vector check, out of line
return scan_string_vector(p, e, plain);
#else
#if NLOHMANN_VIEW_SSSE3_DISPATCH
if (NLOHMANN_VIEW_LIKELY(cpu_has_ssse3()))
#endif
{
// non-ASCII: the vector check, out of line
return scan_string_vector(p, e, plain);
}
#endif
#if !NLOHMANN_VIEW_VECTOR_UTF8 || NLOHMANN_VIEW_SSSE3_DISPATCH
// non-ASCII: a run of well-formed sequences (the library's check, so
// that exactly what json::parse accepts is accepted)
do
+61 -7
View File
@@ -10,6 +10,7 @@
#pragma once
#include <array> // array
#include <atomic> // atomic
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint64_t
@@ -18,11 +19,13 @@
// Vector code for long runs of string bytes. NEON (AArch64) and SSE2 (x86-64)
// belong to the baseline instruction sets and are used by default. The vector
// UTF-8 check needs NEON, or SSSE3 if JSON_VIEW_USE_SSSE3 is defined: SSSE3 is
// not part of x86-64, so it must not depend on the flags of a translation unit
// (two translation units with different flags would have different
// definitions of the same inline functions). JSON_VIEW_NO_SIMD selects the
// portable code.
// UTF-8 check needs NEON or SSSE3. SSSE3 is not part of x86-64, and the code
// must not depend on the flags of a translation unit (two translation units
// with different flags would have different definitions of the same inline
// functions): the check is compiled for SSSE3 with a function attribute and
// used where the CPU has SSSE3 (all x86-64 CPUs since about 2011), else the
// portable check. JSON_VIEW_USE_SSSE3 skips the CPU check (for code compiled
// for SSSE3 anyway); JSON_VIEW_NO_SIMD selects the portable code.
#if !defined(JSON_VIEW_NO_SIMD) && defined(__aarch64__) && (defined(__GNUC__) || defined(__clang__)) && NLOHMANN_VIEW_LITTLE_ENDIAN
#include <arm_neon.h>
#define NLOHMANN_VIEW_NEON 1
@@ -41,8 +44,24 @@
#else
#define NLOHMANN_VIEW_SSSE3 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#endif
#if NLOHMANN_VIEW_SSE2 && !NLOHMANN_VIEW_SSSE3 && ((defined(__clang__) && __clang_major__ >= 4) || (defined(__GNUC__) && !defined(__clang__) && (__GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9))))
// (GCC before 4.9 has no SSSE3 intrinsics without -mssse3)
#include <cpuid.h>
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3_DISPATCH 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET __attribute__((target("ssse3")))
#elif NLOHMANN_VIEW_SSE2 && !NLOHMANN_VIEW_SSSE3 && defined(_MSC_VER)
// (MSVC compiles intrinsics of any instruction set)
#include <intrin.h>
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3_DISPATCH 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET
#else
#define NLOHMANN_VIEW_SSSE3_DISPATCH 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET
#endif
#define NLOHMANN_VIEW_VECTOR (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSE2)
#define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3)
#define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3 || NLOHMANN_VIEW_SSSE3_DISPATCH)
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -90,6 +109,40 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* vector_plain_run(const unsigned
}
#endif
#if NLOHMANN_VIEW_SSSE3_DISPATCH
/// whether the CPU has SSSE3 (CPUID leaf 1, ECX bit 9)
inline bool cpu_ssse3() noexcept
{
#if defined(_MSC_VER) && !defined(__clang__)
std::array<int, 4> regs {{}};
__cpuid(regs.data(), 1);
return (static_cast<unsigned>(regs[2]) & (1u << 9u)) != 0;
#else
unsigned eax = 0;
unsigned ebx = 0;
unsigned ecx = 0;
unsigned edx = 0;
return __get_cpuid(1, &eax, &ebx, &ecx, &edx) != 0 && (ecx & (1u << 9u)) != 0;
#endif
}
/// whether the CPU has SSSE3, asked once: the answer is kept in an atomic
/// that is initialized at compile time, so that neither a guard of a local
/// static nor a global constructor is needed (threads that ask at the same
/// time all store the same answer)
NLOHMANN_VIEW_ALWAYS_INLINE bool cpu_has_ssse3() noexcept
{
static std::atomic<int> known{0}; // 0: not asked yet, 1: no, 2: yes
int state = known.load(std::memory_order_relaxed);
if (NLOHMANN_VIEW_UNLIKELY(state == 0))
{
state = cpu_ssse3() ? 2 : 1;
known.store(state, std::memory_order_relaxed);
}
return state == 2;
}
#endif
#if NLOHMANN_VIEW_VECTOR_UTF8
/// Tables of the UTF-8 check of J. Keiser and D. Lemire, "Validating UTF-8 In
/// Less Than One Instruction Per Byte" (2021), as in simdjson ("lookup4"): each
@@ -192,8 +245,9 @@ compares, and the UTF-8 check covers the bytes up to it. Returns where the
string scan stops, like scan_string_run: before ill-formed UTF-8 and for the
last bytes of the input, the bytes are checked one sequence at a time. Out of
line, so that no constants of the check occupy registers in the parse loop.
On x86-64, it is compiled for SSSE3 (see cpu_has_ssse3()).
*/
NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept
NLOHMANN_VIEW_SSSE3_TARGET NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept
{
using lookup = utf8_lookup4<>;
const unsigned char* block = p;
+7 -5
View File
@@ -8860,8 +8860,9 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc
}
/// eight bytes as a little-endian word (compilers fold this into one load on
/// little-endian targets)
inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
/// little-endian targets; always inlined, as GCC otherwise calls it in the
/// number loops)
JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
{
return static_cast<std::uint64_t>(b[0]) | (static_cast<std::uint64_t>(b[1]) << 8u)
| (static_cast<std::uint64_t>(b[2]) << 16u) | (static_cast<std::uint64_t>(b[3]) << 24u)
@@ -8870,7 +8871,7 @@ inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
}
/// eight bytes as a little-endian word
inline std::uint64_t read_eight_bytes(const char* p) noexcept
JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const char* p) noexcept
{
return read_eight_bytes(reinterpret_cast<const unsigned char*>(p)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
@@ -9479,8 +9480,9 @@ template<typename FloatType>
using native_float_t = typename std::conditional<std::numeric_limits<FloatType>::digits == 24, float, double>::type;
/// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three
/// multiplications instead of eight (after simdjson and fast_float)
inline std::uint32_t parse_eight_digits(std::uint64_t v) noexcept
/// multiplications instead of eight (after simdjson and fast_float); always
/// inlined, as GCC otherwise calls it in the number loops
JSON_HEDLEY_ALWAYS_INLINE std::uint32_t parse_eight_digits(std::uint64_t v) noexcept
{
v = ((v & 0x0F0F0F0F0F0F0F0Fu) * 2561u) >> 8u;
v = ((v & 0x00FF00FF00FF00FFu) * 6553601u) >> 16u;
+72 -10
View File
@@ -542,6 +542,7 @@ NLOHMANN_JSON_NAMESPACE_END
#include <array> // array
#include <atomic> // atomic
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint64_t
@@ -551,11 +552,13 @@ NLOHMANN_JSON_NAMESPACE_END
// Vector code for long runs of string bytes. NEON (AArch64) and SSE2 (x86-64)
// belong to the baseline instruction sets and are used by default. The vector
// UTF-8 check needs NEON, or SSSE3 if JSON_VIEW_USE_SSSE3 is defined: SSSE3 is
// not part of x86-64, so it must not depend on the flags of a translation unit
// (two translation units with different flags would have different
// definitions of the same inline functions). JSON_VIEW_NO_SIMD selects the
// portable code.
// UTF-8 check needs NEON or SSSE3. SSSE3 is not part of x86-64, and the code
// must not depend on the flags of a translation unit (two translation units
// with different flags would have different definitions of the same inline
// functions): the check is compiled for SSSE3 with a function attribute and
// used where the CPU has SSSE3 (all x86-64 CPUs since about 2011), else the
// portable check. JSON_VIEW_USE_SSSE3 skips the CPU check (for code compiled
// for SSSE3 anyway); JSON_VIEW_NO_SIMD selects the portable code.
#if !defined(JSON_VIEW_NO_SIMD) && defined(__aarch64__) && (defined(__GNUC__) || defined(__clang__)) && NLOHMANN_VIEW_LITTLE_ENDIAN
#include <arm_neon.h>
#define NLOHMANN_VIEW_NEON 1
@@ -574,8 +577,24 @@ NLOHMANN_JSON_NAMESPACE_END
#else
#define NLOHMANN_VIEW_SSSE3 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#endif
#if NLOHMANN_VIEW_SSE2 && !NLOHMANN_VIEW_SSSE3 && ((defined(__clang__) && __clang_major__ >= 4) || (defined(__GNUC__) && !defined(__clang__) && (__GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9))))
// (GCC before 4.9 has no SSSE3 intrinsics without -mssse3)
#include <cpuid.h>
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3_DISPATCH 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET __attribute__((target("ssse3")))
#elif NLOHMANN_VIEW_SSE2 && !NLOHMANN_VIEW_SSSE3 && defined(_MSC_VER)
// (MSVC compiles intrinsics of any instruction set)
#include <intrin.h>
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3_DISPATCH 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET
#else
#define NLOHMANN_VIEW_SSSE3_DISPATCH 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET
#endif
#define NLOHMANN_VIEW_VECTOR (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSE2)
#define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3)
#define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3 || NLOHMANN_VIEW_SSSE3_DISPATCH)
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -623,6 +642,40 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* vector_plain_run(const unsigned
}
#endif
#if NLOHMANN_VIEW_SSSE3_DISPATCH
/// whether the CPU has SSSE3 (CPUID leaf 1, ECX bit 9)
inline bool cpu_ssse3() noexcept
{
#if defined(_MSC_VER) && !defined(__clang__)
std::array<int, 4> regs {{}};
__cpuid(regs.data(), 1);
return (static_cast<unsigned>(regs[2]) & (1u << 9u)) != 0;
#else
unsigned eax = 0;
unsigned ebx = 0;
unsigned ecx = 0;
unsigned edx = 0;
return __get_cpuid(1, &eax, &ebx, &ecx, &edx) != 0 && (ecx & (1u << 9u)) != 0;
#endif
}
/// whether the CPU has SSSE3, asked once: the answer is kept in an atomic
/// that is initialized at compile time, so that neither a guard of a local
/// static nor a global constructor is needed (threads that ask at the same
/// time all store the same answer)
NLOHMANN_VIEW_ALWAYS_INLINE bool cpu_has_ssse3() noexcept
{
static std::atomic<int> known{0}; // 0: not asked yet, 1: no, 2: yes
int state = known.load(std::memory_order_relaxed);
if (NLOHMANN_VIEW_UNLIKELY(state == 0))
{
state = cpu_ssse3() ? 2 : 1;
known.store(state, std::memory_order_relaxed);
}
return state == 2;
}
#endif
#if NLOHMANN_VIEW_VECTOR_UTF8
/// Tables of the UTF-8 check of J. Keiser and D. Lemire, "Validating UTF-8 In
/// Less Than One Instruction Per Byte" (2021), as in simdjson ("lookup4"): each
@@ -725,8 +778,9 @@ compares, and the UTF-8 check covers the bytes up to it. Returns where the
string scan stops, like scan_string_run: before ill-formed UTF-8 and for the
last bytes of the input, the bytes are checked one sequence at a time. Out of
line, so that no constants of the check occupy registers in the parse loop.
On x86-64, it is compiled for SSSE3 (see cpu_has_ssse3()).
*/
NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept
NLOHMANN_VIEW_SSSE3_TARGET NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept
{
using lookup = utf8_lookup4<>;
const unsigned char* block = p;
@@ -917,9 +971,15 @@ stop:
return p; // quote, backslash, or control character
}
#if NLOHMANN_VIEW_VECTOR_UTF8
// non-ASCII: the vector check, out of line
return scan_string_vector(p, e, plain);
#else
#if NLOHMANN_VIEW_SSSE3_DISPATCH
if (NLOHMANN_VIEW_LIKELY(cpu_has_ssse3()))
#endif
{
// non-ASCII: the vector check, out of line
return scan_string_vector(p, e, plain);
}
#endif
#if !NLOHMANN_VIEW_VECTOR_UTF8 || NLOHMANN_VIEW_SSSE3_DISPATCH
// non-ASCII: a run of well-formed sequences (the library's check, so
// that exactly what json::parse accepts is accepted)
do
@@ -7891,6 +7951,8 @@ class tuple_element<N, ::nlohmann::detail::view::view_item<View>> // NOLINT(cert
#undef NLOHMANN_VIEW_NEON
#undef NLOHMANN_VIEW_SSE2
#undef NLOHMANN_VIEW_SSSE3
#undef NLOHMANN_VIEW_SSSE3_DISPATCH
#undef NLOHMANN_VIEW_SSSE3_TARGET
#undef NLOHMANN_VIEW_VECTOR
#undef NLOHMANN_VIEW_VECTOR_UTF8