Cleanups of SIMD code and document no Arm port. (#3325)

I spent (a lot) of time working to see if there was any profitable way
to port the SIMD code that scans for identifier length to Arm. There
isn't really. =/ While working on these, I made some cleanups to the
SIMD code that seemed worth landing, and added some benchmarks. All this
PR does is the cleanups, benchmarks, and documents that Arm isn't just
waiting to get attention but doesn't really have good options (so far).

For posterity, here are the core techniques I tried:

1) Direct 32-byte SIMD scanning using pair-wise add trees to build
   a 32-bit mask of valid identifier and then `clz` to compute the
   distance. This is a very good analog to the 16-byte SIMD structure
   used on x86-64. The pair-wise summing technique is the one used in
   simdjson for similar purposes.
2) A 16-byte SIMD scanning similar to the x86 version but using `shrn`
   to produce a 64-bit scalar bitmask with 4 bits per byte, and then
   scaling the bit-count distance.
3) Various hybrid versions of (1) and (2) with short scalar scans to
   identify short identifiers before paying the SIMD start-up cost.
4) A much fancier version of (1) that scanned 64-bytes at a time, but
   cached the resulting 64-bit mask and re-used it until exhausted.

Some good background on these techniques on Arm CPUs is in this blog
post:
https://community.arm.com/arm-community-blogs/b/infrastructure-solutions-blog/posts/porting-x86-vector-bitmask-optimizations-to-arm-neon

Sadly, both (1) and (2) were significantly slower than a scalar loop
over the bytes. Even (3) was consistently slower.

The only approach that came close was (4) and it was very *slightly*
slower in typical examples and very *slightly* faster in extremely
difficult cases like huge identifiers.

Ultimately, the only path I see (suggested by Dougall on a Mastodon
discussion of this whole problem space) is to take (4) to the limit of
computing an identifier-or-not bitmask *for the entire source file*
using a deeply throughput optimized routine (maybe as part of the line
scanning). That should be able to manage the high latency you end up
with when handling these patterns in SIMD on Arm.

The good news is that at least the M1 is *so* fast in the byte-scanning
loop that this isn't hurting nearly as much as I feared.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
This commit is contained in:
Chandler Carruth
2023-10-24 08:32:21 +00:00
committed by GitHub
co-authored by Jon Ross-Perkins
parent 8c28a0494e
commit 1b0e2d3a4b
2 changed files with 169 additions and 100 deletions
@@ -451,6 +451,12 @@ BENCHMARK(BM_ValidIdentifiers<1, 1, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<3, 5, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<3, 16, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<12, 64, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<16, 16, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<24, 24, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<32, 32, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<48, 48, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<64, 64, /*Uniform=*/true>);
BENCHMARK(BM_ValidIdentifiers<80, 80, /*Uniform=*/true>);
// Benchmark to stress the lexing of horizontal whitespace. This sets up what is
// nearly a worst-case scenario of short-but-expensive-to-lex tokens with runs