Similar to #3705, we actually have a mix of `Make` and `Create` in
factory functions too, so this PR is normalizing on `Make`. It's
intended to be consistent with the naming choice for Carbon factory
functions.
Note, MakeSyntheticBlock is the only one I feel a little weird about
because llvm's own APIs use Create, and this is essentially wrapping
LLVM calls. But the flipside is it also feels like a vague line to draw,
when we also differ from LLVM coding style in other ways.
This sets things up to use `bazel` to run `clang-tidy` using
https://github.com/erenon/bazel_clang_tidy.
I'm fixing issues outside of explorer, and disabling clang-tidy for
targets in explorer that have legacy issues. I was going to disable
clang-tidy for targets in explorer such as interpreter anyways, because
they're slow to parse, and just extended that to the currently failing
targets.
I'm looking at this due to the conversation on #3341. Although
diagnostics aren't where they should be, I thought it may help to start
adding raw identifier support (which may also help show how I was
thinking about this).
Note regarding the TODO on how to form the token, `GetTokenText` returns
the `string_id`'s reference value for an `Identifier`. So to make
`GetTokenText` work in a way that returns `r#foo` for a raw identifier,
I think there are a few options:
1. Add additional data indicating the end of the identifier.
2. Add `RawIdentifier` as a token kind to indicate that it's raw and
should be prefixed with `r#` (but also giving later stages one more
token kind to handle)
3. Make the `string_id` correspond to `r#foo`, and have later stages add
`foo` to the strings table whenever `r#foo` is encountered (with map
lookups leading to deduplication).
4. Add `StringId::RawKeyword` special values for each keyword.
- This would mean `self` prints as `self`, `r#self` prints as `r#self`,
but `r#foo` is not a keyword so prints as `foo`.
- This means keywords would need to be listed in a place `StringId` can
depend on them, one way or the other (e.g., a `keywords.def` file in
`base/` should work).
5. Say that it _is_ an `Identifier`, and if it's a keyword spelling, it
must have been a raw identifier.
- Same limitation as above: This would mean `self` prints as `self`,
`r#self` prints as `r#self`, but `r#foo` is not a keyword so prints as
`foo`.
I'm hoping to resolve this issue separately though. :)
I spent (a lot) of time working to see if there was any profitable way
to port the SIMD code that scans for identifier length to Arm. There
isn't really. =/ While working on these, I made some cleanups to the
SIMD code that seemed worth landing, and added some benchmarks. All this
PR does is the cleanups, benchmarks, and documents that Arm isn't just
waiting to get attention but doesn't really have good options (so far).
For posterity, here are the core techniques I tried:
1) Direct 32-byte SIMD scanning using pair-wise add trees to build
a 32-bit mask of valid identifier and then `clz` to compute the
distance. This is a very good analog to the 16-byte SIMD structure
used on x86-64. The pair-wise summing technique is the one used in
simdjson for similar purposes.
2) A 16-byte SIMD scanning similar to the x86 version but using `shrn`
to produce a 64-bit scalar bitmask with 4 bits per byte, and then
scaling the bit-count distance.
3) Various hybrid versions of (1) and (2) with short scalar scans to
identify short identifiers before paying the SIMD start-up cost.
4) A much fancier version of (1) that scanned 64-bytes at a time, but
cached the resulting 64-bit mask and re-used it until exhausted.
Some good background on these techniques on Arm CPUs is in this blog
post:
https://community.arm.com/arm-community-blogs/b/infrastructure-solutions-blog/posts/porting-x86-vector-bitmask-optimizations-to-arm-neon
Sadly, both (1) and (2) were significantly slower than a scalar loop
over the bytes. Even (3) was consistently slower.
The only approach that came close was (4) and it was very *slightly*
slower in typical examples and very *slightly* faster in extremely
difficult cases like huge identifiers.
Ultimately, the only path I see (suggested by Dougall on a Mastodon
discussion of this whole problem space) is to take (4) to the limit of
computing an identifier-or-not bitmask *for the entire source file*
using a deeply throughput optimized routine (maybe as part of the line
scanning). That should be able to manage the high latency you end up
with when handling these patterns in SIMD on Arm.
The good news is that at least the M1 is *so* fast in the byte-scanning
loop that this isn't hurting nearly as much as I feared.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
This updates lexing to use the data. I'll do checking separately, just
to split changes.
Note the ValueStore structure is also set up such that SemIR::File can
use it for other fields.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This isn't really representative of anything, but it should help make it
obvious when the handling of grouping symbols improves or regresses.
Notably, the random source microbenchmark has *no* grouping symbols (in
order to let it be random but always lexically valid), and so it's
especially useful to have something that checks grouping symbols given
their prevalence in realistic source code.
These benchmarks zero-in and stress test horizontal and vertical
whitespace as well as comment lexing performance. They set up
essentially a worst-case scenario of ramping up whitespace between very
sparse tokens to show how the lexer copes with this.
The horizontal whitespace benchmark is perhaps less important as
frequent runs of 50-characters of horizontal whitespace are relatively
rare already, and likely to be exceedingly rare without trailing
comments. But its good to include for completeness and it shows
reasonably strong performance with the current table-dispatch approach.
The blank line and comment line benchmarks are much more important. Lots
of code is relatively line-sparse, especially API files that are perhaps
the most useful to parse quickly. And many of these are a mixture of
sparse with blank lines and sparse with large comment blocks. The
benchmarks show that there are some serious limits here, even falling
below 100k tokens per second throughput on some of the stress tests
here.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
This generalizes the main lexer benchmark's source generation to be
a bit more comprehensive, specifically including whitespace and
comments. Without these, we're missing a key part of the lexer's
performance.
This also tidies up a bit of the code and adds a more specific
distribution of the different factors in lexing based on analysis of
LLVM's source code. The enhancements to the source statistics script
that helped collect the data here will be in a separate PR.
With this, the output on my AMD cloud instance is:
```
-------------------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second tokens_per_second
-------------------------------------------------------------------------------------------------------------------------
BM_ValidKeywords 2809011 ns 2808912 ns 243 212.616M/s 35.601M/s
BM_ValidIdentifiers<1, 64, false> 11461783 ns 11461652 ns 61 128.499M/s 8.72475M/s
BM_ValidIdentifiers<1, 1, true> 3397293 ns 3397240 ns 208 84.2158M/s 29.4357M/s
BM_ValidIdentifiers<3, 5, true> 14165143 ns 14164962 ns 51 40.3956M/s 7.05967M/s
BM_ValidIdentifiers<3, 16, true> 15283128 ns 15282583 ns 47 71.7623M/s 6.5434M/s
BM_ValidIdentifiers<12, 64, true> 17417323 ns 17417109 ns 41 219.007M/s 5.74148M/s
------------------------------------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second lines_per_second tokens_per_second
------------------------------------------------------------------------------------------------------------------------------------------
BM_RandomSource 8132211 ns 8132227 ns 84 137.223M/s 3.9036M/s 12.2968M/s
BM_SpeedOfLightStrCpy 29610 ns 29608 ns 24631 36.8065G/s 1072.17M/s 3.37745G/s
BM_SpeedOfLightDispatch<1> 2144418 ns 2144421 ns 327 520.387M/s 14.8035M/s 46.6326M/s
BM_SpeedOfLightDispatch<2> 1945954 ns 1945827 ns 351 573.498M/s 16.3144M/s 51.392M/s
BM_SpeedOfLightDispatch<4> 2519565 ns 2519467 ns 292 442.923M/s 12.5999M/s 39.6909M/s
BM_SpeedOfLightDispatch<8> 3011965 ns 3011968 ns 238 370.498M/s 10.5396M/s 33.2009M/s
BM_SpeedOfLightDispatch<16> 4379575 ns 4379579 ns 160 254.803M/s 7.24841M/s 22.8332M/s
BM_SpeedOfLightDispatch<32> 6678423 ns 6678353 ns 102 167.096M/s 4.75342M/s 14.9738M/s
BM_SpeedOfLightDispatch<MaxDispatchTargets> 9373075 ns 9372688 ns 75 119.062M/s 3.38697M/s 10.6693M/s
```
I've compared the profile of the `BM_RandomSource` benchmark with this
change and it largely corresponds to what I expect based on profiling
hand-crafted Carbon inputs. And as you can see, we're closing in on
lexing at least hitting the 10-million-lines-per-second mark. =]
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
The goal of these kinds of benchmarks is to help calibrate other
benchmarks and expectations. They benchmark the underlying hardware
capabilities that we can't avoid, and help illustrate bounds for what is
possible. The term "speed-of-light benchmark" references the aspect of
measuring how fast thing could possible run.
The first is a simple memory bandwidth measurement in the best case
scenario -- using `strcpy` over the buffer. This still does a minimal
number of writes to memory and examines each byte of input to see if it
is null, but can cheat in every way possible to run at the maximum speed
of hardware. To a certain extent, we never expect to get close to this
speed, but it's a good illustration of how much headroom the hardware
has available.
The second is potentially more interesting. This illustrates how fast a
byte-by-byte dispatch loop can potentially be. It uses the technique
that I'm hoping to use in the lexer itself of guaranteed tail recursion
to achieve this with a very small code footprint. The performance of
this technique, even when running in this extremely minimal setting to
establish bounds, is hugely dependent on the number of distinct dispatch
targets, and so the benchmark includes a healthy range to show the range
of performance that we might expect when running in a byte-by-byte mode.
Note that we should expect the lexer to be *faster* than this
"speed-of-light" whenever it is able to lex in larger granules than
byte-wise. But for complex, dense token sequences that force looking at
every byte, this shows the "worst case" "speed-of-light" in a sense.
On my recent AMD cloud VM instance, I get the following results running
the main lexer benchmark with these changes included:
```
-------------------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second tokens_per_second
-------------------------------------------------------------------------------------------------------------------------
BM_ValidKeywords 3169403 ns 3169283 ns 221 188.44M/s 31.5529M/s
BM_ValidIdentifiers<1, 64, false> 12486725 ns 12486445 ns 51 117.953M/s 8.00868M/s
BM_ValidIdentifiers<1, 1, true> 3950455 ns 3950298 ns 178 72.4252M/s 25.3145M/s
BM_ValidIdentifiers<3, 5, true> 15562294 ns 15561178 ns 45 36.7712M/s 6.42625M/s
BM_ValidIdentifiers<3, 16, true> 16118656 ns 16118374 ns 44 68.0412M/s 6.2041M/s
BM_ValidIdentifiers<12, 64, true> 19116271 ns 19116258 ns 35 199.541M/s 5.23115M/s
BM_ValidMix/10/40 7074336 ns 7073795 ns 93 140.744M/s 14.1367M/s
BM_ValidMix/25/30 6790722 ns 6790006 ns 102 131.793M/s 14.7275M/s
BM_ValidMix/50/20 5960514 ns 5960443 ns 118 112.594M/s 16.7773M/s
BM_ValidMix/75/10 4325546 ns 4325556 ns 159 102.559M/s 23.1184M/s
BM_SpeedOfLightStrCpy 24339 ns 24339 ns 29650 35.9049G/s 4.10858G/s
BM_SpeedOfLightDispatch<1> 1756051 ns 1755800 ns 398 509.668M/s 56.9541M/s
BM_SpeedOfLightDispatch<2> 1611973 ns 1611725 ns 436 555.228M/s 62.0453M/s
BM_SpeedOfLightDispatch<4> 2064280 ns 2063990 ns 326 433.565M/s 48.4498M/s
BM_SpeedOfLightDispatch<8> 2484055 ns 2483946 ns 280 360.263M/s 40.2585M/s
BM_SpeedOfLightDispatch<16> 4550963 ns 4550894 ns 155 196.637M/s 21.9737M/s
BM_SpeedOfLightDispatch<32> 6507077 ns 6507090 ns 107 137.523M/s 15.3679M/s
BM_SpeedOfLightDispatch<MaxDispatchTargets> 9071198 ns 9071217 ns 77 98.6499M/s 11.0239M/s
```
Even though we're not lexing anything in the speed-of-light benchmark,
the tokens-per-second measure is still meaningful because we *generated*
the token stream and know how many tokens we put into it. The dispatch
technique easily exceeds hits 10-million tokens/second, but we need to
do substantially better than that to lex at 10-million lines/second.
Fortunately, when the lexer is consuming more than one-byte tokens,
we're already faster than this. And the bytes-per-second numbers from
all but the worst case dispatch scenario are promising.
Specifically, rather than nesting them in `Carbon::Testing`, nest them
in `Carbon::Foo` for whatever component they're benchmarking. All our
current benchmarks are lexer benchmarks so its `Carbon::Lex`.
This makes even more sense for benchmarks than unittests I think.
This updates SourceBuffer to diagnostics. Some additional edits to
diagnostics were necessary due to issues moving arguments around, which
seems to stem from a compile error with clang 14 (fixed in later
versions).
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Rearranges driver logic into CompilationUnits in order to associate
artifacts from the various stages of compilation.
Note, I'm not totally sure what the right thing to do is for
lower/codegen, so I'm just doing a rote change there for now that
mirrors prior phases (this is all the code supports anyways, so is
probably right for now regardless).
SourceBuffer error output is moved local for consistency with other
steps, and so that it's less ambiguous whether the error should be
expected to already include a filename.
Continuing with #3070. Just a dir and file rename (only prefix change is
lexer_file_test). Everything in the lex dir should be marked as a move.
Note, I think this closes#3070. There may still be further cleanup
later, but the organizational changes suggested there are being
completed.
---------
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>