Commit Graph
8 Commits
Author SHA1 Message Date
5041a14f59 Add whitespace- and comment-specific benchmarking. (#3276)
These benchmarks zero-in and stress test horizontal and vertical
whitespace as well as comment lexing performance. They set up
essentially a worst-case scenario of ramping up whitespace between very
sparse tokens to show how the lexer copes with this.

The horizontal whitespace benchmark is perhaps less important as
frequent runs of 50-characters of horizontal whitespace are relatively
rare already, and likely to be exceedingly rare without trailing
comments. But its good to include for completeness and it shows
reasonably strong performance with the current table-dispatch approach.

The blank line and comment line benchmarks are much more important. Lots
of code is relatively line-sparse, especially API files that are perhaps
the most useful to parse quickly. And many of these are a mixture of
sparse with blank lines and sparse with large comment blocks. The
benchmarks show that there are some serious limits here, even falling
below 100k tokens per second throughput on some of the stress tests
here.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
2023-10-12 06:54:52 +00:00
Chandler CarruthandRichard Smith 2d735bbc51 Enhance the main lexer benchmark. (#3275)
This generalizes the main lexer benchmark's source generation to be
a bit more comprehensive, specifically including whitespace and
comments. Without these, we're missing a key part of the lexer's
performance.

This also tidies up a bit of the code and adds a more specific
distribution of the different factors in lexing based on analysis of
LLVM's source code. The enhancements to the source statistics script
that helped collect the data here will be in a separate PR.

With this, the output on my AMD cloud instance is:
```
-------------------------------------------------------------------------------------------------------------------------
Benchmark                                            Time             CPU   Iterations bytes_per_second tokens_per_second
-------------------------------------------------------------------------------------------------------------------------
BM_ValidKeywords                               2809011 ns      2808912 ns          243       212.616M/s         35.601M/s
BM_ValidIdentifiers<1, 64, false>             11461783 ns     11461652 ns           61       128.499M/s        8.72475M/s
BM_ValidIdentifiers<1, 1, true>                3397293 ns      3397240 ns          208       84.2158M/s        29.4357M/s
BM_ValidIdentifiers<3, 5, true>               14165143 ns     14164962 ns           51       40.3956M/s        7.05967M/s
BM_ValidIdentifiers<3, 16, true>              15283128 ns     15282583 ns           47       71.7623M/s         6.5434M/s
BM_ValidIdentifiers<12, 64, true>             17417323 ns     17417109 ns           41       219.007M/s        5.74148M/s
------------------------------------------------------------------------------------------------------------------------------------------
Benchmark                                            Time             CPU   Iterations bytes_per_second lines_per_second tokens_per_second
------------------------------------------------------------------------------------------------------------------------------------------
BM_RandomSource                                8132211 ns      8132227 ns           84       137.223M/s        3.9036M/s        12.2968M/s
BM_SpeedOfLightStrCpy                            29610 ns        29608 ns        24631       36.8065G/s       1072.17M/s        3.37745G/s
BM_SpeedOfLightDispatch<1>                     2144418 ns      2144421 ns          327       520.387M/s       14.8035M/s        46.6326M/s
BM_SpeedOfLightDispatch<2>                     1945954 ns      1945827 ns          351       573.498M/s       16.3144M/s         51.392M/s
BM_SpeedOfLightDispatch<4>                     2519565 ns      2519467 ns          292       442.923M/s       12.5999M/s        39.6909M/s
BM_SpeedOfLightDispatch<8>                     3011965 ns      3011968 ns          238       370.498M/s       10.5396M/s        33.2009M/s
BM_SpeedOfLightDispatch<16>                    4379575 ns      4379579 ns          160       254.803M/s       7.24841M/s        22.8332M/s
BM_SpeedOfLightDispatch<32>                    6678423 ns      6678353 ns          102       167.096M/s       4.75342M/s        14.9738M/s
BM_SpeedOfLightDispatch<MaxDispatchTargets>    9373075 ns      9372688 ns           75       119.062M/s       3.38697M/s        10.6693M/s
```

I've compared the profile of the `BM_RandomSource` benchmark with this
change and it largely corresponds to what I expect based on profiling
hand-crafted Carbon inputs. And as you can see, we're closing in on
lexing at least hitting the 10-million-lines-per-second mark. =]

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-10-12 06:35:24 +00:00
Hana Dusíková 48d40aa0f0 Remove unnecessory constexpr for lambdas in tests. (#3272)
After reading Chandler's code I asked him why he has these. Lambdas will
implicitly constexpr anyway.
2023-10-06 07:25:54 +00:00
Chandler Carruth c7e6238fa8 Introduce two speed-of-light benchmarks. (#3270)
The goal of these kinds of benchmarks is to help calibrate other
benchmarks and expectations. They benchmark the underlying hardware
capabilities that we can't avoid, and help illustrate bounds for what is
possible. The term "speed-of-light benchmark" references the aspect of
measuring how fast thing could possible run.

The first is a simple memory bandwidth measurement in the best case
scenario -- using `strcpy` over the buffer. This still does a minimal
number of writes to memory and examines each byte of input to see if it
is null, but can cheat in every way possible to run at the maximum speed
of hardware. To a certain extent, we never expect to get close to this
speed, but it's a good illustration of how much headroom the hardware
has available.

The second is potentially more interesting. This illustrates how fast a
byte-by-byte dispatch loop can potentially be. It uses the technique
that I'm hoping to use in the lexer itself of guaranteed tail recursion
to achieve this with a very small code footprint. The performance of
this technique, even when running in this extremely minimal setting to
establish bounds, is hugely dependent on the number of distinct dispatch
targets, and so the benchmark includes a healthy range to show the range
of performance that we might expect when running in a byte-by-byte mode.
Note that we should expect the lexer to be *faster* than this
"speed-of-light" whenever it is able to lex in larger granules than
byte-wise. But for complex, dense token sequences that force looking at
every byte, this shows the "worst case" "speed-of-light" in a sense.

On my recent AMD cloud VM instance, I get the following results running
the main lexer benchmark with these changes included:

```
-------------------------------------------------------------------------------------------------------------------------
Benchmark                                            Time             CPU   Iterations bytes_per_second tokens_per_second
-------------------------------------------------------------------------------------------------------------------------
BM_ValidKeywords                               3169403 ns      3169283 ns          221        188.44M/s        31.5529M/s
BM_ValidIdentifiers<1, 64, false>             12486725 ns     12486445 ns           51       117.953M/s        8.00868M/s
BM_ValidIdentifiers<1, 1, true>                3950455 ns      3950298 ns          178       72.4252M/s        25.3145M/s
BM_ValidIdentifiers<3, 5, true>               15562294 ns     15561178 ns           45       36.7712M/s        6.42625M/s
BM_ValidIdentifiers<3, 16, true>              16118656 ns     16118374 ns           44       68.0412M/s         6.2041M/s
BM_ValidIdentifiers<12, 64, true>             19116271 ns     19116258 ns           35       199.541M/s        5.23115M/s
BM_ValidMix/10/40                              7074336 ns      7073795 ns           93       140.744M/s        14.1367M/s
BM_ValidMix/25/30                              6790722 ns      6790006 ns          102       131.793M/s        14.7275M/s
BM_ValidMix/50/20                              5960514 ns      5960443 ns          118       112.594M/s        16.7773M/s
BM_ValidMix/75/10                              4325546 ns      4325556 ns          159       102.559M/s        23.1184M/s
BM_SpeedOfLightStrCpy                            24339 ns        24339 ns        29650       35.9049G/s        4.10858G/s
BM_SpeedOfLightDispatch<1>                     1756051 ns      1755800 ns          398       509.668M/s        56.9541M/s
BM_SpeedOfLightDispatch<2>                     1611973 ns      1611725 ns          436       555.228M/s        62.0453M/s
BM_SpeedOfLightDispatch<4>                     2064280 ns      2063990 ns          326       433.565M/s        48.4498M/s
BM_SpeedOfLightDispatch<8>                     2484055 ns      2483946 ns          280       360.263M/s        40.2585M/s
BM_SpeedOfLightDispatch<16>                    4550963 ns      4550894 ns          155       196.637M/s        21.9737M/s
BM_SpeedOfLightDispatch<32>                    6507077 ns      6507090 ns          107       137.523M/s        15.3679M/s
BM_SpeedOfLightDispatch<MaxDispatchTargets>    9071198 ns      9071217 ns           77       98.6499M/s        11.0239M/s
```

Even though we're not lexing anything in the speed-of-light benchmark,
the tokens-per-second measure is still meaningful because we *generated*
the token stream and know how many tokens we put into it. The dispatch
technique easily exceeds hits 10-million tokens/second, but we need to
do substantially better than that to lex at 10-million lines/second.
Fortunately, when the lexer is consuming more than one-byte tokens,
we're already faster than this. And the bytes-per-second numbers from
all but the worst case dispatch scenario are promising.
2023-10-06 00:04:29 +00:00
Chandler Carruth b8802035ed Switch benchmarks to match unit test namespacing. (#3268)
Specifically, rather than nesting them in `Carbon::Testing`, nest them
in `Carbon::Foo` for whatever component they're benchmarking. All our
current benchmarks are lexer benchmarks so its `Carbon::Lex`.

This makes even more sense for benchmarks than unittests I think.
2023-10-05 17:44:31 +00:00
Jon Ross-PerkinsandRichard Smith bc63e6ae0a Switch SourceBuffer to diagnostics. (#3197)
This updates SourceBuffer to diagnostics. Some additional edits to
diagnostics were necessary due to issues moving arguments around, which
seems to stem from a compile error with clang 14 (fixed in later
versions).

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-09-07 23:02:10 +00:00
Jon Ross-Perkins 9ac92ad71b Add support for compiling multiple files at once. (#3182)
Rearranges driver logic into CompilationUnits in order to associate
artifacts from the various stages of compilation.

Note, I'm not totally sure what the right thing to do is for
lower/codegen, so I'm just doing a rote change there for now that
mirrors prior phases (this is all the code supports anyways, so is
probably right for now regardless).

SourceBuffer error output is moved local for consistency with other
steps, and so that it's less ambiguous whether the error should be
expected to already include a filename.
2023-09-06 21:34:59 +00:00
Jon Ross-PerkinsandChandler Carruth ec182fb00d Rename lexer dir to lex (#3179)
Continuing with #3070. Just a dir and file rename (only prefix change is
lexer_file_test). Everything in the lex dir should be marked as a move.

Note, I think this closes #3070. There may still be further cleanup
later, but the organizational changes suggested there are being
completed.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-09-01 02:39:04 +00:00