Can specify which tokens are allowed generally, and any additional
tokens that only occur when the parse node has an error.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Don't produce an error diagnostic if any of the information leading to
identifying that error is itself known to be affected by an
already-diagnosed error.
When originally switching to the table dispatch approach we discussed
that it'd be nice to disentangle the monolithic symbol lexing routine
with this as we'll typically have fairly precise dispatch. This is
especially true for grouping symbols, which in Carbon are all
constructively one-character (at this point).
I think this provides a substantial improvement to the clarity of the
code by disentangling the different paths. It also allowed a bunch of
simplifications / clarifications to exactly what the behavior with
closing invalid groups actually involves currently.
This was initially motivated by code organization improvements, and any
performance wins were speculative. However, when benchmarking it
surfaced a problem that hadn't been clear -- we're generating too many
distinct functions here, and the table-based dispatch slows down in the
face of that.
So this PR also includes a fix for that, removing the template-generated
fan-out of dispatch functions for distinct symbols. Instead, we have
a dedicated table to translate one character into the token kinds. This
seems to work quite well, avoiding the huge branch-y structure and just
do fairly cheap table translation & dispatch for all one-character
symbols. Building the table requires the token kinds to be default
constructable, so this also enables that and arranges for the zero-value
kind to be the error kind.
Combined, this is a modest speedup for *non* grouping symbols (3-4%,
a bit noisy). And in some cases it is a huge speedup for grouping
symbols (>10%).
Raw benchmark data with 20 runs before/after -- despite the # of runs,
the grouping symbols benchmarks were frustratingly noisy in non-uniform
ways that couldn't fully be accounted for here. Still, this seems like
an overall improvement.
```
BM_RandomSource 7.98ms ± 2% 7.73ms ± 3% -3.12% (p=0.000 n=18+19)
BM_GroupingSymbols/1/0/0 5.90ms ± 2% 5.82ms ± 4% -1.38% (p=0.001 n=20+20)
BM_GroupingSymbols/2/0/0 5.21ms ± 2% 5.15ms ± 2% -1.14% (p=0.002 n=20+18)
BM_GroupingSymbols/3/0/0 4.42ms ± 2% 4.34ms ± 2% -1.87% (p=0.000 n=19+18)
BM_GroupingSymbols/4/0/0 4.29ms ± 2% 4.38ms ± 5% ~ (p=0.297 n=17+20)
BM_GroupingSymbols/8/0/0 5.09ms ±10% 5.10ms ± 7% ~ (p=0.919 n=18+20)
BM_GroupingSymbols/16/0/0 6.35ms ± 8% 6.29ms ± 6% ~ (p=0.201 n=20+20)
BM_GroupingSymbols/32/0/0 9.88ms ± 2% 9.83ms ± 1% ~ (p=0.167 n=18+20)
BM_GroupingSymbols/0/1/0 5.12ms ± 2% 5.01ms ± 2% -2.14% (p=0.000 n=20+19)
BM_GroupingSymbols/0/2/0 4.01ms ± 2% 3.93ms ± 4% -2.03% (p=0.000 n=20+19)
BM_GroupingSymbols/0/3/0 2.92ms ± 3% 2.81ms ± 2% -3.87% (p=0.000 n=20+19)
BM_GroupingSymbols/0/4/0 2.61ms ± 3% 2.47ms ± 2% -5.30% (p=0.000 n=20+18)
BM_GroupingSymbols/0/8/0 1.77ms ± 3% 1.61ms ± 2% -8.91% (p=0.000 n=18+19)
BM_GroupingSymbols/0/16/0 1.41ms ± 3% 1.16ms ± 4% -17.66% (p=0.000 n=20+20)
BM_GroupingSymbols/0/32/0 1.10ms ± 2% 0.92ms ± 3% -16.36% (p=0.000 n=20+17)
BM_GroupingSymbols/0/0/1 5.09ms ± 2% 5.03ms ± 3% -1.11% (p=0.001 n=20+18)
BM_GroupingSymbols/0/0/2 4.01ms ± 2% 3.91ms ± 2% -2.67% (p=0.000 n=20+18)
BM_GroupingSymbols/0/0/3 2.93ms ± 3% 2.81ms ± 2% -4.23% (p=0.000 n=20+19)
BM_GroupingSymbols/0/0/4 2.59ms ± 2% 2.48ms ± 3% -4.48% (p=0.000 n=20+19)
BM_GroupingSymbols/0/0/8 1.75ms ± 1% 1.62ms ± 3% -7.65% (p=0.000 n=17+19)
BM_GroupingSymbols/0/0/16 1.40ms ± 2% 1.15ms ± 3% -17.67% (p=0.000 n=19+20)
BM_GroupingSymbols/0/0/32 1.10ms ± 2% 0.92ms ± 3% -15.91% (p=0.000 n=20+19)
BM_GroupingSymbols/32/1/0 9.62ms ± 2% 9.65ms ± 2% ~ (p=0.654 n=18+20)
BM_GroupingSymbols/32/2/0 9.41ms ± 2% 9.37ms ± 2% ~ (p=0.095 n=20+19)
BM_GroupingSymbols/32/3/0 9.13ms ± 2% 9.13ms ± 3% ~ (p=0.687 n=19+20)
BM_GroupingSymbols/32/4/0 8.93ms ± 1% 8.87ms ± 2% -0.69% (p=0.010 n=20+18)
BM_GroupingSymbols/32/8/0 8.15ms ± 2% 8.14ms ± 3% ~ (p=0.729 n=19+19)
BM_GroupingSymbols/32/16/0 7.04ms ± 3% 6.92ms ± 1% -1.71% (p=0.000 n=20+18)
BM_GroupingSymbols/32/32/0 5.48ms ± 2% 5.38ms ± 3% -1.81% (p=0.000 n=20+20)
BM_GroupingSymbols/32/32/1 5.39ms ± 2% 5.29ms ± 2% -1.87% (p=0.000 n=19+19)
BM_GroupingSymbols/32/32/2 5.34ms ± 2% 5.21ms ± 1% -2.45% (p=0.000 n=20+18)
BM_GroupingSymbols/32/32/3 5.27ms ± 3% 5.16ms ± 2% -2.18% (p=0.000 n=20+19)
BM_GroupingSymbols/32/32/4 5.21ms ± 2% 5.10ms ± 3% -2.11% (p=0.000 n=19+20)
BM_GroupingSymbols/32/32/8 4.98ms ± 2% 4.83ms ± 2% -2.85% (p=0.000 n=19+19)
BM_GroupingSymbols/32/32/16 4.55ms ± 2% 4.45ms ± 2% -2.25% (p=0.000 n=18+20)
BM_GroupingSymbols/32/32/32 3.95ms ± 2% 3.84ms ± 2% -2.98% (p=0.000 n=19+20)
```
Fix off-by-one error: the terminating `nullptr` in `argv` is not
included in `argc`. This is currently causing crashing tests to crash
again in their crash handler, meaning we don't get a backtrace or
CHECK-failure message.
This isn't really representative of anything, but it should help make it
obvious when the handling of grouping symbols improves or regresses.
Notably, the random source microbenchmark has *no* grouping symbols (in
order to let it be random but always lexically valid), and so it's
especially useful to have something that checks grouping symbols given
their prevalence in realistic source code.
Bug found by fuzzing. Problem was untyped SemIR nodes had an invalid
type id, which was retrieved by `HandlePrefixOperator` and then passed
to `context.GetUnqualifiedType`, ultimately performing an invalid access
in `semantics_ir_->GetNode`.
We prefer to make a placeholder type for functions and namespaces to
remove the need for checking for the untyped case everywhere. Eventually
functions will have their own types, but this approach will be needed
for namespaces (and perhaps other non-first-class entities like unbound
methods and interface members) long term.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
## Summary ##
Restructures the lexer to first scan the entire source text for newlines
and create all the line structures needed. Doing this up-front makes it
easy to produce an optimized version with minimal complexity. Currently,
it leverages the system `memchr`, but even when expanded to handle more
complex cases like CR+LF line endings, being isolated in this way will
result in a significantly simpler implementation. This change improves
the lexing of comment lines significantly by skipping their contents
immediately. The overhead of the pre-scan is unmeasurable in all
realistic benchmarks, and 10-30% in benchmarks consisting almost
entirely of blank lines or comments. The improvement of comment lexing
with average length comment lines mixed with code starts at 20% and goes
up. Regressing blank line handling for non-empty comments seems like the
right tradeoff (by far).
## Background and details ##
One weak point in the lexer implementation were large runs of comments.
While those aren't terribly common, they shouldn't present a hazard to
the lexer performance.
A bit more common is a pattern of comments like the following:
```carbon
// Some method comment here.
fn SomeMethodName(...) -> ...;
// Some other method comment here.
fn SomeOtherMethodName(...) -> ...;
```
Here, the lexer spends an inordinate amount of time getting from the
`\n` after the first semicolon to the `fn` token. It has to skip a blank
line, scan a line, find the `//` comment start, then scan to find the
next `\n`, and then scan horizontal whitespace, etc.
It is tempting to build a scanner *exactly* for this. In fact, I built
one, and I can publish it in a PR if folks are interested in what it
looks like. For x86-64, the PSHUFB trick used for scanning identifiers
technically works. But it is *complicated*. Amazingly so. 150 lines of
very subtle code with subtle performance pitfalls at ever turn. I felt
very uncomfortable submitting it, but we can always go back to it.
Nothing I've come up with quite matches it for sheer speed.
However, most of the complexity and time is spent walking from a `//` to
the end of the line. And *that* is something we can do very simply. In
fact, there is a tuned function for that in libc: `memchr`. Using this
we can build a very fast and much simpler scanner to split lines
up-front. This PR uses that and a carefully crafted fast loop to first
build up all the line info we need. Getting this to be as fast as
possible required some other subtle changes, for example always creating
a line structure that goes from the last `\n` and the end of the file.
We then back up the EOF token to avoid surfacing this to users. The nice
thing is the EOF token isn't part of any hot loop, and so this removes
branches everywhere else at modest complexity.
Once we have that, the rest of the lexer just needs to keep track of its
current line in order to record column offsets. I've taken some care to
try and optimize the lexer's usage of the line structures but there are
more opportunities here I suspect.
Combined, this gets much but not all of the performance of a huge SIMD
scanner for newline-through-to-next-token. For extreme cases (100s of
blank lines or empty comment lines between tokens) the holistic scanner
is of course still much faster, but those don't seem nearly worth the
cost.
I was initially worried about the overhead of taking two passes over the
source text, but in practice I've not been able to measure any
appreciable cost to this with realistic source files. In some cases
benchmarks with no newlines get *faster* because we use a much more
efficient approach to fetch the source text into cache as
a happenstance. And that in turn makes the byte-wise dispatched loop run
faster as it stalls less.
I'm particularly happy with this approach because it seems very clear
how to extend this to support CR+LF, bare CR, and even complex mixtures
without any significant speed cost. That wasn't at all true for the
other approaches explored.
I may try some further PRs to smooth out the last bits of slowness here,
but already this is working excellent for me in practice. My 10mloc test
case is down to 2.3s to lex.
## Raw benchmark data
Using a tool that runs benchmarks before and after and analyzes the
results, the following summarizes the CPU-time impact, each of these for
lexing 100k tokens:
```
BM_ValidKeywords 2.57ms ± 1% 2.58ms ± 0% ~ (p=0.190 n=5+4)
BM_ValidIdentifiers<1, 64, false> 9.24ms ± 4% 9.31ms ± 4% ~ (p=0.421 n=5+5)
BM_ValidIdentifiers<1, 1, true> 3.05ms ± 4% 3.11ms ± 4% ~ (p=0.222 n=5+5)
BM_ValidIdentifiers<3, 5, true> 10.9ms ± 0% 11.1ms ± 1% +1.76% (p=0.016 n=4+5)
BM_ValidIdentifiers<3, 16, true> 11.1ms ± 7% 11.0ms ± 1% ~ (p=0.310 n=5+5)
BM_ValidIdentifiers<12, 64, true> 12.2ms ± 1% 12.3ms ± 2% ~ (p=0.111 n=4+5)
BM_HorizontalWhitespace/1 11.2ms ± 6% 11.1ms ± 2% ~ (p=0.841 n=5+5)
BM_HorizontalWhitespace/4 12.0ms ± 3% 12.0ms ± 2% ~ (p=0.548 n=5+5)
BM_HorizontalWhitespace/16 16.2ms ± 6% 15.9ms ± 8% ~ (p=0.690 n=5+5)
BM_HorizontalWhitespace/64 27.7ms ± 3% 28.4ms ± 3% ~ (p=0.151 n=5+5)
BM_HorizontalWhitespace/128 44.3ms ± 1% 45.6ms ± 6% +3.15% (p=0.032 n=5+5)
BM_RandomSource 7.75ms ± 2% 7.72ms ± 1% ~ (p=1.000 n=5+5)
BM_BlankLines/1 11.7ms ± 1% 12.1ms ± 1% +3.46% (p=0.008 n=5+5)
BM_BlankLines/4 14.0ms ± 2% 15.2ms ± 3% +8.12% (p=0.008 n=5+5)
BM_BlankLines/16 23.5ms ± 2% 31.1ms ± 4% +32.26% (p=0.008 n=5+5)
BM_BlankLines/64 75.3ms ± 1% 81.2ms ± 3% +7.83% (p=0.008 n=5+5)
BM_BlankLines/128 133ms ± 3% 150ms ± 2% +12.74% (p=0.008 n=5+5)
BM_CommentLines/1/0/0 13.1ms ± 0% 13.7ms ± 1% +5.11% (p=0.008 n=5+5)
BM_CommentLines/4/0/0 16.6ms ± 1% 18.2ms ± 4% +9.56% (p=0.008 n=5+5)
BM_CommentLines/128/0/0 169ms ± 4% 182ms ± 1% +7.24% (p=0.008 n=5+5)
BM_CommentLines/1/30/0 18.7ms ± 5% 14.1ms ± 0% -24.84% (p=0.008 n=5+5)
BM_CommentLines/4/30/0 36.5ms ± 6% 20.6ms ± 3% -43.59% (p=0.008 n=5+5)
BM_CommentLines/128/30/0 525ms ± 4% 198ms ± 1% -62.38% (p=0.008 n=5+5)
BM_CommentLines/1/70/0 23.4ms ± 6% 14.7ms ± 2% -37.15% (p=0.008 n=5+5)
BM_CommentLines/4/70/0 53.3ms ± 7% 22.4ms ± 4% -57.99% (p=0.008 n=5+5)
BM_CommentLines/128/70/0 1.05s ± 4% 0.21s ± 2% -80.31% (p=0.008 n=5+5)
BM_CommentLines/1/0/2 14.1ms ± 6% 14.3ms ± 1% ~ (p=0.151 n=5+5)
BM_CommentLines/4/0/2 19.4ms ± 5% 20.1ms ± 1% ~ (p=0.151 n=5+5)
BM_CommentLines/128/0/2 238ms ± 8% 229ms ± 0% ~ (p=0.151 n=5+5)
BM_CommentLines/1/30/2 19.2ms ± 7% 14.6ms ± 1% -23.87% (p=0.008 n=5+5)
BM_CommentLines/4/30/2 40.3ms ±13% 22.3ms ± 4% -44.63% (p=0.008 n=5+5)
BM_CommentLines/128/30/2 568ms ± 7% 254ms ± 3% -55.28% (p=0.008 n=5+5)
BM_CommentLines/1/70/2 23.3ms ± 1% 15.0ms ± 3% -35.61% (p=0.016 n=4+5)
BM_CommentLines/4/70/2 57.2ms ± 9% 24.1ms ± 2% -57.81% (p=0.008 n=5+5)
BM_CommentLines/128/70/2 1.07s ± 0% 0.26s ± 2% -75.51% (p=0.016 n=4+5)
BM_CommentLines/1/0/8 15.9ms ± 7% 16.0ms ± 1% ~ (p=0.151 n=5+5)
BM_CommentLines/4/0/8 24.2ms ± 6% 27.9ms ± 2% +15.36% (p=0.008 n=5+5)
BM_CommentLines/128/0/8 386ms ± 4% 445ms ± 1% +15.28% (p=0.008 n=5+5)
BM_CommentLines/1/30/8 20.6ms ± 5% 16.3ms ± 1% -20.95% (p=0.008 n=5+5)
BM_CommentLines/4/30/8 45.3ms ± 6% 30.3ms ± 3% -32.98% (p=0.008 n=5+5)
BM_CommentLines/128/30/8 699ms ± 3% 477ms ± 3% -31.83% (p=0.008 n=5+5)
BM_CommentLines/1/70/8 25.7ms ± 5% 16.8ms ± 2% -34.67% (p=0.008 n=5+5)
BM_CommentLines/4/70/8 62.0ms ± 4% 31.6ms ± 2% -49.10% (p=0.008 n=5+5)
BM_CommentLines/128/70/8 1.20s ± 2% 0.48s ± 4% -59.60% (p=0.008 n=5+5)
```
The horizontal whitespace benchmark (and all of the non-line-oriented
ones) are noisier than they appear here but do show some improvements
(surprisingly). My guess is that it has a lot to do with system load, as
the advantage is that we're using a vectorized loop to scan the text
first and then doing the byte-dispatched loop. So when the cache is
a bit slower to populate, the vectorized version starts to be faster.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
These benchmarks zero-in and stress test horizontal and vertical
whitespace as well as comment lexing performance. They set up
essentially a worst-case scenario of ramping up whitespace between very
sparse tokens to show how the lexer copes with this.
The horizontal whitespace benchmark is perhaps less important as
frequent runs of 50-characters of horizontal whitespace are relatively
rare already, and likely to be exceedingly rare without trailing
comments. But its good to include for completeness and it shows
reasonably strong performance with the current table-dispatch approach.
The blank line and comment line benchmarks are much more important. Lots
of code is relatively line-sparse, especially API files that are perhaps
the most useful to parse quickly. And many of these are a mixture of
sparse with blank lines and sparse with large comment blocks. The
benchmarks show that there are some serious limits here, even falling
below 100k tokens per second throughput on some of the stress tests
here.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
This generalizes the main lexer benchmark's source generation to be
a bit more comprehensive, specifically including whitespace and
comments. Without these, we're missing a key part of the lexer's
performance.
This also tidies up a bit of the code and adds a more specific
distribution of the different factors in lexing based on analysis of
LLVM's source code. The enhancements to the source statistics script
that helped collect the data here will be in a separate PR.
With this, the output on my AMD cloud instance is:
```
-------------------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second tokens_per_second
-------------------------------------------------------------------------------------------------------------------------
BM_ValidKeywords 2809011 ns 2808912 ns 243 212.616M/s 35.601M/s
BM_ValidIdentifiers<1, 64, false> 11461783 ns 11461652 ns 61 128.499M/s 8.72475M/s
BM_ValidIdentifiers<1, 1, true> 3397293 ns 3397240 ns 208 84.2158M/s 29.4357M/s
BM_ValidIdentifiers<3, 5, true> 14165143 ns 14164962 ns 51 40.3956M/s 7.05967M/s
BM_ValidIdentifiers<3, 16, true> 15283128 ns 15282583 ns 47 71.7623M/s 6.5434M/s
BM_ValidIdentifiers<12, 64, true> 17417323 ns 17417109 ns 41 219.007M/s 5.74148M/s
------------------------------------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second lines_per_second tokens_per_second
------------------------------------------------------------------------------------------------------------------------------------------
BM_RandomSource 8132211 ns 8132227 ns 84 137.223M/s 3.9036M/s 12.2968M/s
BM_SpeedOfLightStrCpy 29610 ns 29608 ns 24631 36.8065G/s 1072.17M/s 3.37745G/s
BM_SpeedOfLightDispatch<1> 2144418 ns 2144421 ns 327 520.387M/s 14.8035M/s 46.6326M/s
BM_SpeedOfLightDispatch<2> 1945954 ns 1945827 ns 351 573.498M/s 16.3144M/s 51.392M/s
BM_SpeedOfLightDispatch<4> 2519565 ns 2519467 ns 292 442.923M/s 12.5999M/s 39.6909M/s
BM_SpeedOfLightDispatch<8> 3011965 ns 3011968 ns 238 370.498M/s 10.5396M/s 33.2009M/s
BM_SpeedOfLightDispatch<16> 4379575 ns 4379579 ns 160 254.803M/s 7.24841M/s 22.8332M/s
BM_SpeedOfLightDispatch<32> 6678423 ns 6678353 ns 102 167.096M/s 4.75342M/s 14.9738M/s
BM_SpeedOfLightDispatch<MaxDispatchTargets> 9373075 ns 9372688 ns 75 119.062M/s 3.38697M/s 10.6693M/s
```
I've compared the profile of the `BM_RandomSource` benchmark with this
change and it largely corresponds to what I expect based on profiling
hand-crafted Carbon inputs. And as you can see, we're closing in on
lexing at least hitting the 10-million-lines-per-second mark. =]
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This is shorter, more closely connected to code using the typed node
types, and avoids using the ambiguous word `Node` in places referring to
typed `SemIR` nodes.
Replace `SemIR::Node::GetAsFoo` and `SemIR::Node::Foo::Make` with
`SemIR::Foo` class that represents a particular kind of node, with named
fields.
Rename `SemIR::IntegerLiteral` and `SemIR::RealLiteral` to
`IntegerValue` / `RealValue` to better reflect their purpose and avoid a
name collision with the corresponding `SemIR` node kinds.
Remove `NodeKind::Invalid` and the `SemIR::Node` default constructor
entirely, as they were not used for anything.
This is pretty rough both in code-quality and what it tries to cover.
The coverage approximation is probably fine as it mostly just gives some
vague idea of the distributions we should aim at when evaluating
performance. We still want performance with unusual distributions to be
good.
Happy to improve the code quality if folks want, and especially with
suggestions on what would help. For this kind of quick scripting, it
seemed "fine" to me.
The big change is to make the lexer helpers have internal linkage,
making all of them easy to inline into single call sites.
Looking at the profile showed several other cases of unfortunate
out-of-line functions. Two were due to the code size produced for checks
-- those are switched to `DCHECK`s to remove that code from optimized
builds. The loss of coverage seems minor.
A last one was closing open groups. This was a surprising routine to be
hot, but it the paths to discover "nothing to do here" were intertwined
into the code. This PR extracts this common trace into a separate
function that delegates to the looping recovery path. This lets the hot
path inline easily.
At this point, for a large lexing benchmark I'm using, 50% of the time
is in the identifier hash table at this point. The remaining
improvements are to actually make some of the hot routines like symbol
lexing and comment lexing faster.
The latest versions of Clang, at least on an ARM mac, enforce that we
not use static sanitizer runtimes. The flag is also a bit frustrating:
it has to be split out of other flags in order to remove it, you can't
just disable it or ignore a warning about it being unused.
This gets things to be build and run again. However, the runtimes (I'm
guessing ones that ship with the OS?) its using dynamically don't
actually work -- at least one check was hitting pretty obvious false
positives. So I've added a set of sanitizer flag workarounds we can
expand as needed to continue to work around the limitations of
sanitizers on this platform.
Together, this restores fastbuild on my ARM macOS with the latest Clang
installed.
This uses the musttail dispatched table approach to drive the entire
lexing. The result is that there is no main lexer loop at all in a
traditional sense, now everything is driven through tail recursive
dispatch on the next byte of the source text.
This should be easy to extend still -- the design pattern is to add
lexer methods for handling specific cases, and then add a dispatch
function to dispatch to them from the table. For example, we can add a
method that handles decoding UTF-8 outside of the ASCII subset and set
the table entries used by non-ASCII initial bytes to dispatch to it.
The performance is already surprisingly good, benchmarks show a modest
improvement across the board. That's despite there still being some
*serious* performance issues that I'll fix in a separate patch. There
are also opportunities to leverage this structure more heavily as needed
by putting more specialized dispatch targets in for specific bytes.
A follow-up PR will re-organize the functions here, as almost all of the
methods on the `Lexer` should become private, but I wanted to keep that
a separate change since it will probably render the diff even more hard
to read than it already is.
A `let` declaration is represented by a `bind_name` node in SemIR:
```carbon-semir
%b: i32 = bind_name "b", %a
```
Because `Check` encounters the pattern before it sees the value, we
first create the `bind_name` node with an unset value and don't add it
to the block. Then, once we've seen and converted the initializer, we
update the `bind_name` to have the value and add it to the current
block, after the initializer code.
The goal of these kinds of benchmarks is to help calibrate other
benchmarks and expectations. They benchmark the underlying hardware
capabilities that we can't avoid, and help illustrate bounds for what is
possible. The term "speed-of-light benchmark" references the aspect of
measuring how fast thing could possible run.
The first is a simple memory bandwidth measurement in the best case
scenario -- using `strcpy` over the buffer. This still does a minimal
number of writes to memory and examines each byte of input to see if it
is null, but can cheat in every way possible to run at the maximum speed
of hardware. To a certain extent, we never expect to get close to this
speed, but it's a good illustration of how much headroom the hardware
has available.
The second is potentially more interesting. This illustrates how fast a
byte-by-byte dispatch loop can potentially be. It uses the technique
that I'm hoping to use in the lexer itself of guaranteed tail recursion
to achieve this with a very small code footprint. The performance of
this technique, even when running in this extremely minimal setting to
establish bounds, is hugely dependent on the number of distinct dispatch
targets, and so the benchmark includes a healthy range to show the range
of performance that we might expect when running in a byte-by-byte mode.
Note that we should expect the lexer to be *faster* than this
"speed-of-light" whenever it is able to lex in larger granules than
byte-wise. But for complex, dense token sequences that force looking at
every byte, this shows the "worst case" "speed-of-light" in a sense.
On my recent AMD cloud VM instance, I get the following results running
the main lexer benchmark with these changes included:
```
-------------------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations bytes_per_second tokens_per_second
-------------------------------------------------------------------------------------------------------------------------
BM_ValidKeywords 3169403 ns 3169283 ns 221 188.44M/s 31.5529M/s
BM_ValidIdentifiers<1, 64, false> 12486725 ns 12486445 ns 51 117.953M/s 8.00868M/s
BM_ValidIdentifiers<1, 1, true> 3950455 ns 3950298 ns 178 72.4252M/s 25.3145M/s
BM_ValidIdentifiers<3, 5, true> 15562294 ns 15561178 ns 45 36.7712M/s 6.42625M/s
BM_ValidIdentifiers<3, 16, true> 16118656 ns 16118374 ns 44 68.0412M/s 6.2041M/s
BM_ValidIdentifiers<12, 64, true> 19116271 ns 19116258 ns 35 199.541M/s 5.23115M/s
BM_ValidMix/10/40 7074336 ns 7073795 ns 93 140.744M/s 14.1367M/s
BM_ValidMix/25/30 6790722 ns 6790006 ns 102 131.793M/s 14.7275M/s
BM_ValidMix/50/20 5960514 ns 5960443 ns 118 112.594M/s 16.7773M/s
BM_ValidMix/75/10 4325546 ns 4325556 ns 159 102.559M/s 23.1184M/s
BM_SpeedOfLightStrCpy 24339 ns 24339 ns 29650 35.9049G/s 4.10858G/s
BM_SpeedOfLightDispatch<1> 1756051 ns 1755800 ns 398 509.668M/s 56.9541M/s
BM_SpeedOfLightDispatch<2> 1611973 ns 1611725 ns 436 555.228M/s 62.0453M/s
BM_SpeedOfLightDispatch<4> 2064280 ns 2063990 ns 326 433.565M/s 48.4498M/s
BM_SpeedOfLightDispatch<8> 2484055 ns 2483946 ns 280 360.263M/s 40.2585M/s
BM_SpeedOfLightDispatch<16> 4550963 ns 4550894 ns 155 196.637M/s 21.9737M/s
BM_SpeedOfLightDispatch<32> 6507077 ns 6507090 ns 107 137.523M/s 15.3679M/s
BM_SpeedOfLightDispatch<MaxDispatchTargets> 9071198 ns 9071217 ns 77 98.6499M/s 11.0239M/s
```
Even though we're not lexing anything in the speed-of-light benchmark,
the tokens-per-second measure is still meaningful because we *generated*
the token stream and know how many tokens we put into it. The dispatch
technique easily exceeds hits 10-million tokens/second, but we need to
do substantially better than that to lex at 10-million lines/second.
Fortunately, when the lexer is consuming more than one-byte tokens,
we're already faster than this. And the bytes-per-second numbers from
all but the worst case dispatch scenario are promising.
Don't assume that the bazel-supplied environment variable TEST_TARGET is
present. We don't actually need it for anything other than providing
feedback to the developer if the test fails.
This makes it a bit easier to reproduce test failures under gdb,
particularly for multi-file tests.
Specifically, rather than nesting them in `Carbon::Testing`, nest them
in `Carbon::Foo` for whatever component they're benchmarking. All our
current benchmarks are lexer benchmarks so its `Carbon::Lex`.
This makes even more sense for benchmarks than unittests I think.
In C, the signature `f()` is distinct from `f(void)` -- dating from K&R
style prototype-less functions. We need to use the `f(void)` form for
the scanner creation function so that it can be called by function
pointer `void *(*)(void)` without UB.
This was causing a local test that includes treesitter to fail with our
fastbuild that enables most sanitizers. With this, we may be able to
enable treesitter on our CI as well, but it at least fixes my local test
runs.
This file group exists to allow a `genquery` rule and a Python test to
verify our non-test dependency graph. We don't actually need to build
the binaries in the file group as part of that. The `genquery` rule
seems to do the right thing -- building it directly doesn't cause the
binaries in the group to be built. But without a manual tag, the group
itself is part of `:all` and thus part of `//...` and part of the rules
that will be built even with PR #3106. A consequence is that any change
to the toolchain causes several other binaries to be built as well
because this file group is in the impacted set. Making it manual should
avoid all of this, and without breaking the actual use from `genquery`.
For example, before this change, in a fully cached build after a `bazel
clean`:
```
> bazel test //bazel/check_deps:all
INFO: Invocation ID: 2d83ebee-4c00-425d-be33-23f42b079614
INFO: Analyzed 3 targets (103 packages loaded, 7137 targets configured).
INFO: Found 2 targets and 1 test target...
INFO: Elapsed time: 4.081s, Critical Path: 2.61s
INFO: 3111 processes: 2796 disk cache hit, 315 internal.
INFO: Build completed successfully, 3111 total actions
```
After this change:
```
> bazel test //bazel/check_deps:al
INFO: Invocation ID: c94089e8-a420-4d3c-9902-134e6b55b297
INFO: Analyzed 2 targets (92 packages loaded, 568 targets configured).
INFO: Found 1 target and 1 test target...
INFO: Elapsed time: 0.700s, Critical Path: 0.01s
INFO: 7 processes: 2 disk cache hit, 5 internal.
INFO: Build completed successfully, 7 total actions
```
While here, re-generate the file group, and fix several issues it
uncovers: mark test utilities as `testonly` and update our LLVM package
allowlist to include `clangd`'s package.
Also add `name_reference_untyped` for references to non-first-class
names without types, which currently covers namespaces and functions.
This improves the fidelity of the SemIR representation, and fixes some
issues where we would use the wrong location for nodes and diagnostics
downstream of a name reference.
We're still missing a representation for dotted name expressions, such
as `Namespace.Function`, and we don't use the `untyped` node as an
operand of any other node yet.
These routines have a regular signature and are used to build a table of
function pointers for fast dispatch. However, the previous approach
relied on lambdas to build these functions which resulted in very hard
to read functions in the profile and backtrace.
In preparation for expanding dispatch to handle (many) more cases in the
lexer and also enabling more aggressive inlining into the dispatch
routines, I wanted to tidy up how they appear.
This PR alone shouldn't have any interesting functional effect, it's
just re-organizing the code.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This removes a (very) hot branch in the lexer where we need to special
case when a token is the first token and can't look at its previous
token. It also seems like a generally nice change to the structure of
both the token buffer and parse tree as there are now bracketing
elements for both ends and we should be able to avoid similar branching
in the future.
Mostly mechanical updates to the lexer and parser code to handle this,
but also needed to special case the location information in the
autoupdate code. And then the usual large body of auto-updated tests.
No benchmark data for this change alone as in isolation and in the
current lexer structure it doesn't make a big difference. But this
branch was particularly difficult to handle when trying to update the
whitespace skipping code to be faster, and so I think it is worth
systematically avoiding the special case here.
Trust semantics to have put them in the right places.
Many parts of lowering still need to be updated to use the value
representation chosen at the semantics layer, but this is an incremental
step towards that.
Minor changes to the blaze command line executed by our autoupdate
scripts:
- Don't change the convenience symlinks. Running autoupdate shouldn't
cause `./bazel-bin/...` to switch to running a different binary.
- Don't produce so much spam. Bazel will still log its build progress if
necessary, and still report compile and runtime errors, but won't
produce half a dozen lines of INFO at the start of the command.
Continued from part 1: #3231. Second step updating
`docs/design/generics/details.md`. There remains some work to
incorporate proposal #2200.
- The biggest changes are incorporating much of the text of proposals:
- #2173
- #2687
- It incorporates changes from proposals:
- #989
- #1178
- #2138
- #2200
- #2360
- #2964
- #3162
- It also updates the text to reflect the latest thinking from leads
issues:
- #996
- #2153 -- most notably deleting the section on `TypeId`.
- Update to rule for prioritization blocks with mixed type structures
from [discussion on
2023-07-18](https://docs.google.com/document/d/1gnJBTfY81fZYvI_QXjwKk1uQHYBNHGqRLI2BS_cYYNQ/edit?resourcekey=0-ql1Q1WvTcDvhycf8LbA9DQ#heading=h.7jxges9ojgy3)
- Adds reference links to proposals, issues, and discussions relevant to
the text.
- Also tries to use more precise language when talking about
implementations, to avoid confusing `impl` declaration and definitions
with the `impls` operator used in `where` clauses, an issue brought up
in
- #2495
- #2483
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Combine the initialization, implicit conversion, and value category
conversion functions into a single function.
This substantially reduces the duplication between these steps, and
ensures that we support the same set of conversions in all these
contexts. This also fixes some issues where we would not use the proper
value representation for tuples and structs after performing implicit
conversions.
This makes the difference between errors and lower-level diagnostics
visible to users, and aligns the toolchain's behavior with the
expectations in `driver_fuzzer.cpp`.
Fix a bug where we would perform the computation of the return location
in SemIR after we have already used it in some cases, leading to
assertion failures during lowering. Instead, accumulate a sequence of
instructions to compute the return location in a temporary block, and
overwrite the return slot with those instructions when we perform
initialization.
StubReference is replaced by a more general SpliceBlock node, that takes
a code block and a result value, executes the instructions in the block,
and produces the result. This is used in the uncommon case where more
than one instruction is required to compute the return slot, which can
happen if we need to first emit a temporary and then index into it, or
if we need to perform multiple levels of indexing before we reach an
entity to initialize.
Includes proposals:
- #990
- #2188
- #2138
- #2200
- #2360
- #2760
- #2964
- #3162
Also tries to use more precise language when talking about:
- implementations, to avoid confusing `impl` declaration and definitions
with the `impls` operator used in `where` clauses, an issue brought up
in #2495 and #2483;
- "binding patterns", like `x: i32`, and "bindings" like `x`.
---------
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Factor out the common code to find the return slot for an initializing
expression, and use it to simplify the two different ways we finalize
initializers, as either initializing a temporary or initializing some
specific object.
This slightly changes the SemIR we create for function calls: instead of
rewriting the Call node to have a different destination and replacing
its temporary with a `no_op`, we now replace its temporary with a
`stub_reference` to the new destination. This results in the same amount
of SemIR being produced, but allows calls and other kinds of
initializers to be handled uniformly.
The speculative insertion of StubReferences after elements in an
argument list turned out to not be necessary, because we decided we want
to insert per-argument initialization steps after all arguments are
evaluated, rather than interleaving them. The StubReferences we insert
are causing some minor code complexity, so remove them.
We still create StubReferences when performing patch-ups of
already-emitted code, but we no longer ever need to look through them
when determining whether an initializer was a literal or when evaluating
a type expression.
This implements initializing expression semantics for structs and
tuples, following #2006 and discussions since.
Tuple and (and analogously, struct) literals are treated as having a
mixed expression category that is later resolved based on how the
literal is used, as either a tuple initializer or a tuple value, at
which point we create a `TupleInit` or `TupleValue` that represents the
formation of the tuple initializer or tuple value from the tuple
literal.
There's quite a lot of TODOs here, and the SemIR representation is still
not quite right, but this seems like a good place to checkpoint some
incremental progress.
First step in updating `docs/design/generics/details.md`. It
incorporates changes from proposals: #989#2138#2173#2200#2360#2964#3162 , but there are still more changes from those proposals to be
made.
It also switches away from suggesting static-dispatch witness tables,
and creates an appendix to describe that decision.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Rationale: this convention avoids forcing closely-related code to be far
apart in the namespace hierarchy, and vice versa. By the same token, it
makes the namespace hierarchy more consistent with the directory
hierarchy.