mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-10-01 16:55:51 +01:00
a79ea4b28d50169d3af9362f868f1aded3fcd049
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a79ea4b28d |
Do more precise table dispatch for symbols. (#3287)
When originally switching to the table dispatch approach we discussed that it'd be nice to disentangle the monolithic symbol lexing routine with this as we'll typically have fairly precise dispatch. This is especially true for grouping symbols, which in Carbon are all constructively one-character (at this point). I think this provides a substantial improvement to the clarity of the code by disentangling the different paths. It also allowed a bunch of simplifications / clarifications to exactly what the behavior with closing invalid groups actually involves currently. This was initially motivated by code organization improvements, and any performance wins were speculative. However, when benchmarking it surfaced a problem that hadn't been clear -- we're generating too many distinct functions here, and the table-based dispatch slows down in the face of that. So this PR also includes a fix for that, removing the template-generated fan-out of dispatch functions for distinct symbols. Instead, we have a dedicated table to translate one character into the token kinds. This seems to work quite well, avoiding the huge branch-y structure and just do fairly cheap table translation & dispatch for all one-character symbols. Building the table requires the token kinds to be default constructable, so this also enables that and arranges for the zero-value kind to be the error kind. Combined, this is a modest speedup for *non* grouping symbols (3-4%, a bit noisy). And in some cases it is a huge speedup for grouping symbols (>10%). Raw benchmark data with 20 runs before/after -- despite the # of runs, the grouping symbols benchmarks were frustratingly noisy in non-uniform ways that couldn't fully be accounted for here. Still, this seems like an overall improvement. ``` BM_RandomSource 7.98ms ± 2% 7.73ms ± 3% -3.12% (p=0.000 n=18+19) BM_GroupingSymbols/1/0/0 5.90ms ± 2% 5.82ms ± 4% -1.38% (p=0.001 n=20+20) BM_GroupingSymbols/2/0/0 5.21ms ± 2% 5.15ms ± 2% -1.14% (p=0.002 n=20+18) BM_GroupingSymbols/3/0/0 4.42ms ± 2% 4.34ms ± 2% -1.87% (p=0.000 n=19+18) BM_GroupingSymbols/4/0/0 4.29ms ± 2% 4.38ms ± 5% ~ (p=0.297 n=17+20) BM_GroupingSymbols/8/0/0 5.09ms ±10% 5.10ms ± 7% ~ (p=0.919 n=18+20) BM_GroupingSymbols/16/0/0 6.35ms ± 8% 6.29ms ± 6% ~ (p=0.201 n=20+20) BM_GroupingSymbols/32/0/0 9.88ms ± 2% 9.83ms ± 1% ~ (p=0.167 n=18+20) BM_GroupingSymbols/0/1/0 5.12ms ± 2% 5.01ms ± 2% -2.14% (p=0.000 n=20+19) BM_GroupingSymbols/0/2/0 4.01ms ± 2% 3.93ms ± 4% -2.03% (p=0.000 n=20+19) BM_GroupingSymbols/0/3/0 2.92ms ± 3% 2.81ms ± 2% -3.87% (p=0.000 n=20+19) BM_GroupingSymbols/0/4/0 2.61ms ± 3% 2.47ms ± 2% -5.30% (p=0.000 n=20+18) BM_GroupingSymbols/0/8/0 1.77ms ± 3% 1.61ms ± 2% -8.91% (p=0.000 n=18+19) BM_GroupingSymbols/0/16/0 1.41ms ± 3% 1.16ms ± 4% -17.66% (p=0.000 n=20+20) BM_GroupingSymbols/0/32/0 1.10ms ± 2% 0.92ms ± 3% -16.36% (p=0.000 n=20+17) BM_GroupingSymbols/0/0/1 5.09ms ± 2% 5.03ms ± 3% -1.11% (p=0.001 n=20+18) BM_GroupingSymbols/0/0/2 4.01ms ± 2% 3.91ms ± 2% -2.67% (p=0.000 n=20+18) BM_GroupingSymbols/0/0/3 2.93ms ± 3% 2.81ms ± 2% -4.23% (p=0.000 n=20+19) BM_GroupingSymbols/0/0/4 2.59ms ± 2% 2.48ms ± 3% -4.48% (p=0.000 n=20+19) BM_GroupingSymbols/0/0/8 1.75ms ± 1% 1.62ms ± 3% -7.65% (p=0.000 n=17+19) BM_GroupingSymbols/0/0/16 1.40ms ± 2% 1.15ms ± 3% -17.67% (p=0.000 n=19+20) BM_GroupingSymbols/0/0/32 1.10ms ± 2% 0.92ms ± 3% -15.91% (p=0.000 n=20+19) BM_GroupingSymbols/32/1/0 9.62ms ± 2% 9.65ms ± 2% ~ (p=0.654 n=18+20) BM_GroupingSymbols/32/2/0 9.41ms ± 2% 9.37ms ± 2% ~ (p=0.095 n=20+19) BM_GroupingSymbols/32/3/0 9.13ms ± 2% 9.13ms ± 3% ~ (p=0.687 n=19+20) BM_GroupingSymbols/32/4/0 8.93ms ± 1% 8.87ms ± 2% -0.69% (p=0.010 n=20+18) BM_GroupingSymbols/32/8/0 8.15ms ± 2% 8.14ms ± 3% ~ (p=0.729 n=19+19) BM_GroupingSymbols/32/16/0 7.04ms ± 3% 6.92ms ± 1% -1.71% (p=0.000 n=20+18) BM_GroupingSymbols/32/32/0 5.48ms ± 2% 5.38ms ± 3% -1.81% (p=0.000 n=20+20) BM_GroupingSymbols/32/32/1 5.39ms ± 2% 5.29ms ± 2% -1.87% (p=0.000 n=19+19) BM_GroupingSymbols/32/32/2 5.34ms ± 2% 5.21ms ± 1% -2.45% (p=0.000 n=20+18) BM_GroupingSymbols/32/32/3 5.27ms ± 3% 5.16ms ± 2% -2.18% (p=0.000 n=20+19) BM_GroupingSymbols/32/32/4 5.21ms ± 2% 5.10ms ± 3% -2.11% (p=0.000 n=19+20) BM_GroupingSymbols/32/32/8 4.98ms ± 2% 4.83ms ± 2% -2.85% (p=0.000 n=19+19) BM_GroupingSymbols/32/32/16 4.55ms ± 2% 4.45ms ± 2% -2.25% (p=0.000 n=18+20) BM_GroupingSymbols/32/32/32 3.95ms ± 2% 3.84ms ± 2% -2.98% (p=0.000 n=19+20) ``` |
||
|
|
6ba8712fbd |
Predetermine all the line splits in the lexer. (#3278)
## Summary ## Restructures the lexer to first scan the entire source text for newlines and create all the line structures needed. Doing this up-front makes it easy to produce an optimized version with minimal complexity. Currently, it leverages the system `memchr`, but even when expanded to handle more complex cases like CR+LF line endings, being isolated in this way will result in a significantly simpler implementation. This change improves the lexing of comment lines significantly by skipping their contents immediately. The overhead of the pre-scan is unmeasurable in all realistic benchmarks, and 10-30% in benchmarks consisting almost entirely of blank lines or comments. The improvement of comment lexing with average length comment lines mixed with code starts at 20% and goes up. Regressing blank line handling for non-empty comments seems like the right tradeoff (by far). ## Background and details ## One weak point in the lexer implementation were large runs of comments. While those aren't terribly common, they shouldn't present a hazard to the lexer performance. A bit more common is a pattern of comments like the following: ```carbon // Some method comment here. fn SomeMethodName(...) -> ...; // Some other method comment here. fn SomeOtherMethodName(...) -> ...; ``` Here, the lexer spends an inordinate amount of time getting from the `\n` after the first semicolon to the `fn` token. It has to skip a blank line, scan a line, find the `//` comment start, then scan to find the next `\n`, and then scan horizontal whitespace, etc. It is tempting to build a scanner *exactly* for this. In fact, I built one, and I can publish it in a PR if folks are interested in what it looks like. For x86-64, the PSHUFB trick used for scanning identifiers technically works. But it is *complicated*. Amazingly so. 150 lines of very subtle code with subtle performance pitfalls at ever turn. I felt very uncomfortable submitting it, but we can always go back to it. Nothing I've come up with quite matches it for sheer speed. However, most of the complexity and time is spent walking from a `//` to the end of the line. And *that* is something we can do very simply. In fact, there is a tuned function for that in libc: `memchr`. Using this we can build a very fast and much simpler scanner to split lines up-front. This PR uses that and a carefully crafted fast loop to first build up all the line info we need. Getting this to be as fast as possible required some other subtle changes, for example always creating a line structure that goes from the last `\n` and the end of the file. We then back up the EOF token to avoid surfacing this to users. The nice thing is the EOF token isn't part of any hot loop, and so this removes branches everywhere else at modest complexity. Once we have that, the rest of the lexer just needs to keep track of its current line in order to record column offsets. I've taken some care to try and optimize the lexer's usage of the line structures but there are more opportunities here I suspect. Combined, this gets much but not all of the performance of a huge SIMD scanner for newline-through-to-next-token. For extreme cases (100s of blank lines or empty comment lines between tokens) the holistic scanner is of course still much faster, but those don't seem nearly worth the cost. I was initially worried about the overhead of taking two passes over the source text, but in practice I've not been able to measure any appreciable cost to this with realistic source files. In some cases benchmarks with no newlines get *faster* because we use a much more efficient approach to fetch the source text into cache as a happenstance. And that in turn makes the byte-wise dispatched loop run faster as it stalls less. I'm particularly happy with this approach because it seems very clear how to extend this to support CR+LF, bare CR, and even complex mixtures without any significant speed cost. That wasn't at all true for the other approaches explored. I may try some further PRs to smooth out the last bits of slowness here, but already this is working excellent for me in practice. My 10mloc test case is down to 2.3s to lex. ## Raw benchmark data Using a tool that runs benchmarks before and after and analyzes the results, the following summarizes the CPU-time impact, each of these for lexing 100k tokens: ``` BM_ValidKeywords 2.57ms ± 1% 2.58ms ± 0% ~ (p=0.190 n=5+4) BM_ValidIdentifiers<1, 64, false> 9.24ms ± 4% 9.31ms ± 4% ~ (p=0.421 n=5+5) BM_ValidIdentifiers<1, 1, true> 3.05ms ± 4% 3.11ms ± 4% ~ (p=0.222 n=5+5) BM_ValidIdentifiers<3, 5, true> 10.9ms ± 0% 11.1ms ± 1% +1.76% (p=0.016 n=4+5) BM_ValidIdentifiers<3, 16, true> 11.1ms ± 7% 11.0ms ± 1% ~ (p=0.310 n=5+5) BM_ValidIdentifiers<12, 64, true> 12.2ms ± 1% 12.3ms ± 2% ~ (p=0.111 n=4+5) BM_HorizontalWhitespace/1 11.2ms ± 6% 11.1ms ± 2% ~ (p=0.841 n=5+5) BM_HorizontalWhitespace/4 12.0ms ± 3% 12.0ms ± 2% ~ (p=0.548 n=5+5) BM_HorizontalWhitespace/16 16.2ms ± 6% 15.9ms ± 8% ~ (p=0.690 n=5+5) BM_HorizontalWhitespace/64 27.7ms ± 3% 28.4ms ± 3% ~ (p=0.151 n=5+5) BM_HorizontalWhitespace/128 44.3ms ± 1% 45.6ms ± 6% +3.15% (p=0.032 n=5+5) BM_RandomSource 7.75ms ± 2% 7.72ms ± 1% ~ (p=1.000 n=5+5) BM_BlankLines/1 11.7ms ± 1% 12.1ms ± 1% +3.46% (p=0.008 n=5+5) BM_BlankLines/4 14.0ms ± 2% 15.2ms ± 3% +8.12% (p=0.008 n=5+5) BM_BlankLines/16 23.5ms ± 2% 31.1ms ± 4% +32.26% (p=0.008 n=5+5) BM_BlankLines/64 75.3ms ± 1% 81.2ms ± 3% +7.83% (p=0.008 n=5+5) BM_BlankLines/128 133ms ± 3% 150ms ± 2% +12.74% (p=0.008 n=5+5) BM_CommentLines/1/0/0 13.1ms ± 0% 13.7ms ± 1% +5.11% (p=0.008 n=5+5) BM_CommentLines/4/0/0 16.6ms ± 1% 18.2ms ± 4% +9.56% (p=0.008 n=5+5) BM_CommentLines/128/0/0 169ms ± 4% 182ms ± 1% +7.24% (p=0.008 n=5+5) BM_CommentLines/1/30/0 18.7ms ± 5% 14.1ms ± 0% -24.84% (p=0.008 n=5+5) BM_CommentLines/4/30/0 36.5ms ± 6% 20.6ms ± 3% -43.59% (p=0.008 n=5+5) BM_CommentLines/128/30/0 525ms ± 4% 198ms ± 1% -62.38% (p=0.008 n=5+5) BM_CommentLines/1/70/0 23.4ms ± 6% 14.7ms ± 2% -37.15% (p=0.008 n=5+5) BM_CommentLines/4/70/0 53.3ms ± 7% 22.4ms ± 4% -57.99% (p=0.008 n=5+5) BM_CommentLines/128/70/0 1.05s ± 4% 0.21s ± 2% -80.31% (p=0.008 n=5+5) BM_CommentLines/1/0/2 14.1ms ± 6% 14.3ms ± 1% ~ (p=0.151 n=5+5) BM_CommentLines/4/0/2 19.4ms ± 5% 20.1ms ± 1% ~ (p=0.151 n=5+5) BM_CommentLines/128/0/2 238ms ± 8% 229ms ± 0% ~ (p=0.151 n=5+5) BM_CommentLines/1/30/2 19.2ms ± 7% 14.6ms ± 1% -23.87% (p=0.008 n=5+5) BM_CommentLines/4/30/2 40.3ms ±13% 22.3ms ± 4% -44.63% (p=0.008 n=5+5) BM_CommentLines/128/30/2 568ms ± 7% 254ms ± 3% -55.28% (p=0.008 n=5+5) BM_CommentLines/1/70/2 23.3ms ± 1% 15.0ms ± 3% -35.61% (p=0.016 n=4+5) BM_CommentLines/4/70/2 57.2ms ± 9% 24.1ms ± 2% -57.81% (p=0.008 n=5+5) BM_CommentLines/128/70/2 1.07s ± 0% 0.26s ± 2% -75.51% (p=0.016 n=4+5) BM_CommentLines/1/0/8 15.9ms ± 7% 16.0ms ± 1% ~ (p=0.151 n=5+5) BM_CommentLines/4/0/8 24.2ms ± 6% 27.9ms ± 2% +15.36% (p=0.008 n=5+5) BM_CommentLines/128/0/8 386ms ± 4% 445ms ± 1% +15.28% (p=0.008 n=5+5) BM_CommentLines/1/30/8 20.6ms ± 5% 16.3ms ± 1% -20.95% (p=0.008 n=5+5) BM_CommentLines/4/30/8 45.3ms ± 6% 30.3ms ± 3% -32.98% (p=0.008 n=5+5) BM_CommentLines/128/30/8 699ms ± 3% 477ms ± 3% -31.83% (p=0.008 n=5+5) BM_CommentLines/1/70/8 25.7ms ± 5% 16.8ms ± 2% -34.67% (p=0.008 n=5+5) BM_CommentLines/4/70/8 62.0ms ± 4% 31.6ms ± 2% -49.10% (p=0.008 n=5+5) BM_CommentLines/128/70/8 1.20s ± 2% 0.48s ± 4% -59.60% (p=0.008 n=5+5) ``` The horizontal whitespace benchmark (and all of the non-line-oriented ones) are noisier than they appear here but do show some improvements (surprisingly). My guess is that it has a lot to do with system load, as the advantage is that we're using a vectorized loop to scan the text first and then doing the byte-dispatched loop. So when the cache is a bit slower to populate, the vectorized version starts to be faster. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> Co-authored-by: josh11b <josh11b@users.noreply.github.com> |
||
|
|
6f5934a505 |
Unblock more lexer inlining. (#3274)
The big change is to make the lexer helpers have internal linkage, making all of them easy to inline into single call sites. Looking at the profile showed several other cases of unfortunate out-of-line functions. Two were due to the code size produced for checks -- those are switched to `DCHECK`s to remove that code from optimized builds. The loss of coverage seems minor. A last one was closing open groups. This was a surprising routine to be hot, but it the paths to discover "nothing to do here" were intertwined into the code. This PR extracts this common trace into a separate function that delegates to the looping recovery path. This lets the hot path inline easily. At this point, for a large lexing benchmark I'm using, 50% of the time is in the identifier hash table at this point. The remaining improvements are to actually make some of the hot routines like symbol lexing and comment lexing faster. |
||
|
|
03c3b86758 |
Switch lexer to fully table-driven design. (#3273)
This uses the musttail dispatched table approach to drive the entire lexing. The result is that there is no main lexer loop at all in a traditional sense, now everything is driven through tail recursive dispatch on the next byte of the source text. This should be easy to extend still -- the design pattern is to add lexer methods for handling specific cases, and then add a dispatch function to dispatch to them from the table. For example, we can add a method that handles decoding UTF-8 outside of the ASCII subset and set the table entries used by non-ASCII initial bytes to dispatch to it. The performance is already surprisingly good, benchmarks show a modest improvement across the board. That's despite there still being some *serious* performance issues that I'll fix in a separate patch. There are also opportunities to leverage this structure more heavily as needed by putting more specialized dispatch targets in for specific bytes. A follow-up PR will re-organize the functions here, as almost all of the methods on the `Lexer` should become private, but I wanted to keep that a separate change since it will probably render the diff even more hard to read than it already is. |
||
|
|
d552545c6d |
Move dispatch routines to be static member functions. (#3265)
These routines have a regular signature and are used to build a table of function pointers for fast dispatch. However, the previous approach relied on lambdas to build these functions which resulted in very hard to read functions in the profile and backtrace. In preparation for expanding dispatch to handle (many) more cases in the lexer and also enabling more aggressive inlining into the dispatch routines, I wanted to tidy up how they appear. This PR alone shouldn't have any interesting functional effect, it's just re-organizing the code. Co-authored-by: Richard Smith <richard@metafoo.co.uk> |
||
|
|
a46ca6bf7a |
Add a start-of-file token and parse node. (#3263)
This removes a (very) hot branch in the lexer where we need to special case when a token is the first token and can't look at its previous token. It also seems like a generally nice change to the structure of both the token buffer and parse tree as there are now bracketing elements for both ends and we should be able to avoid similar branching in the future. Mostly mechanical updates to the lexer and parser code to handle this, but also needed to special case the location information in the autoupdate code. And then the usual large body of auto-updated tests. No benchmark data for this change alone as in isolation and in the current lexer structure it doesn't make a big difference. But this branch was particularly difficult to handle when trying to update the whitespace skipping code to be faster, and so I think it is worth systematically avoiding the special case here. |
||
|
|
2ecab78297 |
Support multi-file lex printing and testing. (#3214)
Lex now prints its yaml as: ``` - filename: name tokens: [ ... ] ``` New support in file_test allows the `filename` marker at the top to define the default file number for later lines, meaning multi-file output from lexing is now associated with the appropriate file. Similar support will probably also apply to lowering, semir, and other places that print a filename once for the full dump. This hammers a bit at how line number replacements work in file_test, allowing stacking them so that lex errors and stdout can both be line-associated properly. I've tried to make the autoupdate more frequently work in one pass, now also taking into account the file index when doing line replacements. There are still some issues with EndOfFile that it may be good to discuss: because CHECK lines are appended to the end of the file now, and the EndOfFile token points at the last line including comments, new lex tests now take two runs to autoupdate (because without CHECK lines, the EndOfFile points at a content line, which content is then inserted after). Note that removing CHECK lines from the test is not a solution: autoupdate also started inserting blank lines, which breaks this for a similar reason. One solution here might be to not have EndOfFile associate with a line or column, which has been a bit of an issue regardless. Also fixes a small issue with toolchain's autoupdate script. |
||
|
|
ec182fb00d |
Rename lexer dir to lex (#3179)
Continuing with #3070. Just a dir and file rename (only prefix change is lexer_file_test). Everything in the lex dir should be marked as a move. Note, I think this closes #3070. There may still be further cleanup later, but the organizational changes suggested there are being completed. --------- Co-authored-by: Chandler Carruth <chandlerc@gmail.com> |