Switch lexer to fully table-driven design. (#3273)

This uses the musttail dispatched table approach to drive the entire
lexing. The result is that there is no main lexer loop at all in a
traditional sense, now everything is driven through tail recursive
dispatch on the next byte of the source text.

This should be easy to extend still -- the design pattern is to add
lexer methods for handling specific cases, and then add a dispatch
function to dispatch to them from the table. For example, we can add a
method that handles decoding UTF-8 outside of the ASCII subset and set
the table entries used by non-ASCII initial bytes to dispatch to it.

The performance is already surprisingly good, benchmarks show a modest
improvement across the board. That's despite there still being some
*serious* performance issues that I'll fix in a separate patch. There
are also opportunities to leverage this structure more heavily as needed
by putting more specialized dispatch targets in for specific bytes.

A follow-up PR will re-organize the functions here, as almost all of the
methods on the `Lexer` should become private, but I wanted to keep that
a separate change since it will probably render the diff even more hard
to read than it already is.
This commit is contained in:
Chandler Carruth
2023-10-07 03:35:24 +00:00
committed by GitHub
parent 48d40aa0f0
commit 03c3b86758
2 changed files with 242 additions and 189 deletions
+3
View File
@@ -571,6 +571,9 @@ TEST_F(LexerTest, Whitespace) {
false};
int pos = 0;
for (Token token : buffer.tokens()) {
SCOPED_TRACE(
llvm::formatv("Token #{0}: '{1}'", token, buffer.GetTokenText(token)));
ASSERT_LT(pos, std::size(space));
EXPECT_THAT(buffer.HasLeadingWhitespace(token), Eq(space[pos]));
++pos;