mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-10-04 22:02:52 +01:00
Bit-pack the lexer's token info (#4270)
This makes each token info consist of 8 bytes of data: - 1 byte of the kind - 1 bit for whitespace tracking - 23 bits of payload - 32 bits for byte offset in the file This builds directly on representing the location of the token as a single 32-bit offset, now compressing the rest of the data into a single 32-bit bitfield. This adds some implementation limits: we can no longer lex more than 2^23 tokens in a single source file. Nor can we have more than 2^23 string literals, integer literals, real literals, or identifiers. Only the first of these is even close to an issue, and even then seems unlikely to ever be a problem in practice. The memory efficiency here is great and the motivating goal. But to make this work well, we also need to streamline how we create the tokens. Otherwise, all the bit fiddling can end up erasing our gains. This PR adds a number of APIs to manage creating and accessing the now significantly more complex storage of token infos to try and help with this. One big change required to simplify the writes here is to switch from computing whether a token has trailing space after-the-fact to pre-computing whether a token will have leading space. That lets us have the leading space information available immediately when forming the token, and avoids doing a single bit flip afterward. Another change that helps with this representation is to minimize the updating of groups after-the-fact. The code now tries to set the opening index directly when creating the closing token and only updates the opening group afterward. Because of the bit packing, this is a reduction of 0.5% of dynamic instructions in the compile benchmark, and has dramatic improvements for the grouping symbol focused benchmarks. All combined, this is a significant improvement on the lexer-focused benchmarks despite the added complexity, and a significant win on our compile time benchmarks due to both the lexer improvements and downstream memory density improvements: 5-12% reduction in lex time, growing larger as files get larger. About a 4.5% reduction in parse time, and even a 1-2% reduction in total check time. =D --------- Co-authored-by: Jon Ross-Perkins <jperkins@google.com> Co-authored-by: Geoff Romer <gromer@google.com>
This commit is contained in:
co-authored by
Jon Ross-Perkins
Geoff Romer
parent
187a3608df
commit
c43fa3a8a5
@@ -637,7 +637,7 @@ TEST_F(LexerTest, Whitespace) {
|
||||
auto buffer = Lex("{( } {(");
|
||||
|
||||
// Whether there should be whitespace before/after each token.
|
||||
bool space[] = {true,
|
||||
bool space[] = {false,
|
||||
// start-of-file
|
||||
true,
|
||||
// {
|
||||
@@ -1126,32 +1126,31 @@ TEST_F(LexerTest, PrintingOutputYaml) {
|
||||
Yaml::Value::FromText(print_stream.TakeStr()),
|
||||
IsYaml(ElementsAre(Yaml::Sequence(ElementsAre(Yaml::Mapping(ElementsAre(
|
||||
Pair("filename", source_storage_.front().filename().str()),
|
||||
Pair("tokens",
|
||||
Yaml::Sequence(ElementsAre(
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "0"), Pair("kind", "FileStart"),
|
||||
Pair("line", "1"), Pair("column", "1"),
|
||||
Pair("indent", "1"), Pair("spelling", ""),
|
||||
Pair("has_trailing_space", "true"))),
|
||||
Yaml::Mapping(
|
||||
ElementsAre(Pair("index", "1"), Pair("kind", "Semi"),
|
||||
Pair("line", "2"), Pair("column", "2"),
|
||||
Pair("indent", "2"), Pair("spelling", ";"),
|
||||
Pair("has_trailing_space", "true"))),
|
||||
Yaml::Mapping(
|
||||
ElementsAre(Pair("index", "2"), Pair("kind", "Semi"),
|
||||
Pair("line", "5"), Pair("column", "1"),
|
||||
Pair("indent", "1"), Pair("spelling", ";"),
|
||||
Pair("has_trailing_space", "true"))),
|
||||
Yaml::Mapping(
|
||||
ElementsAre(Pair("index", "3"), Pair("kind", "Semi"),
|
||||
Pair("line", "5"), Pair("column", "3"),
|
||||
Pair("indent", "1"), Pair("spelling", ";"),
|
||||
Pair("has_trailing_space", "true"))),
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "4"), Pair("kind", "FileEnd"),
|
||||
Pair("line", "15"), Pair("column", "1"),
|
||||
Pair("indent", "1"), Pair("spelling", "")))))))))))));
|
||||
Pair("tokens", Yaml::Sequence(ElementsAre(
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "0"), Pair("kind", "FileStart"),
|
||||
Pair("line", "1"), Pair("column", "1"),
|
||||
Pair("indent", "1"), Pair("spelling", ""))),
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "1"), Pair("kind", "Semi"),
|
||||
Pair("line", "2"), Pair("column", "2"),
|
||||
Pair("indent", "2"), Pair("spelling", ";"),
|
||||
Pair("has_leading_space", "true"))),
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "2"), Pair("kind", "Semi"),
|
||||
Pair("line", "5"), Pair("column", "1"),
|
||||
Pair("indent", "1"), Pair("spelling", ";"),
|
||||
Pair("has_leading_space", "true"))),
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "3"), Pair("kind", "Semi"),
|
||||
Pair("line", "5"), Pair("column", "3"),
|
||||
Pair("indent", "1"), Pair("spelling", ";"),
|
||||
Pair("has_leading_space", "true"))),
|
||||
Yaml::Mapping(ElementsAre(
|
||||
Pair("index", "4"), Pair("kind", "FileEnd"),
|
||||
Pair("line", "15"), Pair("column", "1"),
|
||||
Pair("indent", "1"), Pair("spelling", ""),
|
||||
Pair("has_leading_space", "true")))))))))))));
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
Reference in New Issue
Block a user