mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-10-05 22:02:55 +01:00
Currently, for long identifiers, a huge (>30%) fraction of time is spent finding the end of the identifier. We can speed this up with a fun application of SIMD and in-register lookup tables. With this, the BM_ValidIdentifiers/12/64 benchmark goes from around 4 million tokens/second to around 6 mt/s, so roughly 1.5x improvement. However, there was a decent amount of noise in the measurement and I didn't study it too closely as I was very happy with the overall result. The profile shifted from >30% of the time in this loop to <10% of the time, so the scan itself is 3x or more faster with this. One concern with optimizing the lexer right now is that we don't have full Unicode support from the design. This PR takes some steps to at least try and avoid this pitfall -- the new routine works to classify UTF-8 code units, and has a fallback in that case that can grow the needed logic. Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Toolchain
A design is currently maintained in Google Drive. It'll be migrated to markdown once we are confident in its stability.