Trying to make split file tests of lex functionality shorter and easier
to read. numeric_literals.carbon in particular has an example of why I'm
interested in this (at the bottom). This also switches from `[]` list
format to `-` list format so that the trailing `]` is removed.
Trimming comments in tokenized_buffer.h because (1) it feels like it's
giving too much detail about what's printed, which has drifted slightly
and (2) it also feels like it's trying to justify YAML output, when
that's just what we're doing in general.
---------
Co-authored-by: Geoff Romer <gromer@google.com>
Also surround it in square brackets rather than parentheses. This
matches the format used by Clang and GCC, and means diagnostics will
still match the `file:line:col: error: ` pattern used by some IDE tools.
Before:
```console
fail_builtins.carbon:11:11: error(AliasRequiresNameRef): alias initializer must be a name reference
```
After:
```console
fail_builtins.carbon:11:11: error: alias initializer must be a name reference [AliasRequiresNameRef]
```
Also tighten up test regex to only match on `STDERR` lines that list a
file name.
This is to help identify which diagnostics we're actually using.
Note that driver/testdata still has tests which don't pass this flag,
and so continue to test the kind-less (default) behavior.
This is a primarily automated change:
- Search & replace for capitalization
-
`(CARBON_DIAGNOSTIC\((?:\n\s+)?\w+,(?:\n\s+)?\s\w+,(?:\n\s+)?\s")([A-Z])`
- `$1\L$2`
- Search & replace for period
-
`(CARBON_DIAGNOSTIC\((?:\n\s+)?\w+,(?:\n\s+)?\s\w+,(?:\n\s+)?\s"(?:[^)]|\n)+)\.("[,)])`
- `$1$2`
- Limited search & replace for `ERROR: ` -> `error: ` in streamed things
- Leaving a TODO for command_line because there's more cleanup that can
be done there
- Modify diagnostic_consumer.cpp
- ERROR -> error
- WARNING -> warning
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This makes each token info consist of 8 bytes of data:
- 1 byte of the kind
- 1 bit for whitespace tracking
- 23 bits of payload
- 32 bits for byte offset in the file
This builds directly on representing the location of the token as
a single 32-bit offset, now compressing the rest of the data into
a single 32-bit bitfield.
This adds some implementation limits: we can no longer lex more than
2^23 tokens in a single source file. Nor can we have more than 2^23
string literals, integer literals, real literals, or identifiers. Only
the first of these is even close to an issue, and even then seems
unlikely to ever be a problem in practice.
The memory efficiency here is great and the motivating goal. But to make
this work well, we also need to streamline how we create the tokens.
Otherwise, all the bit fiddling can end up erasing our gains. This PR
adds a number of APIs to manage creating and accessing the now
significantly more complex storage of token infos to try and help with
this.
One big change required to simplify the writes here is to switch from
computing whether a token has trailing space after-the-fact to
pre-computing whether a token will have leading space. That lets us have
the leading space information available immediately when forming the
token, and avoids doing a single bit flip afterward.
Another change that helps with this representation is to minimize the
updating of groups after-the-fact. The code now tries to set the opening
index directly when creating the closing token and only updates the
opening group afterward. Because of the bit packing, this is a reduction
of 0.5% of dynamic instructions in the compile benchmark, and has
dramatic improvements for the grouping symbol focused benchmarks.
All combined, this is a significant improvement on the lexer-focused
benchmarks despite the added complexity, and a significant win on our
compile time benchmarks due to both the lexer improvements and
downstream memory density improvements: 5-12% reduction in lex time,
growing larger as files get larger. About a 4.5% reduction in parse
time, and even a 1-2% reduction in total check time. =D
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Geoff Romer <gromer@google.com>
This has is a nice-to-have for me. Frequently I want to run a specific
test, and end up digging through output to be able to copy-paste the run
line. This uses TIP lines to inject the command into the file when using
AUTOUPDATE.
Note, one of the reasons I want this is because "bazel test
//toolchain/testing:file_test --test_output=all" has been regularly
exceeding bazel's output limit for me (workaround is either opening the
output file or specifying an obscure output limit flag), making it a
little harder for me to get the commands. However, frequently I'm adding
a file and want to iterate on it, so that's really the use case I have
in mind here.
The purpose of the newline is to make it clearer where a given
diagnostic begins and ends, particularly as the first message of a
diagnostic may not be the error.
This is a trivial code change, but ripples edits through test files.
I was suggesting this because `FloatingPoint` is pretty long. `int` and
`float` should be familiar abbreviations. `unsigned` should be familiar
to developers too, but `UnsignedInt` still feels usefully clearer for
the additional chars.
Specifically, after lexing a comment line, look at the next line and see
if it starts with an identical sequence of indent, comment '/'s and
character after the '/'s. If so, skip it as part of a block of comments.
This skips repeatedly diagnosing the same erroneous comment introducer
after the first one in a block, but that seems like a feature rather
than a bug.
The big motivation is to make sure the lexer is minimally impacted by
the length of comment blocks and skips them as efficiently as possible.
While they aren't exactly common, large block comments do come up and
it'd be unfortunate for those to actually slow down the toolchain.
It also happens that this is particularly easy to do because we're just
looking to see if we see the same prefix byte sequence. With SIMD we can
typically handle the most common indents with just a few instructions.
Because of the diagnostic differences, I've included a scalar fallback
that replicates the functionality but has no limit on indent size or CPU
features. I've also added testing to cover this behavior.
The only non-noise benchmark changes are as expected the comment ones,
with a nice improvement across the board:
```
BM_CommentLines/1/0/0 15.8ms ± 2% 15.6ms ± 2% -0.87% (p=0.004 n=19+19)
BM_CommentLines/4/0/0 20.2ms ± 1% 18.8ms ± 1% -6.75% (p=0.000 n=18+18)
BM_CommentLines/128/0/0 221ms ± 1% 167ms ± 1% -24.44% (p=0.000 n=20+19)
BM_CommentLines/1/30/0 16.6ms ± 3% 16.5ms ± 3% ~ (p=0.175 n=19+20)
BM_CommentLines/4/30/0 26.1ms ± 1% 24.8ms ± 2% -5.05% (p=0.000 n=18+19)
BM_CommentLines/128/30/0 233ms ± 1% 185ms ± 1% -20.38% (p=0.000 n=19+20)
BM_CommentLines/1/70/0 19.2ms ± 1% 19.0ms ± 2% -0.66% (p=0.016 n=19+20)
BM_CommentLines/4/70/0 27.9ms ± 1% 26.6ms ± 1% -4.63% (p=0.000 n=19+19)
BM_CommentLines/128/70/0 251ms ± 1% 213ms ± 1% -15.18% (p=0.000 n=20+18)
BM_CommentLines/1/0/2 15.9ms ± 1% 15.8ms ± 2% ~ (p=0.061 n=19+19)
BM_CommentLines/4/0/2 20.5ms ± 2% 19.0ms ± 2% -7.53% (p=0.000 n=20+20)
BM_CommentLines/128/0/2 213ms ± 1% 153ms ± 1% -28.18% (p=0.000 n=19+20)
BM_CommentLines/1/30/2 16.8ms ± 2% 16.7ms ± 3% ~ (p=0.134 n=20+20)
BM_CommentLines/4/30/2 26.6ms ± 1% 25.2ms ± 3% -5.50% (p=0.000 n=20+20)
BM_CommentLines/128/30/2 238ms ± 1% 187ms ± 2% -21.49% (p=0.000 n=17+19)
BM_CommentLines/1/70/2 19.3ms ± 1% 19.4ms ± 3% ~ (p=0.407 n=17+20)
BM_CommentLines/4/70/2 28.2ms ± 1% 26.9ms ± 2% -4.70% (p=0.000 n=19+19)
BM_CommentLines/128/70/2 257ms ± 2% 214ms ± 1% -16.52% (p=0.000 n=20+18)
BM_CommentLines/1/0/8 16.3ms ± 2% 16.1ms ± 2% -1.22% (p=0.001 n=20+20)
BM_CommentLines/4/0/8 22.7ms ± 2% 20.4ms ± 2% -10.20% (p=0.000 n=20+20)
BM_CommentLines/128/0/8 244ms ± 1% 153ms ± 1% -37.26% (p=0.000 n=20+18)
BM_CommentLines/1/30/8 17.3ms ± 2% 17.2ms ± 3% ~ (p=0.192 n=20+20)
BM_CommentLines/4/30/8 28.0ms ± 2% 25.6ms ± 3% -8.46% (p=0.000 n=19+18)
BM_CommentLines/128/30/8 272ms ± 1% 196ms ± 2% -27.90% (p=0.000 n=18+20)
BM_CommentLines/1/70/8 19.9ms ± 2% 19.9ms ± 2% ~ (p=0.531 n=20+19)
BM_CommentLines/4/70/8 29.3ms ± 1% 27.3ms ± 1% -6.87% (p=0.000 n=19+19)
BM_CommentLines/128/70/8 292ms ± 1% 228ms ± 1% -21.97% (p=0.000 n=20+19)
```
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>