As discussed around #3792, identify the import a diagnostic message came
from prior to the diagnostic message itself. This occurs during location
translation so that the logic can be central.
I'd considered associating the parse node with ImportRef instructions,
but I realized about halfway through that because I need to store the
ImportDirectiveId on the ImportIR for cross-package imports, it's there
for use in location translation without extra work. That saves a fair
amount of stringing it through declarations, as well as an oddity where
ImportRef instructions would have a node that didn't really represent
them.
Since the addition of TranslateArg, I don't think this type is going to
go away (cutting a TODO). Refactoring names slightly to fit the current
role, and adding const to ConvertLocation.
This doesn't matter too much as these were expanded by the LLVM iterator
facade, but eventually this should enable that facade to be a little
less fancy and should also simplify the dispatch to directly use the
three-way comparison.
One case is a bit subtle and didn't have a comment so I added one to
explain a bit what is going on there.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Move handling of mismatched brackets out of the main lexing loop into a
separate pass that is only run if there are mismatched brackets This is
done in preparation for using both lookahead and lookbehind to work out
how to match brackets, and to get this code far away from the hot lexing
loop.
Fix bracket insertion location to be immediately after the token that
we're inserting the bracket after, rather than potentially at the end of
a comment. When there are open brackets at the end of the file, say that
there are open brackets, not that there's a closing bracket without a
matching opening bracket.
We have `StringLiteral`s in multiple other `Carbon` sub-namespaces.
Rename to a more specific name to avoid collisions.
We should likely also rename `Carbon::IntId` -> `Carbon::IntValueId` and
`Carbon::RealId` -> `Carbon::RealValueId`, but this collision is
prioritized because it was blocking work on typed parse nodes which
introduces a `Carbon::Parse::StringLiteralId`.
This reflects how we're naming classes that derive from these classes,
and matches usage for each existing `Id` and `Index` type, except:
- `Parse::NodeId` previously inherited from `ComparableIndexBase`, and
is no longer comparable.
- `SemIR::MemberIndex` previously inherited from `IndexBase`, and is now
comparable.
Making `Parse::NodeId` non-comparable reflects that it's intended to be
an opaque identifier for a node and that the ordering is an
implementation detail rather than part of the intended public interface.
`PostorderIterator` and `SiblingIterator` still rely on the numerical
meaning of `NodeId`s, but that's OK since they're part of the node
implementation.
I was suggesting this because `FloatingPoint` is pretty long. `int` and
`float` should be familiar abbreviations. `unsigned` should be familiar
to developers too, but `UnsignedInt` still feels usefully clearer for
the additional chars.
Per [#toolchain
discussion](https://discord.com/channels/655572317891461132/655578254970716160/1176632520834560211)
We'd at one point been trying to put `[[nodiscard]]` everywhere, but
then we stopped because it had felt verbose without finding many issues
(plus, people plain forgot to add it). Some history in #888.
Since newer code gets added without it, we now have code like:
```
auto GetLineInfo(Line line) -> LineInfo&;
[[nodiscard]] auto GetLineInfo(Line line) const -> const LineInfo&;
auto AddLine(LineInfo info) -> Line;
auto GetTokenInfo(Token token) -> TokenInfo&;
[[nodiscard]] auto GetTokenInfo(Token token) const -> const TokenInfo&;
auto AddToken(TokenInfo info) -> Token;
[[nodiscard]] auto GetTokenPrintWidths(Token token) const -> PrintWidths;
```
Here, the lack of `[[nodiscard]]` doesn't mean anything: for example,
`GetLineInfo` should not have its result discarded if it's called. But
the mix could be confusing for readers.
As a resolution, remove the attribute. `[[nodiscard]]` should be treated
like other attributes going forward, which essentially means "avoid in
general, add a comment to explain why the attribute is needed" rather
than use-as-default.
- Treat an input of `-` as meaning stdin.
- Fix building of an llvm::MemoryBuffer from a non-regular file.
- Do not enforce filename restrictions on non-regular files.
- Do not invent an output file name based on the name of a non-regular
file.
---------
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Following up on discussion yesterday regarding this split.
Note, I'm expecting #3341 to do IdentifierId -> NameId in SemIR. It
might be worth adding NameId creation directly to StringStore if you're
content with this setup though.
This detects ordering issues with the `package` and `import` statements.
`library` is changed from package-specific to instead be generic between
the two, since structurally it's non-specific.
The next step would be to start exposing the results for the driver to
make ordering decisions for checking. That'll involve further
modifications to this code, but this felt like a reasonable change point
because it's the extent of the parser enforcement, and still causes
significant refactoring.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Building on #3311, which started moving the result string into a
`unique_ptr`, instead have `StringLiteral` use a `BumpPtrAllocator` to
manage memory. But also, detect when a string is really trivial during
`Lex` and, if so, return `contents_` directly.
Building on #3311, change SemIR to use the SharedValueStore. Since this
removes hermeticity, raw output no longer prints ints, reals, and
strings. TokenizedBuffer accessors are modified to return IDs because
values are often passed through in semantics without needing to read
them.
I would've put SharedValueStores on Context, except for the
GetArrayBoundValue convenience method. I felt awkward removing that, so
it's on File, at least for now. That's then used by the formatter and
Lower too. The flipside of this is that TokenizedBuffer has a
SharedValueStores only for printing, so maybe that's similar enough to
what File is doing.
This doesn't start shifting other SemIR members to ValueStore, but that
seems like a next step.
This updates lexing to use the data. I'll do checking separately, just
to split changes.
Note the ValueStore structure is also set up such that SemIR::File can
use it for other fields.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
## Summary ##
Restructures the lexer to first scan the entire source text for newlines
and create all the line structures needed. Doing this up-front makes it
easy to produce an optimized version with minimal complexity. Currently,
it leverages the system `memchr`, but even when expanded to handle more
complex cases like CR+LF line endings, being isolated in this way will
result in a significantly simpler implementation. This change improves
the lexing of comment lines significantly by skipping their contents
immediately. The overhead of the pre-scan is unmeasurable in all
realistic benchmarks, and 10-30% in benchmarks consisting almost
entirely of blank lines or comments. The improvement of comment lexing
with average length comment lines mixed with code starts at 20% and goes
up. Regressing blank line handling for non-empty comments seems like the
right tradeoff (by far).
## Background and details ##
One weak point in the lexer implementation were large runs of comments.
While those aren't terribly common, they shouldn't present a hazard to
the lexer performance.
A bit more common is a pattern of comments like the following:
```carbon
// Some method comment here.
fn SomeMethodName(...) -> ...;
// Some other method comment here.
fn SomeOtherMethodName(...) -> ...;
```
Here, the lexer spends an inordinate amount of time getting from the
`\n` after the first semicolon to the `fn` token. It has to skip a blank
line, scan a line, find the `//` comment start, then scan to find the
next `\n`, and then scan horizontal whitespace, etc.
It is tempting to build a scanner *exactly* for this. In fact, I built
one, and I can publish it in a PR if folks are interested in what it
looks like. For x86-64, the PSHUFB trick used for scanning identifiers
technically works. But it is *complicated*. Amazingly so. 150 lines of
very subtle code with subtle performance pitfalls at ever turn. I felt
very uncomfortable submitting it, but we can always go back to it.
Nothing I've come up with quite matches it for sheer speed.
However, most of the complexity and time is spent walking from a `//` to
the end of the line. And *that* is something we can do very simply. In
fact, there is a tuned function for that in libc: `memchr`. Using this
we can build a very fast and much simpler scanner to split lines
up-front. This PR uses that and a carefully crafted fast loop to first
build up all the line info we need. Getting this to be as fast as
possible required some other subtle changes, for example always creating
a line structure that goes from the last `\n` and the end of the file.
We then back up the EOF token to avoid surfacing this to users. The nice
thing is the EOF token isn't part of any hot loop, and so this removes
branches everywhere else at modest complexity.
Once we have that, the rest of the lexer just needs to keep track of its
current line in order to record column offsets. I've taken some care to
try and optimize the lexer's usage of the line structures but there are
more opportunities here I suspect.
Combined, this gets much but not all of the performance of a huge SIMD
scanner for newline-through-to-next-token. For extreme cases (100s of
blank lines or empty comment lines between tokens) the holistic scanner
is of course still much faster, but those don't seem nearly worth the
cost.
I was initially worried about the overhead of taking two passes over the
source text, but in practice I've not been able to measure any
appreciable cost to this with realistic source files. In some cases
benchmarks with no newlines get *faster* because we use a much more
efficient approach to fetch the source text into cache as
a happenstance. And that in turn makes the byte-wise dispatched loop run
faster as it stalls less.
I'm particularly happy with this approach because it seems very clear
how to extend this to support CR+LF, bare CR, and even complex mixtures
without any significant speed cost. That wasn't at all true for the
other approaches explored.
I may try some further PRs to smooth out the last bits of slowness here,
but already this is working excellent for me in practice. My 10mloc test
case is down to 2.3s to lex.
## Raw benchmark data
Using a tool that runs benchmarks before and after and analyzes the
results, the following summarizes the CPU-time impact, each of these for
lexing 100k tokens:
```
BM_ValidKeywords 2.57ms ± 1% 2.58ms ± 0% ~ (p=0.190 n=5+4)
BM_ValidIdentifiers<1, 64, false> 9.24ms ± 4% 9.31ms ± 4% ~ (p=0.421 n=5+5)
BM_ValidIdentifiers<1, 1, true> 3.05ms ± 4% 3.11ms ± 4% ~ (p=0.222 n=5+5)
BM_ValidIdentifiers<3, 5, true> 10.9ms ± 0% 11.1ms ± 1% +1.76% (p=0.016 n=4+5)
BM_ValidIdentifiers<3, 16, true> 11.1ms ± 7% 11.0ms ± 1% ~ (p=0.310 n=5+5)
BM_ValidIdentifiers<12, 64, true> 12.2ms ± 1% 12.3ms ± 2% ~ (p=0.111 n=4+5)
BM_HorizontalWhitespace/1 11.2ms ± 6% 11.1ms ± 2% ~ (p=0.841 n=5+5)
BM_HorizontalWhitespace/4 12.0ms ± 3% 12.0ms ± 2% ~ (p=0.548 n=5+5)
BM_HorizontalWhitespace/16 16.2ms ± 6% 15.9ms ± 8% ~ (p=0.690 n=5+5)
BM_HorizontalWhitespace/64 27.7ms ± 3% 28.4ms ± 3% ~ (p=0.151 n=5+5)
BM_HorizontalWhitespace/128 44.3ms ± 1% 45.6ms ± 6% +3.15% (p=0.032 n=5+5)
BM_RandomSource 7.75ms ± 2% 7.72ms ± 1% ~ (p=1.000 n=5+5)
BM_BlankLines/1 11.7ms ± 1% 12.1ms ± 1% +3.46% (p=0.008 n=5+5)
BM_BlankLines/4 14.0ms ± 2% 15.2ms ± 3% +8.12% (p=0.008 n=5+5)
BM_BlankLines/16 23.5ms ± 2% 31.1ms ± 4% +32.26% (p=0.008 n=5+5)
BM_BlankLines/64 75.3ms ± 1% 81.2ms ± 3% +7.83% (p=0.008 n=5+5)
BM_BlankLines/128 133ms ± 3% 150ms ± 2% +12.74% (p=0.008 n=5+5)
BM_CommentLines/1/0/0 13.1ms ± 0% 13.7ms ± 1% +5.11% (p=0.008 n=5+5)
BM_CommentLines/4/0/0 16.6ms ± 1% 18.2ms ± 4% +9.56% (p=0.008 n=5+5)
BM_CommentLines/128/0/0 169ms ± 4% 182ms ± 1% +7.24% (p=0.008 n=5+5)
BM_CommentLines/1/30/0 18.7ms ± 5% 14.1ms ± 0% -24.84% (p=0.008 n=5+5)
BM_CommentLines/4/30/0 36.5ms ± 6% 20.6ms ± 3% -43.59% (p=0.008 n=5+5)
BM_CommentLines/128/30/0 525ms ± 4% 198ms ± 1% -62.38% (p=0.008 n=5+5)
BM_CommentLines/1/70/0 23.4ms ± 6% 14.7ms ± 2% -37.15% (p=0.008 n=5+5)
BM_CommentLines/4/70/0 53.3ms ± 7% 22.4ms ± 4% -57.99% (p=0.008 n=5+5)
BM_CommentLines/128/70/0 1.05s ± 4% 0.21s ± 2% -80.31% (p=0.008 n=5+5)
BM_CommentLines/1/0/2 14.1ms ± 6% 14.3ms ± 1% ~ (p=0.151 n=5+5)
BM_CommentLines/4/0/2 19.4ms ± 5% 20.1ms ± 1% ~ (p=0.151 n=5+5)
BM_CommentLines/128/0/2 238ms ± 8% 229ms ± 0% ~ (p=0.151 n=5+5)
BM_CommentLines/1/30/2 19.2ms ± 7% 14.6ms ± 1% -23.87% (p=0.008 n=5+5)
BM_CommentLines/4/30/2 40.3ms ±13% 22.3ms ± 4% -44.63% (p=0.008 n=5+5)
BM_CommentLines/128/30/2 568ms ± 7% 254ms ± 3% -55.28% (p=0.008 n=5+5)
BM_CommentLines/1/70/2 23.3ms ± 1% 15.0ms ± 3% -35.61% (p=0.016 n=4+5)
BM_CommentLines/4/70/2 57.2ms ± 9% 24.1ms ± 2% -57.81% (p=0.008 n=5+5)
BM_CommentLines/128/70/2 1.07s ± 0% 0.26s ± 2% -75.51% (p=0.016 n=4+5)
BM_CommentLines/1/0/8 15.9ms ± 7% 16.0ms ± 1% ~ (p=0.151 n=5+5)
BM_CommentLines/4/0/8 24.2ms ± 6% 27.9ms ± 2% +15.36% (p=0.008 n=5+5)
BM_CommentLines/128/0/8 386ms ± 4% 445ms ± 1% +15.28% (p=0.008 n=5+5)
BM_CommentLines/1/30/8 20.6ms ± 5% 16.3ms ± 1% -20.95% (p=0.008 n=5+5)
BM_CommentLines/4/30/8 45.3ms ± 6% 30.3ms ± 3% -32.98% (p=0.008 n=5+5)
BM_CommentLines/128/30/8 699ms ± 3% 477ms ± 3% -31.83% (p=0.008 n=5+5)
BM_CommentLines/1/70/8 25.7ms ± 5% 16.8ms ± 2% -34.67% (p=0.008 n=5+5)
BM_CommentLines/4/70/8 62.0ms ± 4% 31.6ms ± 2% -49.10% (p=0.008 n=5+5)
BM_CommentLines/128/70/8 1.20s ± 2% 0.48s ± 4% -59.60% (p=0.008 n=5+5)
```
The horizontal whitespace benchmark (and all of the non-line-oriented
ones) are noisier than they appear here but do show some improvements
(surprisingly). My guess is that it has a lot to do with system load, as
the advantage is that we're using a vectorized loop to scan the text
first and then doing the byte-dispatched loop. So when the cache is
a bit slower to populate, the vectorized version starts to be faster.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Continuing with #3070. Just a dir and file rename (only prefix change is
lexer_file_test). Everything in the lex dir should be marked as a move.
Note, I think this closes#3070. There may still be further cleanup
later, but the organizational changes suggested there are being
completed.
---------
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>