Commit Graph
17 Commits
Author SHA1 Message Date
Jon Ross-Perkins b8ceb8dd8b Print a blank line after a diagnostic. (#3806)
The purpose of the newline is to make it clearer where a given
diagnostic begins and ends, particularly as the first message of a
diagnostic may not be the error.

This is a trivial code change, but ripples edits through test files.
2024-03-22 18:10:49 +00:00
Jon Ross-Perkins 551a6d385e Augment the file_test framework to allow per-file fail checks. (#3747)
This handles toolchain failures per-file. The intent is to allow placing
both "success" and "fail" tests in the same file, using splits. However,
this PR only adds support and updates existing tests to continue
passing.
2024-03-08 20:28:25 +00:00
Richard Smith 0a06fceb5f Improve diagnosis of mismatched brackets. (#3282)
Move handling of mismatched brackets out of the main lexing loop into a
separate pass that is only run if there are mismatched brackets This is
done in preparation for using both lookahead and lookbehind to work out
how to match brackets, and to get this code far away from the hot lexing
loop.

Fix bracket insertion location to be immediately after the token that
we're inserting the bracket after, rather than potentially at the end of
a comment. When there are open brackets at the end of the file, say that
there are open brackets, not that there's a closing bracket without a
matching opening bracket.
2023-12-21 08:49:37 +00:00
josh11bandJon Ross-Perkins fada410559 Support declaration modifier keywords (#3412)
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-12-05 22:45:57 +00:00
Jon Ross-Perkins 0db63ff17a Abbreviate Integer and FloatingPoint (#3435)
I was suggesting this because `FloatingPoint` is pretty long. `int` and
`float` should be familiar abbreviations. `unsigned` should be familiar
to developers too, but `UnsignedInt` still feels usefully clearer for
the additional chars.
2023-11-29 23:29:48 +00:00
Jon Ross-Perkins 3f208e27f9 Align on FileStart/FileEnd for naming. (#3428)
The lexer has been using EndOfFile form (stemming from EOF), parser went
to FileEnd form. This consolidates on FileEnd form.
2023-11-29 16:36:57 +00:00
Jacob Schneider 482d233def fix crash caused by unicode chars (#3387)
Prevent a crash during lexing for unicode chars.
2023-11-11 06:27:15 +00:00
Jon Ross-Perkins d096655cc6 Split out the SharedValueStores to be per-compilation unit. (#3353)
Advantages:

- Allows lexing/parsing in parallel, since they are modifying fully
separate ValueStores.
- Allows SemIR to reliably be stored hermetically.

Disadvantages:

- Creates overlapping storage of duplicate strings when multiple files
are compiled together.
- Prevents Ids from being uniquely compared cross-file.

Per discussion, the decision is that the advantages are more important.

The looser ownership remains because both SemIR checking and
metaprogramming may still generate things we would want to deduplicate.
It would be somewhat odd if TokenizedBuffer owned something that
checking modified.
2023-11-01 19:44:17 +00:00
Jon Ross-Perkins 6742d0d048 Add partial raw identifier support. (#3344)
I'm looking at this due to the conversation on #3341. Although
diagnostics aren't where they should be, I thought it may help to start
adding raw identifier support (which may also help show how I was
thinking about this).

Note regarding the TODO on how to form the token, `GetTokenText` returns
the `string_id`'s reference value for an `Identifier`. So to make
`GetTokenText` work in a way that returns `r#foo` for a raw identifier,
I think there are a few options:

1. Add additional data indicating the end of the identifier.
2. Add `RawIdentifier` as a token kind to indicate that it's raw and
should be prefixed with `r#` (but also giving later stages one more
token kind to handle)
3. Make the `string_id` correspond to `r#foo`, and have later stages add
`foo` to the strings table whenever `r#foo` is encountered (with map
lookups leading to deduplication).
4. Add `StringId::RawKeyword` special values for each keyword.
- This would mean `self` prints as `self`, `r#self` prints as `r#self`,
but `r#foo` is not a keyword so prints as `foo`.
- This means keywords would need to be listed in a place `StringId` can
depend on them, one way or the other (e.g., a `keywords.def` file in
`base/` should work).
5. Say that it _is_ an `Identifier`, and if it's a keyword spelling, it
must have been a raw identifier.
- Same limitation as above: This would mean `self` prints as `self`,
`r#self` prints as `r#self`, but `r#foo` is not a keyword so prints as
`foo`.

I'm hoping to resolve this issue separately though. :)
2023-10-30 18:11:55 +00:00
Jon Ross-PerkinsandRichard Smith d13f76e001 Add value store to be shared across compile stages. (#3311)
This updates lexing to use the data. I'll do checking separately, just
to split changes.

Note the ValueStore structure is also set up such that SemIR::File can
use it for other fields.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-10-20 15:00:53 +00:00
Chandler CarruthandRichard Smith 3015135a52 Skip blocks of comments with identical prefixes. (#3299)
Specifically, after lexing a comment line, look at the next line and see
if it starts with an identical sequence of indent, comment '/'s and
character after the '/'s. If so, skip it as part of a block of comments.
This skips repeatedly diagnosing the same erroneous comment introducer
after the first one in a block, but that seems like a feature rather
than a bug.

The big motivation is to make sure the lexer is minimally impacted by
the length of comment blocks and skips them as efficiently as possible.
While they aren't exactly common, large block comments do come up and
it'd be unfortunate for those to actually slow down the toolchain.

It also happens that this is particularly easy to do because we're just
looking to see if we see the same prefix byte sequence. With SIMD we can
typically handle the most common indents with just a few instructions.

Because of the diagnostic differences, I've included a scalar fallback
that replicates the functionality but has no limit on indent size or CPU
features. I've also added testing to cover this behavior.

The only non-noise benchmark changes are as expected the comment ones,
with a nice improvement across the board:

```
BM_CommentLines/1/0/0                          15.8ms ± 2%  15.6ms ± 2%   -0.87%  (p=0.004 n=19+19)
BM_CommentLines/4/0/0                          20.2ms ± 1%  18.8ms ± 1%   -6.75%  (p=0.000 n=18+18)
BM_CommentLines/128/0/0                         221ms ± 1%   167ms ± 1%  -24.44%  (p=0.000 n=20+19)
BM_CommentLines/1/30/0                         16.6ms ± 3%  16.5ms ± 3%     ~     (p=0.175 n=19+20)
BM_CommentLines/4/30/0                         26.1ms ± 1%  24.8ms ± 2%   -5.05%  (p=0.000 n=18+19)
BM_CommentLines/128/30/0                        233ms ± 1%   185ms ± 1%  -20.38%  (p=0.000 n=19+20)
BM_CommentLines/1/70/0                         19.2ms ± 1%  19.0ms ± 2%   -0.66%  (p=0.016 n=19+20)
BM_CommentLines/4/70/0                         27.9ms ± 1%  26.6ms ± 1%   -4.63%  (p=0.000 n=19+19)
BM_CommentLines/128/70/0                        251ms ± 1%   213ms ± 1%  -15.18%  (p=0.000 n=20+18)
BM_CommentLines/1/0/2                          15.9ms ± 1%  15.8ms ± 2%     ~     (p=0.061 n=19+19)
BM_CommentLines/4/0/2                          20.5ms ± 2%  19.0ms ± 2%   -7.53%  (p=0.000 n=20+20)
BM_CommentLines/128/0/2                         213ms ± 1%   153ms ± 1%  -28.18%  (p=0.000 n=19+20)
BM_CommentLines/1/30/2                         16.8ms ± 2%  16.7ms ± 3%     ~     (p=0.134 n=20+20)
BM_CommentLines/4/30/2                         26.6ms ± 1%  25.2ms ± 3%   -5.50%  (p=0.000 n=20+20)
BM_CommentLines/128/30/2                        238ms ± 1%   187ms ± 2%  -21.49%  (p=0.000 n=17+19)
BM_CommentLines/1/70/2                         19.3ms ± 1%  19.4ms ± 3%     ~     (p=0.407 n=17+20)
BM_CommentLines/4/70/2                         28.2ms ± 1%  26.9ms ± 2%   -4.70%  (p=0.000 n=19+19)
BM_CommentLines/128/70/2                        257ms ± 2%   214ms ± 1%  -16.52%  (p=0.000 n=20+18)
BM_CommentLines/1/0/8                          16.3ms ± 2%  16.1ms ± 2%   -1.22%  (p=0.001 n=20+20)
BM_CommentLines/4/0/8                          22.7ms ± 2%  20.4ms ± 2%  -10.20%  (p=0.000 n=20+20)
BM_CommentLines/128/0/8                         244ms ± 1%   153ms ± 1%  -37.26%  (p=0.000 n=20+18)
BM_CommentLines/1/30/8                         17.3ms ± 2%  17.2ms ± 3%     ~     (p=0.192 n=20+20)
BM_CommentLines/4/30/8                         28.0ms ± 2%  25.6ms ± 3%   -8.46%  (p=0.000 n=19+18)
BM_CommentLines/128/30/8                        272ms ± 1%   196ms ± 2%  -27.90%  (p=0.000 n=18+20)
BM_CommentLines/1/70/8                         19.9ms ± 2%  19.9ms ± 2%     ~     (p=0.531 n=20+19)
BM_CommentLines/4/70/8                         29.3ms ± 1%  27.3ms ± 1%   -6.87%  (p=0.000 n=19+19)
BM_CommentLines/128/70/8                        292ms ± 1%   228ms ± 1%  -21.97%  (p=0.000 n=20+19)
```

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-10-17 16:07:03 +00:00
Chandler Carruth a46ca6bf7a Add a start-of-file token and parse node. (#3263)
This removes a (very) hot branch in the lexer where we need to special
case when a token is the first token and can't look at its previous
token. It also seems like a generally nice change to the structure of
both the token buffer and parse tree as there are now bracketing
elements for both ends and we should be able to avoid similar branching
in the future.

Mostly mechanical updates to the lexer and parser code to handle this,
but also needed to special case the location information in the
autoupdate code. And then the usual large body of auto-updated tests.

No benchmark data for this change alone as in isolation and in the
current lexer structure it doesn't make a big difference. But this
branch was particularly difficult to handle when trying to update the
whitespace skipping code to be faster, and so I think it is worth
systematically avoiding the special case here.
2023-10-04 23:36:35 +00:00
Geoff Romer 7899154a21 Add "ERROR" to all error diagnostics (#3251)
This makes the difference between errors and lower-level diagnostics
visible to users, and aligns the toolchain's behavior with the
expectations in `driver_fuzzer.cpp`.
2023-09-26 16:52:49 +00:00
Jon Ross-Perkins 75282462d4 When starting a split file, try adding stdout lines. (#3233)
Addresses
https://github.com/carbon-language/carbon-lang/pull/3217#discussion_r1323581989
2023-09-15 21:33:48 +00:00
Jon Ross-Perkins 2ecab78297 Support multi-file lex printing and testing. (#3214)
Lex now prints its yaml as:
```
- filename: name
  tokens: [ ... ]
```

New support in file_test allows the `filename` marker at the top to
define the default file number for later lines, meaning multi-file
output from lexing is now associated with the appropriate file. Similar
support will probably also apply to lowering, semir, and other places
that print a filename once for the full dump.

This hammers a bit at how line number replacements work in file_test,
allowing stacking them so that lex errors and stdout can both be
line-associated properly. I've tried to make the autoupdate more
frequently work in one pass, now also taking into account the file index
when doing line replacements.

There are still some issues with EndOfFile that it may be good to
discuss: because CHECK lines are appended to the end of the file now,
and the EndOfFile token points at the last line including comments, new
lex tests now take two runs to autoupdate (because without CHECK lines,
the EndOfFile points at a content line, which content is then inserted
after). Note that removing CHECK lines from the test is not a solution:
autoupdate also started inserting blank lines, which breaks this for a
similar reason. One solution here might be to not have EndOfFile
associate with a line or column, which has been a bit of an issue
regardless.

Also fixes a small issue with toolchain's autoupdate script.
2023-09-13 16:14:09 +00:00
Jon Ross-Perkins 87d4c2dfc6 Improve testing and handling of unsigned APInt values (#3202) 2023-09-07 18:53:38 +00:00
Jon Ross-PerkinsandChandler Carruth ec182fb00d Rename lexer dir to lex (#3179)
Continuing with #3070. Just a dir and file rename (only prefix change is
lexer_file_test). Everything in the lex dir should be marked as a move.

Note, I think this closes #3070. There may still be further cleanup
later, but the organizational changes suggested there are being
completed.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-09-01 02:39:04 +00:00