Commit Graph
679 Commits
Author SHA1 Message Date
Richard Smith 0fb2924c99 Extend description of parser states to show which tokens they consume. (#3378)
Also some minor improvements and typo fixes to the parser code for
issues found while writing these descriptions.
2023-11-10 22:08:58 +00:00
josh11b b953fbf314 Subexpr -> SubExpr (#3383) 2023-11-10 20:48:46 +00:00
Richard Smith 4aa6a6894d Add missing periods to diagnostics. (#3380) 2023-11-10 20:29:22 +00:00
Richard Smith afd6d85610 Support for returned var and return var. (#3374)
Implement toolchain support for `returned var` and `return var`.

- Modeled `returned` in the parse tree as a `ReturnedSpecifier`
appearing after the `VariableIntroducer`.
- Modeled `return var` in the parse tree as a `ReturnVarSpecifier`
appearing after the `ReturnStatementStart`.
- Factored out the implementation of `return` statement and `returned
var` handling in check into a new `return.{h,cpp}`. The parse nodes
themselves are still handled in `handle_*.cpp`. This allows easy code
reuse between `return` and `returned var`.
2023-11-10 19:52:30 +00:00
josh11b 5020fdb3be Use abbreviation "decl" instead of "declaration" (#3382)
Part of switching to the [abbreviations we've decided to
use](https://docs.google.com/document/d/1RRYMm42osyqhI2LyjrjockYCutQ5dOf8Abu50kTrkX0/edit?resourcekey=0-kHyqOESbOHmzZphUbtLrTw#heading=h.pph7i5m5un7q).

I will rename files in a follow-up PR.
2023-11-10 10:43:25 +00:00
josh11b 11ca083855 Use abbreviation "expr" instead of "expression" (#3375)
Part of switching to the [abbreviations we've decided to
use](https://docs.google.com/document/d/1RRYMm42osyqhI2LyjrjockYCutQ5dOf8Abu50kTrkX0/edit?resourcekey=0-kHyqOESbOHmzZphUbtLrTw#heading=h.pph7i5m5un7q).

I will rename files in a follow-up PR.
2023-11-10 01:32:32 +00:00
Richard Smith 6d5e62974c Add SemIR instruction to track that a conversion was performed. (#3363)
Instead of ad-hoc conversion tracking on some kinds of nodes that
conversion creates, consolidate tracking into a single node kind. This
frees up an operand on `Init` instructions that can be used to store the
destination.
2023-11-09 23:58:54 +00:00
josh11b 681fbf9da2 Add missing #include of base/value_store.h (#3377) 2023-11-09 20:58:18 +00:00
Richard Smith 71aa4a45be Distinguish between name IDs and string IDs in the type system. (#3341)
Add a `NameId` that is effectively just a wrapper around a `StringId`,
with
some additional predefined values for names that don't correspond to
strings, such as the name of `self` or the function's return slot.
2023-11-09 16:51:36 +00:00
Jon Ross-Perkins 84bc8cc4bf Fix crash when if expressions aren't in a function. (#3373)
Another issue found while trying to make `package` work, lurking in
fuzzer inputs. This leaves TODOs because we probably do want to support
this, it's just non-trivial to fix.
2023-11-07 22:14:15 +00:00
Jon Ross-Perkins 50a614aaf3 Fix crash when an expression cannot convert to a type. (#3372)
This turns out to somewhat block `package` support because there's a
fuzzer test-case that does similar. The parse is valid so we should
probably handle it reasonably.

I suspect the `let` test case might work with a little effort, given it
shouldn't really require much evaluation. On the other hand, `var`
definitely shouldn't, and `fn` will probably require something like
`constexpr` plus more substantial compile-time evaluation support.
2023-11-07 21:36:35 +00:00
josh11b f9fb27bdfc Split handlers for aggregates out of lower/handle.cpp into its own file (#3368) 2023-11-06 22:20:16 +00:00
Richard Smith 3bee8932a9 Rework name lookup to handle non-lexical scoping. (#3354)
When declaring a name such as `fn Ns.Class.F() { ... }`, enter the
scopes of `Ns` and `Ns.Class` as we form the name, and remain in those
non-lexical scopes until the end of the declaration.

When performing an unqualified lookup, look in any enclosing non-lexical
scopes in addition to looking into the lexical name table.

We now track a scope index with each lookup result in the lexical name
lookup table. This is used to determine whether a lexical or non-lexcial
result is the innermost result and whether a declared name is in the
same scope as some previous introduction of that name or in a nested
scope. For now, this could just be the index into the scope_stack, but
the intent is to also use this to detect names being declared after they
are first looked up, which requires the indexes to outlive their scopes,
so we use a persistent numbering of all scopes instead. The persistent
numbering also permits more invariant checking.
2023-11-06 19:48:41 +00:00
Richard Smith 184eafd521 Add a separate store for computed constant values. (#3362)
This moves the instructions generated for type values out of the block
in which they happen to first be referenced, and into shared storage.
2023-11-04 23:48:13 +00:00
josh11b 0318631d1a Clarify some comments (#3360) 2023-11-03 00:25:05 +00:00
Richard SmithandJon Ross-Perkins 545a5b3679 Support initializing a class from a struct. (#3358)
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-11-02 20:27:01 +00:00
josh11b 737162cc8f Rename sem_ir files node->inst, follow up to #3355 (#3361) 2023-11-02 19:40:41 +00:00
Jon Ross-Perkins 3401eed8d8 Split IdentifierId and StringLiteralId from StringId (#3352)
Following up on discussion yesterday regarding this split.

Note, I'm expecting #3341 to do IdentifierId -> NameId in SemIR. It
might be worth adding NameId creation directly to StringStore if you're
content with this setup though.
2023-11-02 18:44:32 +00:00
josh11bandChandler Carruth 7edfd8e02a Rename SemIR::Node to SemIR::Inst (#3355)
And generally replace "node" by "inst" in the code and "instruction" in
comments.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-11-02 17:58:30 +00:00
Richard Smith 35c2142392 Treat field access into a class value expression as a value expression. (#3357)
Per the design, field access into a class value expression is a value
expression, even though we could produce an ephemeral reference
expression instead and avoid performing a value binding. This slightly
pessimizes class member access in some cases, but we should be able to
restore the old generated code by deferring actually performing the
value binding until a value expression is needed.
2023-11-02 16:49:23 +00:00
Richard Smith 544f802746 Fix member accesses that name a nested class. (#3356)
The value in the name lookup table is the class declaration. Map it to
the class type when it's found by lookup in an expression. This differs
from the behavior in a declaration name, where we want to find the class
declaration itself.
2023-11-02 15:57:52 +00:00
Jon Ross-Perkins d096655cc6 Split out the SharedValueStores to be per-compilation unit. (#3353)
Advantages:

- Allows lexing/parsing in parallel, since they are modifying fully
separate ValueStores.
- Allows SemIR to reliably be stored hermetically.

Disadvantages:

- Creates overlapping storage of duplicate strings when multiple files
are compiled together.
- Prevents Ids from being uniquely compared cross-file.

Per discussion, the decision is that the advantages are more important.

The looser ownership remains because both SemIR checking and
metaprogramming may still generate things we would want to deduplicate.
It would be somewhat odd if TokenizedBuffer owned something that
checking modified.
2023-11-01 19:44:17 +00:00
Jon Ross-PerkinsandRichard Smith c9458fe30a Add parse support for 'import', brush up 'package' a little. (#3347)
This detects ordering issues with the `package` and `import` statements.
`library` is changed from package-specific to instead be generic between
the two, since structurally it's non-specific.

The next step would be to start exposing the results for the driver to
make ordering decisions for checking. That'll involve further
modifications to this code, but this felt like a reasonable change point
because it's the extent of the parser enforcement, and still causes
significant refactoring.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-10-31 21:26:49 +00:00
Richard Smith 428d11323a Directly convert from an initializer to a value where possible. (#3306)
If the initializing representation is the same as the value
representation, don't materialize a temporary and perform a value
binding. Instead, directly extract the value, using a new
`value_of_initializer` node.

This removes a lot of redundant `alloca`s from our generated LLVM IR.
2023-10-31 20:59:28 +00:00
Richard Smith bc8db9406d Support for Self expressions. (#3351) 2023-10-31 19:59:50 +00:00
Richard SmithandJon Ross-Perkins a2e7d5e008 Take the address of the object when calling an addr self method. (#3349)
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-10-31 19:29:44 +00:00
Richard Smith 3ea3cc38b3 Support for as expressions. (#3348) 2023-10-31 00:24:11 +00:00
Jon Ross-Perkins 6742d0d048 Add partial raw identifier support. (#3344)
I'm looking at this due to the conversation on #3341. Although
diagnostics aren't where they should be, I thought it may help to start
adding raw identifier support (which may also help show how I was
thinking about this).

Note regarding the TODO on how to form the token, `GetTokenText` returns
the `string_id`'s reference value for an `Identifier`. So to make
`GetTokenText` work in a way that returns `r#foo` for a raw identifier,
I think there are a few options:

1. Add additional data indicating the end of the identifier.
2. Add `RawIdentifier` as a token kind to indicate that it's raw and
should be prefixed with `r#` (but also giving later stages one more
token kind to handle)
3. Make the `string_id` correspond to `r#foo`, and have later stages add
`foo` to the strings table whenever `r#foo` is encountered (with map
lookups leading to deduplication).
4. Add `StringId::RawKeyword` special values for each keyword.
- This would mean `self` prints as `self`, `r#self` prints as `r#self`,
but `r#foo` is not a keyword so prints as `foo`.
- This means keywords would need to be listed in a place `StringId` can
depend on them, one way or the other (e.g., a `keywords.def` file in
`base/` should work).
5. Say that it _is_ an `Identifier`, and if it's a keyword spelling, it
must have been a raw identifier.
- Same limitation as above: This would mean `self` prints as `self`,
`r#self` prints as `r#self`, but `r#foo` is not a keyword so prints as
`foo`.

I'm hoping to resolve this issue separately though. :)
2023-10-30 18:11:55 +00:00
Richard SmithandJon Ross-Perkins d42c1a3a7c Diagnose attempts to copy a non-copyable type. (#3345)
For now, we treat class types and `String` as non-copyable, because we
don't know how to emit SemIR to copy them yet. This will change as we
add support for copying those types when appropriate.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-10-28 01:00:59 +00:00
Richard Smith ae22338468 Allow fields to be reordered in struct initialization. (#3346) 2023-10-28 00:55:06 +00:00
Richard SmithandJon Ross-Perkins 57f3c553b8 Support for type-checking and lowering method calls. (#3343)
Adds a `BoundMethod` SemIR node to represent an `x.F` bound method, with
a new builtin type `BoundMethodType`. Reorganized conversion of call
expression arguments to also check and convert a `self` parameter in the
implicit parameters list.

In passing, improved diagnostics and error recovery for bad call
expressions. We now build a `call` node with the appropriate type and
value category, but with invalid arguments, if the argument conversion
failed, and diagnose calls to non-callable expressions.

`addr self` methods don't work properly yet; the `addr` is ignored for
now.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-10-27 20:41:13 +00:00
Richard Smith 04ae5a0531 Support for functions with a self parameter. (#3338)
So far, such functions can only be defined; calls are not supported yet.
2023-10-26 21:01:46 +00:00
josh11b e7b3d395b3 Fix typo (#3340) 2023-10-26 20:59:28 +00:00
Jon Ross-Perkins 3af7eb2672 Refactor YAML handling to use the llvm::yaml API. (#3337)
Provides an adapter for the llvm::yaml API because it otherwise needs a
bunch of const/non-const definitions, and the traits are difficult to
diagnose issues with. The current approach is pretty simple to use, even
if it's not super efficient (which, yaml output is more of a debugging
thing so I'm not really expecting it to be an issue).

Changes the format of yaml output to provide more index information,
just as reminders when seeing something like `node+0`. Note this would
create more churn in deltas if we were reliant on the output yaml in
tests, but we aren't so it should be okay.
2023-10-26 18:50:30 +00:00
Richard Smith 387a1711af Rename Deduced parameters to Implicit parameters. (#3336)
In preparation for supporting `self`, which is implicit but not deduced.
2023-10-25 21:46:12 +00:00
Richard SmithandJon Ross-Perkins ab575cf32a Support for member access into classes. (#3335)
Add support for member access into classes, for both non-instance
members and for fields.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-10-25 21:06:45 +00:00
Richard Smith 620408b999 Basic lowering support for classes. (#3334)
Track the fields in a class, and generate a corresponding struct type as
the object representation for the class. For now, we always use a
pointer as the value representation for a class.
2023-10-25 20:44:24 +00:00
Richard Smith 7f2f4bec4f Add a Field node for fields in a class. (#3332)
This replaces the use of `VarStorage` in this case.

Add an `UnboundFieldType` type as the type of a field, in cases where
it's referenced without an accompanying object.

Add a `BindName` node to describe the name binding performed for both
variables and fields so that we can handle them more uniformly.
2023-10-25 18:49:36 +00:00
Jon Ross-Perkins 74c3c665fa Refactor SemIR YAML printing to use dashed lists. (#3330)
This standardizes on having ValueStore and related structures provide
printing, removing the handlers in file.cpp.

The `[]` is provided for empty sequences versus if there was simply
nothing, in which case it would be a sequence when non-empty, and a null
value when empty. Consistently (and explicitly) providing sequences
feels easier to understand.

The changes to the output yaml are overall more terse. My hope is that
this is an improvement for most readers.

Also fixes printing of APInt, defaulting to unsigned for consistency
with Carbon's use.
2023-10-25 16:32:24 +00:00
Richard Smith 35721dc3d0 sem-ir: Write references to the file block as file. not package. (#3333)
This matches the renaming of the `package { ... }` block to `file { ...
}`.
2023-10-25 00:35:16 +00:00
Jon Ross-Perkins e6634d240f Make SemIR::File access more terse. (#3331)
1. In general, `semantics_ir` -> `sem_ir`, to match the directory name.
2. For the list of `ValueStore`-related accessors on `SemIR::File`, add
them to `check`'s `Context` object, shortening access.
2023-10-24 21:05:46 +00:00
Jon Ross-Perkins b2cfd5a8a8 Change StringLiteral to less frequently allocate a new string. (#3314)
Building on #3311, which started moving the result string into a
`unique_ptr`, instead have `StringLiteral` use a `BumpPtrAllocator` to
manage memory. But also, detect when a string is really trivial during
`Lex` and, if so, return `contents_` directly.
2023-10-24 19:15:49 +00:00
Jon Ross-Perkins 1d6298290f Add more value store types to File. (#3317)
Finishing what #3316 started, add more bespoke ValueStore-like
structures to File. With this, the things which previously had somewhat
boilerplate Add/Get functions are now all on side classes, giving a
uniform style of API for calling.

Note, I was on the fence about making things public on ValueStore. If
it's preferred that I make some things there protected I certainly can,
there's just a trade-off that may mean more distinct child/wrapper
types.
2023-10-24 18:23:40 +00:00
Chandler CarruthandJon Ross-Perkins 1b0e2d3a4b Cleanups of SIMD code and document no Arm port. (#3325)
I spent (a lot) of time working to see if there was any profitable way
to port the SIMD code that scans for identifier length to Arm. There
isn't really. =/ While working on these, I made some cleanups to the
SIMD code that seemed worth landing, and added some benchmarks. All this
PR does is the cleanups, benchmarks, and documents that Arm isn't just
waiting to get attention but doesn't really have good options (so far).

For posterity, here are the core techniques I tried:

1) Direct 32-byte SIMD scanning using pair-wise add trees to build
   a 32-bit mask of valid identifier and then `clz` to compute the
   distance. This is a very good analog to the 16-byte SIMD structure
   used on x86-64. The pair-wise summing technique is the one used in
   simdjson for similar purposes.
2) A 16-byte SIMD scanning similar to the x86 version but using `shrn`
   to produce a 64-bit scalar bitmask with 4 bits per byte, and then
   scaling the bit-count distance.
3) Various hybrid versions of (1) and (2) with short scalar scans to
   identify short identifiers before paying the SIMD start-up cost.
4) A much fancier version of (1) that scanned 64-bytes at a time, but
   cached the resulting 64-bit mask and re-used it until exhausted.

Some good background on these techniques on Arm CPUs is in this blog
post:
https://community.arm.com/arm-community-blogs/b/infrastructure-solutions-blog/posts/porting-x86-vector-bitmask-optimizations-to-arm-neon

Sadly, both (1) and (2) were significantly slower than a scalar loop
over the bytes. Even (3) was consistently slower.

The only approach that came close was (4) and it was very *slightly*
slower in typical examples and very *slightly* faster in extremely
difficult cases like huge identifiers.

Ultimately, the only path I see (suggested by Dougall on a Mastodon
discussion of this whole problem space) is to take (4) to the limit of
computing an identifier-or-not bitmask *for the entire source file*
using a deeply throughput optimized routine (maybe as part of the line
scanning). That should be able to manage the high latency you end up
with when handling these patterns in SIMD on Arm.

The good news is that at least the M1 is *so* fast in the byte-scanning
loop that this isn't hurting nearly as much as I feared.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-10-24 08:32:21 +00:00
Jonathan B. CoeandChandler Carruth 8c28a0494e Add size="small" to test targets where advised (#3326)
Running `bazel test //...` reported:

```
Test execution time outside of range for MODERATE tests.
Consider setting timeout="short" or size="small".
```

This change adds size="small" to avoid such warnings being reported.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-10-24 06:52:00 +00:00
Richard Smith 7d9340880e Separate ClassType from ClassDeclaration. (#3329)
Retain the `ClassDeclaration` node to represent a syntactic declaration
of a class (including possibly a declaration of a generic class), but
use a separate SemIR node to represent the class type itself. This
allows us to give the two separate treatment.

The `ClassDeclaration` is still entered into the name lookup table for
its enclosing scope, but when it is named in an expression, the class
type is produced instead. When the class declaration is named in a
declaration name, it can be used to define members of the class, but an
expression that resolves to the class type cannot be used to define
members of the class.

In order to distinguish these cases, use `Name` rather than
`NameExpression` for the left-hand side of a `QualifiedName` parse node.
This removes the only use of the `Expression` form of a declaration
name, so that is also removed.

In the future, `ClassType` will also be used to describe types such as
`Vector(T)`, for which there is no corresponding `ClassDeclaration`.
2023-10-24 01:26:44 +00:00
Jon Ross-Perkins ce248239d4 Fix missing include for ids.h (#3328)
I think this is only noticeable in a more modular build, but we try to
keep that working.
2023-10-23 22:10:10 +00:00
Richard Smith 85e9642d18 Lower types in the order they were completed. (#3324)
This is a prerequisite for class support, where a class can be
referenced as a type before it becomes complete. For example, given:

```carbon
class A {
  fn F(a: A);

  class B {}
  var b: B;
}

fn A.F(a: A) {}
```

we need to lower `B` before we lower `A`, even though `A` is used as a
type first.

This will also start catching some cases where we don't require a type
to be complete despite using it, as we now only lower types that are
required to be complete.

Remove the poison values for struct and tuple literals. We don't need
those any more, because we never generate references to those literals
as values, and we don't have a type to use for them because we never
require the type of a literal to be complete, only the type of the
entity initialized by the literal, which can be different, for example
when initializing an array from a tuple literal or a class from a struct
literal.

This doesn't affect the output: `llvm::Type` objects that are not
referenced by an LLVM module don't affect the IR for that module, and
the order in which `llvm::Type`s are created doesn't affect anything
either.
2023-10-23 18:04:01 +00:00
josh11b 5971f80826 Add consistency test between NodeKind and Definition (#3322) 2023-10-21 00:13:40 +00:00
Geoff Romer 9e0d831092 Handle let with no closing semicolon (#3323)
Also add a test for the `var` case
2023-10-21 00:12:42 +00:00