In `Class` we have:
```
// The following members always have values, and do not change throughout the
// lifetime of the class.
// The following members are set at the `{` of the class definition.
// The following members are accumulated throughout the class definition.
// The following members are set at the `}` of the class definition.
```
`Interface` has similar, minus the "accumulated" members. I'm echoing
this, except `Function` has no `}` members; at present, nothing
differentiates between "started definition" and "completed definition",
unlike the other two.
Also removing a slightly inconsistent default value for decl_id.
Interface support is pretty skeletal so this may need additions later,
but I think it's still worthwhile to fill in the necessary bits now.
With this change, the expectation is then that everything we have right
now which _can_ be imported, is supported for import (at least for the
"current package, no overlap" case).
I'm proposing a different split, along the line of "what does this
relate to". I view impl.h as having started down this route. Moving the
inst store stuff to inst.h feels odd to me given how much else is there
right now, but maybe it's still the best approach. Some files only
contain a store, no structured class, but I felt the consistency in file
naming (without _store suffixes) might help.
I noticed that there are some empty `__init__.py` files in the repo,
with no immediately clear reason to have them. I asked [in
Discord](https://discord.com/channels/655572317891461132/655578254970716160/1211088334894534686)
and it looks like these are likely artifacts from when `lit` was used
for testing, now left over and obsolete from the migration. To keep the
repo tidy, this PR deletes the files.
By adding a constant to ClassDecl/InterfaceDecl, we're able to remove
name reference special-casing. Use TryEvalInst on the Decl to generate
the Type. For ClassDecl, then use the generated constant for
self_type_id.
To make the implementation simpler, make `PopWithParseNodeIf` return
`pair<NodeId, optional<value>>` rather than `optional<pair<NodeId,
value>>`. While wrapping the whole result in `optional` seems more
principled, it's significantly harder to work with.
I believe this PR is sufficient to pull in all current class features,
including the current bits of inheritance which have been implemented.
Because a class declaration can reference its own type, this creates an
incomplete type prior to constant loading.
Right now, the object representation is imported proactively, but
individual fields are left as ImportRefUnused. This means that member
functions and similar will only be imported if called.
This also adjusts how function parameters are being handled, to match
the expectations of Self param structure.
When formatting, I'm starting to look into constants. Otherwise we get
"unexpected instref".
Overall, there are a few things that may be worth further discussion:
- The lack of a constant corresponding to the ClassType on ClassDecl is
inconvenient -- I'd like to see how zygoloid feels about trying to
restructure this. i.e., I'm setting a constant in order to be able to
track things down later, it'd be nice if the normal IR did this simply
for consistency, or if we were able to combine these rather than having
separate instructions.
- Should we shift the parse node tracking further, and go with a setup
wherein imports can embed import references into that? e.g., negative
values go to another array which includes a ImportIRId for printing
diagnostics, replacing the invalid NodeId.
- Can the formatter switch to a more general scan of instructions for
naming, to eliminate the ImportRef constant approach added here?
- GetExprValueForLookupResult special-casing instructions felt
surprising, I might see if there's a way to restructure to avoid that.
But I think these issues are things that can be separated out.
This undoes parts of #3515 in order to allow PushGlobalInit to be called
when the initializer is called, instead of at the end of the binding
pattern. The current approach is fragile because supported patterns will
become more complex. We also will likely want similar support in `let`,
which puts the initializer first, so this offers a consistent approach
for both.
[Looking
back](https://discord.com/channels/655572317891461132/655578254970716160/1184237511766179840),
this is more or less the second option in that message, but using the
PeekNextIs to avoid vagueness about what's being popped first.
Note I'm putting in PeekNextIs for what I'm hoping will be a pretty
narrow use-case. I could've added depth arguments to the Peek functions,
but that would've rippled through a number of APIs and it's not clear to
me that this has generic utility. I mean, right now it could just be
PeekNextIsVariableInitializer, since it's only optional in that case.
This works by creating a faux FunctionDecl in the context of the current
IR, which seems to be working for function calls. Deduced params are
there, but won't really be tested until classes are up and running. Also
I may need to look further at return_slot_id to ensure it's working. But
the basics, I think, are here.
Reorganizes some other ImportRef work from `has_unresolved` that'd
relied on manual calls to a more detection-based `HasUnresolved`
approach that doesn't require as much checking.
Unqualified names don't handle scopes the way that typical names do, so
a name conflict with a namespace needs to be handled specially. I'm
still favoring keeping code close as much as possible, particularly
since long-term this syntax will probably shift to be more consistent.
For now I'm just flagging when we shouldn't push scopes, so that
MakeUnqualifiedName doesn't need to clean up.
Note, a different approach would basically be:
```
PushScopeAndStartName
ApplyNameQualifierTo
result = decl_name_stack_.back();
decl_name_stack_.back().state = NameContext::State::Finished;
PopScope
return result;
```
But that approach feels worse to me, due to the additional stack
manipulations and the need to duplicate some of the FinishName logic
just to be able to pop the scope that didn't really need to be added.
Adds `BindAlias` with a hybrid of `BindName` and `NameRef` semantics. I
think it's slightly closer to `BindName` because it introduces a name,
so I'm going more in that direction. This also matches the need for
`bind_name_id` with imports on enclosing scopes.
Note, only things that look like a name reference are being allowed on
the RHS of `alias`. This includes builtins that look like name
references, such as `bool`, but not ones that turn into values
underneath, such as `false`.
- File::StringifyTypeExpr now has a case that hits the ImportRefUsed
TODO, so implementing that. I think the `static` approach will be
helpful in ensuring there aren't access bugs, particularly when future
support is added.
- `let` wasn't adding to exports because it doesn't use
`decl_name_stack` the way `var` does. This now adds to exports, but we
might want to unify logic for issues such as this.
- BuildImportRefUsedValueRepr is now called for another inst kind, and
it seemed like calling back to BuildValueRepr was the best way to
resolve this. I don't *think* that's going to cause recursion.
Note we may also want to do this with NameId, maybe some other things,
but the TypeId use is pretty broad and repetitive -- I thought I'd start
with it first.
`GlobalInit` block is now static block within a `SemIR` which will be
used to emit initialization instructions for variables in the `Package`
scope.
inst_block_stack now has additional methods to handle `GlobalInit` block
separately, this block can be popped without being finalized allowing to
accumulate between all instances of variables.
At the end of the `check` phase, if this block is not empty , the
function `__global_init` will be added with this block being inserted
into it.
This block is pushed to `inst_block_scope` at the end `BindName`,
allowing instruction to be emitted into it, then popped at the semicolon
(VariableDecl).
This significantly changes the `SemIR` output, that's why this commit
updates a lot of the test cases.
This provides support for Const, Pointer, Struct, and Tuple types. It
does not cover Class, Function, or Interface which have their own Id and
are tracked slightly differently.
They don't have names, but using the DeclNameStack anyway keeps our
behavior more consistent, and keeps track of the enclosing name scope
and the prior state of the scope stack for us.
Depends on #3683.
Collect the contents of an `impl` into a scope, and start doing very
basic checking for `impl` declarations and definitions.
This change adds two new `Id` types to the set of type that `NodeStack`
supports -- `ImplId` and `NameScopeId`.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Add a type representing a non-discriminated union of IDs. Refactor the
node stack to use it. Plus a few other refactorings aiming to clean up
and simplify the code. The overall goal here is that adding a new kind
of ID, node category, or instruction should only require changing one
place in the node stack rather than a bunch of different changes.
One minor functionality change: crash backtraces now use the correct
type for IDs when dumping the node stack rather than using `InstID`
printing for all but one case.
Right now, ConstantValueStore defaults to having unknown values use
NotConstant. This generally works for the current IR, but with imports
we're expecting sparse entries which are generally unknown -- and
distinguishing between NotConstant and simply unset would be helpful. As
a consequence, add Invalid.
We discussed whether to simply have ConstantValueStore default to
Invalid going forward, or to make the default flexible. The upside to
the former is consistency, the upside to the latter is that it should
result in fewer Sets when operating on the current IR (which will more
frequently have known non-constant values). This PR offers both
approaches in separate commits, but I somewhat lean towards the latter
for fewer array resizes.
Note, a totally different approach would be to use a different class
(not ConstantValueStore) for imported IRs -- then the default of Invalid
versus NotConstant would be type-dependent. However, I expect we're
going to want to do at least somewhat consistent lookups, and using the
same ConstantValueStore for both cases allows avoiding a virtual
interface or templating. Also, I'm hoping to only maintain the
ConstantValueStore for an imported IR as part of Context (not File),
which would mean the SmallVector storage overhead is ephemeral,
mitigating one of the potential advantages of using a different type for
imported IRs.
Use two different nodes for "<type> followed by `as`" and "<type>
omitted before `as`, use `self`", so it is easier to determine which
case. Later the second case will push the type id for `self` onto the
node stack, making the two paths more similar.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Consume the components of the `impl` declaration, and set up scopes for
the child elements. We don't yet build a representation for the impl
itself.
Also, add an interface type value. This is necessary so that we have a
value for the expression on the right-hand side of `as` in an `impl`.
Previously, we created scopes for implicit parameter lists and tuple
patterns, but that meant that bindings went out of scope too soon. We
now keep them in scope until the end of the enclosing declaration. This
is accomplished by pushing a scope for parameters when we handle a name
that might have them, and then popping the scope again if it turns out
that there were no parameters.
For a case such as:
```carbon
fn A(T:! type).B(U:! type).F(x: T, y: U) {
var z: T;
}
```
... we now have the following scopes in the stack:
- A parameter scope containing `T`.
- A class scope for `A(T:! type)`.
- A parameter scope containing `U`.
- A class scope for `A(T:! type).B(U:! type)`.
- A parameter scope containing `x: T` and `y: U`.
- A function body scope containing `z: T`.
The innermost scope when check processes a declaration of a function,
class, or similar is now often a parameter scope rather than the
enclosing scope in which the class or function is declared, so the
target scope is now passed explicitly into the modifier checking code
that wants to inspect that enclosing scope.
I'm basically just nudging down the path I think is right here. Adding a
small bit more support, but also more tests to capture cases that I
think will need to be verified as working.
With diagnostics like "Value of type `<function>` is not callable.",
that's because it expects a FunctionDecl but is instead finding a
ImportRefUsed. I'll need to work out the necessary support for a
callable function.
One more (hopefully last) rename on the Import instruction renaming.
I was kind of tempted to rename to just "IRId", since the IRs aren't all
imports. However, this felt easier to read, and a better choice than
CrossRef because it's more consistent with the other ways imports exist
in code. (even if IRs aren't all imports, most use-cases are derived
from imports)
Note though that import_irs may include IRs not just from direct
imports. Beyond the builtin IR, I'm thinking that for indirect imports,
or the prelude, we may end up adding them. e.g., so that constants can
be generated for indirect imports and still correspond to a directly
known IR, and for a given IR that's indirectly imported multiple times
to be deduplicated locally. I'm not there yet, I'm just mentioning this
to help give background for naming thoughts.
This required adding a few headers that were found transitively before,
but not too many. This is sadly a fairly manual process of opening every
file in my IDE, but I think I got everything in `//common` and
`//toolchain`.
There are a few cases where technically we don't need `foo.h` to be
included into `foo.cpp`, but I've forced those to stay with a pragma.
I've tried to catch the places where we can cut deps in Bazel as well,
but not sure I got all of those.
I had been noticing these in other PRs and it seemed better to isolate
the change.
The categories `Expr`, `MemberName`, `Decl`, `Statement`, and `Modifier`
are usable since they are associated with a consistent `IdKind`. The
mapping to `IdKind` for NodeKinds that have those categories are no
longer listed explicitly, ensuring that the `NodeCategory` mapping is
the source of truth.
Also: fixes the category of the `FunctionDefinitionStart` and
`ArrayExprStart` node kinds.
Note: I've added [a section on defining constexpr constants to the
Toolchain architecture
doc](https://docs.google.com/document/d/1RRYMm42osyqhI2LyjrjockYCutQ5dOf8Abu50kTrkX0/edit?resourcekey=0-kHyqOESbOHmzZphUbtLrTw&tab=t.0#heading=h.f7682a2tpvxr).
FUTURE:
* We should switch `TuplePattern` to put an `InstId` on the `NodeStack`
instead of an `InstBlockId`, so we can handle the pattern category.
* We should make a category for names to replace uses of the `NameId`
`IdKind`.
* We should make use of these new APIs more, and propagate more-precise
types through the codebase.
QUESTION: Should I use a different approach for determining the number
of members of the `NodeKind` enum?
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This is a bit of a cleanup; I probably should've just renamed CrossRef
instead of adding ImportRefUsed.
Adding `is_builtin` to InstId is more about providing a standard API for
the check, which I expect to add a little more of.
Shifts import tests to validate that the BuildValueRepr CHECK isn't
accidentally hit.
Recent runs of `clang-tidy` for me started showing more errors, and this
is a collection of changes to address them.
First, I've systematically applied the disabling tag to all C++ rules
under //explorer/... with `buildozer` so we don't spend time analyzing
this code or reporting errors from it. Not sure this was strictly
necessary, but it seemed like a nice consistency improvement.
Next, I disabled a buggy check for missing `default` cases in
`switch`es. It seems to get confused by the fancy conversions in our
`enum_base.h`. We don't miss much with this as the Clang compiler
warnings for `switch` catch most of our actual bugs. I also removed the
local disabling of this now that it is turned off centrally.
Lastly, I added error checking to two file descriptor manipulating calls
in the `file_test` infrastructure.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Builds on #3656.
Under the prior "lazy" model, we had been planning to copy instructions.
Under the current "unused+used" model, we're restricting that to
constants. I'm trimming back some of the ResolveIfImportRefUnused logic
because it was more appropriate for the former model.
Right now, I'm adding ImportRefUsed direct creation in import.cpp for
namespaces. I think I'll need to do something similar in
RseolveIfImportRefUnused... I'm still trying to think about how to
manage type information there (which needs to come in as an import
reference itself, and probably have some amount of deduplication before
forming a TypeId). So ImportRefUnused lacks a type because I'm hesitant
to aggressively load it, whereas ImportRefUsed should *always* have a
type but it's just an error while I think things through.
For reference, ImportRefUnused and ImportRefUsed are mainly split in
order to track the boolean "used" without making fundamental
modifications to Inst for bit packing (this effectively instead packs a
bit into InstKind). AnyImportRef currently excludes the type because
it's mainly for diagnostic printing at the moment. It could end up with
a TypeId that would end up Invalid for ImportRefUnused, though it could
also be that the TypeId is only accessed when using ImportRefUsed
explicitly.
This makes some changes to the formatter so that ImportRefUnused and
ImportRefUsed will both be labeled as "import_ref" with an "unused ->
used" argument change in textual IR, but is otherwise not changing
logic.
I'd excluded these initially just because I was thinking towards copies,
but under the current model I'm trying to catch all the decl types just
for consistency. Note references will still be a TODO error
(LazyImportRef is already tested for this, it just didn't feel necessary
to add individual tests while I try to sort out behavior).
Fixes an oversight where declarations in an entity's scope were being
added to the list of exports.
Note I'm trimming some Import API arguments as now-unused.
Per discussion with zygoloid, namespace declarations will merge both
when repeated in a given file, and across files. This echoes how forward
declarations of other entities are allowed to repeat. When merging with
an imported namespace, fill in the parse node so that future diagnostics
point at the declaration in the same file rather than a declaration in a
different file, just for locality.
Although it might be desirable to issue a diagnostic when a namespace
declaration is repeated within a given file, that's similarly true in
other cases, but may be more desirable as a tidy-style issue rather than
preventing code from compiling. Allowing the repetition also makes the
import versus non-import cases more consistent: the namespace
declarations merge regardless of the source.
To avoid bouncing through `constant_values()` to determine whether a
type is symbolic or template, store the `ConstantId` on the `TypeInfo`
not just the `InstId`.
In addition to propagating the symbolic / template phase, this also
propagates whether a type contains an error, resulting in our no longer
producing types such as `<error>*` -- these now evaluate to simply
`<error>`. While this makes our types less precise after an error, it
also removes some follow-on diagnostics, so it seems to be an
improvement on the whole.
Building on #3636 which handles the general import case, add special
casing for namespaces. Namespaces can be combined cross-IR, so it's a
little more complex.
The implementation adds import_id to the Namespace instruction as a
reference to find the original using the normal structure. This is
achieved by moving the name_id to NameScope to free up space.
---
I considered a few alternatives...
I considered adding import_id to NameScope, but:
1. It's more consistent with things such as Function or Class that
provide name_id on the info object rather than the instruction.
2. I thought it more likely that there would be more NameScope cases
that might want a name_id rather than the import_id, since ImportRef
will typically be used.
I considered putting the import source (cross-ref IR id + inst id) on
the NameScope versus a separate ImportRef, which seems like the
strongest argument towards the NameScope approach because it removes an
instruction. That just felt inconsistent though, and the overhead of
instruction-per-imported-namespace should be low (theoretically few
namespaces should be used). Plus I feel a bit odd adding two
generally-unused ids to NameScope.
A specialized Namespace structure could also have been created to store
the import_id, but that would add an indirection to the NameScope.