C++ interop used `GetFormatted` to turn a Carbon name into a C++ name.
For an identifier that is a Carbon keyword, such as `base`,
`GetFormatted` adds an `r#` prefix, so the C++ name came out as `r#base`
instead of `base`. This showed up in three places:
- A failed call to the C++ function `base` was reported as "no matching
function for call to 'r#base'".
- An exported Carbon function was mangled as `_ZN6CarbonL6r#baseEv`
instead of `_ZN6CarbonL4baseEv`.
- Its thunk was named `r#base__carbon_thunk` instead of
`base__carbon_thunk`.
All three now use `GetClangIdentifierInfo`, which `export.cpp` already
uses for other C++ names. The name is always an identifier in these
places: an overload set is only imported by looking up an identifier in
C++, and only functions declared in Carbon are exported, which are
always named by an identifier. So they check for an identifier, as the
export of a class name already does, rather than falling back to
`GetFormatted`. This removes a TODO about naming a `NameId::CppOperator`
overload set, which can't happen: C++ operators are resolved in
`operators.cpp` without an overload set.
Assisted-by: Claude Code
C++ pointers to Carbon methods are supported using the same mechanisms
as for Carbon functions: we export the method to C++ and then form a
pointer to the exported method. This change does not support invoking a
C++ method pointer from Carbon, which will be more complicated because
we need to model instance binding on the Carbon side.
Base classes will be destroyed in a dedicated change, so we can
trivially confirm that the base is being destroyed.
This is a partial implementation of #7362.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
A specific's definition block may contain `InstId::None`, indicating it
has no constant value. Display this as `<not constant>` instead of
`invalid`, to avoid it sounding like an error.
Noticed these issues when trying to query raw semir with
[yq](https://github.com/mikefarah/yq)
- Removed additional "}" within PrintClassFields
- Removed additional "}" after is_frozen_period_self
- Surround `int(value: 123)` in quotes to avoid issues parsing the
additional ':'
- Updated DeclaredFacetType::Print to have valid YAML structure
- Constraints are added to inline list (i.e. wrapping in `[ ... ]`)
- Switched Outer seperator to Comma.
The last two may be up for debate, maybe there's no intention for raw
semir to be true valid yaml but I think its quite useful, for example I
can now do the following:
```sh
carbon compile examples/sieve.carbon --phase=check --dump-raw-sem-ir | yq '
.sem_ir |
select(.names as $n | .classes[] | .name as $k | $n[$k] == "Sieve") |
. as $ir |
$ir.classes[] |
select(.name as $k | $ir.names[$k] == "Sieve" and .body_block_id != "inst_block_empty") |
. as $c |
($ir.inst_blocks[$c.body_block_id][] as $i | $ir.insts[$i])
'
```
outputs:
```yaml
{kind: ImplDecl, arg0: impl41000000, arg1: inst_block41000008}
{kind: ImplSelfWitness, arg0: inst41000014, arg1: specific_interface41000000, type: type(inst(WitnessType))}
{kind: FunctionDecl, arg0: function41000000, arg1: inst_block41000013, type: type(inst41000030)}
{kind: FunctionDecl, arg0: function41000001, arg1: inst_block41000023, type: type(inst4100005B)}
{kind: FieldDecl, arg0: field41000000, arg1: region41000002, type: type(inst41000065)}
{kind: CompleteTypeWitness, arg0: inst41000067, type: type(inst(WitnessType))}
```
This allows us to collect rewrites from named constraints and use them
to initialize the witness table for an `impl as` statement.
Only rewrites from extend constraints are tracked, as other constraints
should turn into equality constraints, as they don't modify the witness
table.
This change introduces a new inst kind `CppFunctionPointerType`, which
represents an imported C++ function pointer that can be invoked from
Carbon (support for forming such a pointer from a Carbon function is in
a follow-up PR). This is implemented by treating the operation of
invoking a function pointer in C++ as if it were a call to an `__invoke`
method on the function pointer type, and extending the existing
function-import logic to support importing this fictitious method.
---------
Co-authored-by: Nicholas Bishop <nbishop@nbishop.net>
This is a step toward supporting C++ function pointer types, which need
to be
imported and thunked in much the same way as C++ functions, but have a
different
underlying representation. `CalleeFunctionInfo` gives us a way to
abstract
away the representation differences, so expressing import and thunking
in terms
of `CalleeFunctionInfo` lets us reuse that code for function pointers.
Actual support for function pointers will come in a follow-up PR, but
the
API design choices I've made here are driven by that use case.
Implements `typeof(expr)`, as described in #7697. `expr` is treated as
an unevaluated operand, and is kept in an `ExprRegion` separate from the
enclosing scope.
Assisted-by: Claude via Antigravity
The AST doesn't really matter here, since it's NRVO, there's no actual
code to generate to return the value. But having a correct AST does
address at least one clang false-positive diagnostic:
```
error: stack memory associated with local variable 'return_storage' is returned
3 | fn F() -> bool {
| ^
```
The `self` parameter block is presumably a relic from when `self` was in
the implicit parameter list; now it's just SemIR bloat.
This also factors out some common code between `self` and the other
parameters.
We use `LocId`s to refer to physical locations in source code. Those
don't exist for toolchain-generated entities, so we instead choose a
related location that can stand in for a physical location. This inlines
the generated entities' constants into their points of use, and reduces
the total amount of generated SemIR.
This is especially important for calls to `Destroy.SelfDestruct` because
these are automatically generated when any destroyable object reaches
the end its lifetime.
Reconstruct the call syntax from the callee's explicit parameter
patterns, the callee specific, and the call arguments.
Assisted-by: Claude via Antigravity.
Removes support for unspecified default values. Fixes canonicalization
of the `DefaultValuePattern` instruction by making them immutable after
they are issued, and by removing the `DefaultValueId` operand which
wasn't being canonicalized.
Add support for deferring initialization as a template action, and
performing the deferred initialization during template instantiation.
This is substantially more complex than other conversion actions, for
two primary reasons:
* The initializer in the generic may have storage arguments as inputs.
We model an initializing expression as having a "slot" where
initialization writes the location that should be initialized by that
initializing expression, and that needs to be an output of the
initialization action.
* Initialization from a tuple or struct literal needs to recurse into
that literal, and the literal will have been spelled in the generic,
meaning we don't have an `InstId` that can be used to name the specific
version of the initializer as input for nested conversions.
These issues are addressed by introducing two new features to the action
machinery:
In addition to `InstAction`, we now have `MultiInstAction`, which is an
action that produces a tuple of instruction values instead of a single
instruction value. Initialization actions produce one instruction for
the final result, which is spliced at the point of initialization, plus
one instruction for each storage argument, which are spliced into the
storage argument slots in the original generic. During initialization,
if we find one of those splices in the storage argument of an
initializing expression, we return the new storage argument back to the
initialization action to be included in the specific, instead of
overwriting the storage argument in the generic.
Actions whose `PerforrmAction` takes a `SpecificId` as input no longer
perform automatic refinement of their operands to specific instructions.
Instead, the action is given control over when and where it performs
that refinement. In `InitializeAction`, we use this freedom to form a
`SpecificInst` for the initializer in the primary output block, and form
a `SpecificInst` for the target in the target block. When detecting
whether we are initializing from a tuple or struct literal, we step over
the `SpecificInst` and track its `SpecificId`, and if necessary create a
new `SpecificInst` wrapping the sub-initializer when we recurse into the
nested element conversion.
Assisted-by: Claude and Gemini via Antigravity
This change partially implements [PR #7362], which revises how objects
are destroyed. It is a partial implementation for two reasons:
1. This change moves `Destroy.Op`'s current behaviour into
`Destroy.SubobjectDestroy`, but it doesn't add support for objects with
non-trivial destruction.
2. `Destroy.SubobjectDestroy` is a workaround for `require impls
SubobjectDestroy`. We aren't able to use the latter until the dependents
add their requirements' implementations to their own witness tables.
[PR #7362]: https://github.com/carbon-language/carbon-lang/pulls/7362
When a generic impl uses a default or final fn, it picks the specific
function value out of the interface to put in the witness table.
However, because this is done by modifying an existing instruction
block, the generics machinery has no hook to convert the function
constant into an attached constant, and because it was found in a
specific for a different generic, the constant inst will be unattached.
Fix this by manually mapping to an attached constant inst in the current
generic when building the witness table.
The type `type` is now a `FacetType` inst with no constraints. This
brings the model implemented in the toolchain into better alignment with
the language design. The `SemIR::TypeType` struct remains as a scope for
holding the `TypeInstId`, `ConstantId`, and `TypeId` constants, but is
not an `InstKind` anymore.
The `TypeType` inst looks a lot like singletons, but there are many
`FacetType` insts so it doesn't quite fit that model. So we put it
alongside singletons with a fixed inst id but refer to it as a more
general "builtin" inst that is not a singleton.
`Namespace::PackageInstId` is similar, and we group it with `TypeType`
conceptually as another builtin instruction with a fixed id.
No conversion is needed anymore to use a `type` as a facet, since types
also have a `FacetType` type. This simplifies and removes a number of
helpers and branches throughout the code.
The `TypeType` inst is now part of the constant store, so we end up
printing it in the constants block in every test. But it's also named
`type` rather than `%type` to preserve the majority of existing
formatting behaviour, though this does look different from other
constants.
Assisted-by: Opus 5 was used to generate a first draft and validate the
refactoring. Though nearly everything non-trivial the tool wrote has
been modified or rewritten.
Per #7737 we move the default value `InstId` storage from
a block in the `SemIR::Function` data structure to a
`SemIR::File` scoped `ValueStore`.
Moves the default value consistency checking to the general
merge argument pattern matching logic, which changes the
error message issued to the generic one.
In the `EntityName` for a binding, preserve the `TypeInstId` describing
how the type was written. When a diagnostic refers to that type via
`TypeOfInstId`, use the type-as-written in the diagnostic rather than
the canonical type.
Assisted-by: Claude Opus via Antigravity
When evaluating a deferred member access action, the scope stack cannot
be relied on, so `LookupUnqualifiedName` cannot be used in
`GetHighestAllowedAccess` to get the `Self` type.
Instead, store the `Self` type in the `Context` when evaluating a
method, and use that in `GetHighestAllowedAccess`.
If a Carbon class overrides virtual functions from a C++ base class but
is never referenced from C++, it is never exported to Clang. During
lowering, `BuildVtable` then fails to find a `CXXRecordDecl` and crashes
when attempting to get the vtable from Clang's code generator.
Ensure dynamic classes with foreign vtables are exported to Clang when
completing the class definition in `CheckCompleteClassType`, and look up
`first_decl_id()` in `BuildVtable`.
Fixes#7721
---------
Co-authored-by: Dana Jansens <danakj@orodu.net>
Per feedback on #7665, this PR switches the default value
table storage from canonical constant inst_ids to
non-canonical.
Furthermore, this PR simplifies the default value support in
check by requiring that the first owned declaration of a
function completely specify all of its default values.
Updates the diagnostic code and tests to reflect this new
stricter requirement.
This action was created to wrap any `MetaInstId` operand of an action
instruction. This served two purposes:
1) It had a special hook in `OperandIsDependent` to allow it to be
performed while it had a dependent operand (the reference to the
instruction in the generic).
2) It created a `specific_inst` so that the downstream action saw an
instruction in the specific instead of one in the generic.
These are both replaced: the special case in `OperandIsDependent` for
`RefineInstAction` is replaced by a special case for `MetaInstId`s in
general, and the `SpecificInst` is now created as part of performing the
downstream action, rather than as a separate step carried out
beforehand.
This simplifies the produced SemIR and reduces the number of splices
significantly. It also prepares us to handle actions like
initialization, where we don't actually want to create `SpecificInst`s
immediately in the location where the action is performed, because they
actually belong somewhere else in the IR.
Assisted-by: Claude Opus 5 and Gemini via Antigravity
Instead of always printing types as canonical, attempt to find a sugared
type where possible, and include that type in the diagnostic. We can
only do this when given the instruction whose type is being printed
(`TypeOfInstId`) rather than the canonical type ID.
Initial support here is intentionally minimal: just looking through
calls to the callee's declared return type, and looking through pointer
dereferences and corresponding pointer types, to build out the initial
infrastructure. More cases can be added later; this degrades gracefully
to using the canonical type if a better type can't be found.
Assisted-by: Claude Opus 5 via Antigravity
We use the same conversion codepath to handle both qualification
conversions and derived-to-base conversions, because we allow both to be
performed at once. However, we were previously modeling the
qualification conversion as happening *first*, and producing a result
whose type is the target type of the overall conversion (that is, the
base class type). That led to bogus SemIR, where a `Derived` -> `const
Base` conversion would first have a "compatible" conversion from
`Derived` to `const Base`, *then* an access of the base subobject (of
type `const Base`, within an object of type `const Base`).
We now reverse the order: first we do a derived-to-base conversion,
which already has logic to preserve qualifiers, and then we do any
necessary qualification conversions on the result to reach the overall
target type.
In passing, we now skip forming the `as_compatible` instruction at all
for a pure derived-to-base conversion that has no qualification
conversion, simplifying the SemIR by one instruction in the common case.
Fixes#7731, at least under `--share-cpp-ast` which is expected to be
the future direction.
When --share-cpp-ast is enabled and any compilation unit has C++
imports, include all compilation units in the shared CppDomain inputs
and assign the domain to every unit. In ImportCpp, when a unit has no
direct C++ imports but is covered by a shared CppDomain, initialize
its C++ AST context and import namespace. This ensures units without
direct C++ imports have access to the C++ AST and code generator when
instantiating generics or referencing declarations from units that do.
Assisted-by: Antigravity with Gemini
Use a single `SemIR::Function` per `Core` interface method, whether it's
generated locally or imported. This prevents generating duplicate
functions, which lead to different types when the witness appears in a
`FacetValue` as part of a specific for a class.
We use a `CanonicalValueStore` of `GeneratedFunction` objects that allow
finding an existing FunctionId for a `Generated` special function before
(re-)generating it. Mangling for `Generated` functions is also moved to
use the values from the `GeneratedFunction`'s canonicalization key, so
that we have a consistent source of truth for the unique ID of a
`Generated` function across all files.
New tests are in
`toolchain/check/testdata/impl/custom_witness/destroy.carbon`.
### Description
When looking up default initializers for class elements in
`ConvertStructToClass`, the compiler previously assumed that every
member looked up from the class scope was a `FieldDecl` and called
`GetAs<SemIR::FieldDecl>` directly.
For a derived class with a base class, looking up `base` returns a
`BaseDecl`, which caused a `CHECK` assertion failure when casting to
`FieldDecl`. Use `TryGetAs<SemIR::FieldDecl>` instead so non-`FieldDecl`
entries like `BaseDecl` are recognized as having no default initializer,
cleanly diagnosing that the `base` field is missing.
Fixes#7722
Assisted-by: Google Deepmind Antigravity
Avoid CHECK failure when performing member access on a runtime type
value. We will just fail to find the CanonicalFacetOrTypeValue and then
fail lookup.
Test passing a template argument or a symbolic argument to a template.
Fix a bug in the template case where we'd crash when instantiating a
dependent discarded expression, because conversion produced an
`InstId::None` which the actions machinery did not expect and crashed
on.
Updates the pattern matching code to support unspecified default values.
Adds logic to decl and def merge code to diagnose mismatches in defaults
if specified in both places, or if let entirely unspecified.
Per https://github.com/carbon-language/carbon-lang/pull/7521.