We use `LocId`s to refer to physical locations in source code. Those
don't exist for toolchain-generated entities, so we instead choose a
related location that can stand in for a physical location. This inlines
the generated entities' constants into their points of use, and reduces
the total amount of generated SemIR.
This is especially important for calls to `Destroy.SelfDestruct` because
these are automatically generated when any destroyable object reaches
the end its lifetime.
We've made quite a few changes to how associated constants are
implemented since this was written; update the doc to match.
Assisted-by: Claude via Antigravity
Reconstruct the call syntax from the callee's explicit parameter
patterns, the callee specific, and the call arguments.
Assisted-by: Claude via Antigravity.
* Adds function begin/end logic for non-trivial `SubobjectDestroy.Op`
* Reorganises cases in `MakeSubobjectDestroyOpBody` to use `CARBON_KIND`
so each case can be made as an independent change.
As I was starting to try to play around with carbon, I wanted nvim
integration. Unfortunately, a lot of the code in the integration had
bitrotted, but not unbearably so, so I updated it 😄
It works now! Syntax highlighting and the LSP work out of the box with
nvim 0.12 (and should also work on nvim 0.11)
Hopefully all good that I did All The Commits, I come from the school of
"every commit should do one thing"
Closes#7821
---------
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Removes support for unspecified default values. Fixes canonicalization
of the `DefaultValuePattern` instruction by making them immutable after
they are issued, and by removing the `DefaultValueId` operand which
wasn't being canonicalized.
If a line or column is unknown, we internally represent it as line or
column -1. When mapped from our 1-based numbering to LSP's 0-based
numbering, it comes out as -2, which is out of range since LSP requires
lines and columns to be >= 0.
Detect this case and produce a fallback.
Assisted-by: Claude via Antigravity
The '...Exec...' version of this benchmark runs subprocesses in a tight
benchmarking loop which seems like the likely culprit for timeouts we're
seeing on GitHub. Reduce testing to only a single one of those
benchmarks to hopefully reduce the frequency.
Add support for deferring initialization as a template action, and
performing the deferred initialization during template instantiation.
This is substantially more complex than other conversion actions, for
two primary reasons:
* The initializer in the generic may have storage arguments as inputs.
We model an initializing expression as having a "slot" where
initialization writes the location that should be initialized by that
initializing expression, and that needs to be an output of the
initialization action.
* Initialization from a tuple or struct literal needs to recurse into
that literal, and the literal will have been spelled in the generic,
meaning we don't have an `InstId` that can be used to name the specific
version of the initializer as input for nested conversions.
These issues are addressed by introducing two new features to the action
machinery:
In addition to `InstAction`, we now have `MultiInstAction`, which is an
action that produces a tuple of instruction values instead of a single
instruction value. Initialization actions produce one instruction for
the final result, which is spliced at the point of initialization, plus
one instruction for each storage argument, which are spliced into the
storage argument slots in the original generic. During initialization,
if we find one of those splices in the storage argument of an
initializing expression, we return the new storage argument back to the
initialization action to be included in the specific, instead of
overwriting the storage argument in the generic.
Actions whose `PerforrmAction` takes a `SpecificId` as input no longer
perform automatic refinement of their operands to specific instructions.
Instead, the action is given control over when and where it performs
that refinement. In `InitializeAction`, we use this freedom to form a
`SpecificInst` for the initializer in the primary output block, and form
a `SpecificInst` for the target in the target block. When detecting
whether we are initializing from a tuple or struct literal, we step over
the `SpecificInst` and track its `SpecificId`, and if necessary create a
new `SpecificInst` wrapping the sub-initializer when we recurse into the
nested element conversion.
Assisted-by: Claude and Gemini via Antigravity
This change partially implements [PR #7362], which revises how objects
are destroyed. It is a partial implementation for two reasons:
1. This change moves `Destroy.Op`'s current behaviour into
`Destroy.SubobjectDestroy`, but it doesn't add support for objects with
non-trivial destruction.
2. `Destroy.SubobjectDestroy` is a workaround for `require impls
SubobjectDestroy`. We aren't able to use the latter until the dependents
add their requirements' implementations to their own witness tables.
[PR #7362]: https://github.com/carbon-language/carbon-lang/pulls/7362
Add hover cards and jump to declaration / definition / reference for the
formatted SemIR that appears in check tests. This is done by adding a
heuristic "parser" for SemIR to the language server. The
cross-references are strictly best-effort, since this is just a tool for
Carbon developers, not a user-facing facility.
A couple of other changes made along the way:
* file_test tests with an AUTOUPDATE-SPLIT no longer look for CHECK:
lines outside that split. This was motivated by the tests for this new
facility including CHECK: lines as part of the test input.
* An agent skill for working on the language server, tracking some
things that cost Claude time when working on this.
Assisted-by: Claude via Antigravity
It had a few subtle GNU extensions in it that didn't work on macOS.
There are simple portable alternatives so it was easy to adapt.
Assisted-by: Antigravity with Gemini
When an LLM tool is used to prototype a change, one of the major tasks
it must undertake is validating filetest output changes. This skill
helps to ground that validation in some best practices and explain what
sorts of changes should or should not be expected, and how to judge
STDOUT vs STDERR changes.
Assisted-by: Opus 5
When a generic impl uses a default or final fn, it picks the specific
function value out of the interface to put in the witness table.
However, because this is done by modifying an existing instruction
block, the generics machinery has no hook to convert the function
constant into an attached constant, and because it was found in a
specific for a different generic, the constant inst will be unattached.
Fix this by manually mapping to an attached constant inst in the current
generic when building the witness table.
When there's a wildcard in a path, git treats the path as matching
exactly, unless the path also ends in a wildcard. So
`toolchain/*/testdata` only matches the testdata directory names,
whereas `toolchain/*/testdata/*` matches all the files under them.
The type `type` is now a `FacetType` inst with no constraints. This
brings the model implemented in the toolchain into better alignment with
the language design. The `SemIR::TypeType` struct remains as a scope for
holding the `TypeInstId`, `ConstantId`, and `TypeId` constants, but is
not an `InstKind` anymore.
The `TypeType` inst looks a lot like singletons, but there are many
`FacetType` insts so it doesn't quite fit that model. So we put it
alongside singletons with a fixed inst id but refer to it as a more
general "builtin" inst that is not a singleton.
`Namespace::PackageInstId` is similar, and we group it with `TypeType`
conceptually as another builtin instruction with a fixed id.
No conversion is needed anymore to use a `type` as a facet, since types
also have a `FacetType` type. This simplifies and removes a number of
helpers and branches throughout the code.
The `TypeType` inst is now part of the constant store, so we end up
printing it in the constants block in every test. But it's also named
`type` rather than `%type` to preserve the majority of existing
formatting behaviour, though this does look different from other
constants.
Assisted-by: Opus 5 was used to generate a first draft and validate the
refactoring. Though nearly everything non-trivial the tool wrote has
been modified or rewritten.
Per #7737 we move the default value `InstId` storage from
a block in the `SemIR::Function` data structure to a
`SemIR::File` scoped `ValueStore`.
Moves the default value consistency checking to the general
merge argument pattern matching logic, which changes the
error message issued to the generic one.
In the `EntityName` for a binding, preserve the `TypeInstId` describing
how the type was written. When a diagnostic refers to that type via
`TypeOfInstId`, use the type-as-written in the diagnostic rather than
the canonical type.
Assisted-by: Claude Opus via Antigravity
`scripts/jj_push.sh` takes the same arguments as `jj git push`, runs
prek over the commits that push would send, and pushes only if they
pass. It learns what is being sent by running `jj git push --dry-run`
and reading back the plan, so `--bookmark`, `--change`, `--all` and the
rest work without reimplementing how they select commits.
Hooks that rewrite files need a commit to write into, so the checks run
with the working copy on top of the commit being pushed. When the
working copy is already an empty commit there, which is the common case,
it is used directly; otherwise one is created, and named in the error so
the fixes can be squashed.
`scripts/jj_prek.sh` gets two changes. It forwards its arguments to
`prek run`, so `jj_push.sh` can ask for a specific range, and it now
changes to the workspace root before running. It exports `GIT_DIR`,
which makes git treat the current directory as the work tree, so prek
could not find its configuration from a subdirectory.
`jj` does not expand aliases when completing arguments, so `jj push`
completed file names. `scripts/completions` has Bash, Zsh, and Fish
completions that give it the same completions as `jj git push`.
`docs/project/contribution_tools.md` documents the `push` alias, and a
`prek` alias for `jj_prek.sh`, with the other per-repository `jj`
configuration. Both are opt-in.
Assisted-by: Claude Code
When evaluating a deferred member access action, the scope stack cannot
be relied on, so `LookupUnqualifiedName` cannot be used in
`GetHighestAllowedAccess` to get the `Self` type.
Instead, store the `Self` type in the `Context` when evaluating a
method, and use that in `GetHighestAllowedAccess`.
Add rules to not overwrite git/jj history without asking, since this
destroys the reviewer's view of things. And some information on dealing
with stacks of commits within a single bookmark/PR.
Prek can make fixes for whatever caused a failure, and then pass when
you run it again, even though the user didn't change anything, and that
is now explained.
Assisted-by: Opus 5
If a Carbon class overrides virtual functions from a C++ base class but
is never referenced from C++, it is never exported to Clang. During
lowering, `BuildVtable` then fails to find a `CXXRecordDecl` and crashes
when attempting to get the vtable from Clang's code generator.
Ensure dynamic classes with foreign vtables are exported to Clang when
completing the class definition in `CheckCompleteClassType`, and look up
`first_decl_id()` in `BuildVtable`.
Fixes#7721
---------
Co-authored-by: Dana Jansens <danakj@orodu.net>
Per feedback on #7665, this PR switches the default value
table storage from canonical constant inst_ids to
non-canonical.
Furthermore, this PR simplifies the default value support in
check by requiring that the first owned declaration of a
function completely specify all of its default values.
Updates the diagnostic code and tests to reflect this new
stricter requirement.
Remove claim from script that it takes minutes; this caused Claude to
decide to not run it. It only takes a few seconds these days. Add note
in toolchain development skill that lints won't be accurate if the
compilation database is outdated.
Assisted-by: Claude Opus 5 via Antigravity
This action was created to wrap any `MetaInstId` operand of an action
instruction. This served two purposes:
1) It had a special hook in `OperandIsDependent` to allow it to be
performed while it had a dependent operand (the reference to the
instruction in the generic).
2) It created a `specific_inst` so that the downstream action saw an
instruction in the specific instead of one in the generic.
These are both replaced: the special case in `OperandIsDependent` for
`RefineInstAction` is replaced by a special case for `MetaInstId`s in
general, and the `SpecificInst` is now created as part of performing the
downstream action, rather than as a separate step carried out
beforehand.
This simplifies the produced SemIR and reduces the number of splices
significantly. It also prepares us to handle actions like
initialization, where we don't actually want to create `SpecificInst`s
immediately in the location where the action is performed, because they
actually belong somewhere else in the IR.
Assisted-by: Claude Opus 5 and Gemini via Antigravity
Replaces several quadratic-time steps in the algorithm with linear or n
log n implementations. Worst-case runtime before hitting the complexity
bailout drops from 25ms to about 12ms, and typical runtime is
single-digit ms on my development machine.
Assisted-by: Claude Code
Instead of always printing types as canonical, attempt to find a sugared
type where possible, and include that type in the diagnostic. We can
only do this when given the instruction whose type is being printed
(`TypeOfInstId`) rather than the canonical type ID.
Initial support here is intentionally minimal: just looking through
calls to the callee's declared return type, and looking through pointer
dereferences and corresponding pointer types, to build out the initial
infrastructure. More cases can be added later; this degrades gracefully
to using the canonical type if a better type can't be found.
Assisted-by: Claude Opus 5 via Antigravity
We use the same conversion codepath to handle both qualification
conversions and derived-to-base conversions, because we allow both to be
performed at once. However, we were previously modeling the
qualification conversion as happening *first*, and producing a result
whose type is the target type of the overall conversion (that is, the
base class type). That led to bogus SemIR, where a `Derived` -> `const
Base` conversion would first have a "compatible" conversion from
`Derived` to `const Base`, *then* an access of the base subobject (of
type `const Base`, within an object of type `const Base`).
We now reverse the order: first we do a derived-to-base conversion,
which already has logic to preserve qualifiers, and then we do any
necessary qualification conversions on the result to reach the overall
target type.
In passing, we now skip forming the `as_compatible` instruction at all
for a pure derived-to-base conversion that has no qualification
conversion, simplifying the SemIR by one instruction in the common case.
Fixes the malformed parse tree produced for an invalid let struct
pattern containing a single identifier (e.g., `let {s};`).
As pointed out by @DavidLoftus, the parser should produce a parse tree
similar to that of `let {ref s};`, since both are missing a binding
power operator `:`, and both do not have a `.` preceding the identifier
(i.e., `state.in_field_shorthand_pattern == true`).
This means that the parser can produce an the `InvalidParse` node just
as it does for `let {ref s};`.
Closes#7674
Four regions had `end` patterns that could fail to match, so `return
var;`, a `fn` with no parameter list, and an unterminated `"` each
swallowed the rest of the file; 77 of 1696 testdata files lost their
highlighting partway through. Operators were wrapped in `\b`, which only
holds next to a word character, so `a + b` highlighted nothing. And the
keywords had drifted about two years behind the lexer.
Identifiers are now classified by naming convention plus a call-site
lookahead, the way the Rust grammar does it, so nothing carries between
lines. Regions survive only for strings and embedded C++, where a
terminator reliably turns up, and raw strings spell out hash levels 0
through 2, so `\n` is an escape in `"..."` and plain text in `#"..."#`.
Trailing comments, character literals, raw identifiers, `$0`, `0o`
octal, arbitrary integer widths and six missing keywords are covered
now, `destructor` is gone, and `i32` reads as a type rather than as a
keyword.
Highlighting stays forgiving rather than diagnostic: anything after `//`
is a comment and odd numeric spellings still read as numbers. Pointing
out mistakes is the toolchain's job, and lenient rules hold steady while
you are still typing.
Every keyword and symbol in `token_kind.def` is covered, and unscoped
tokens across examples and the toolchain drop from 39% to 21%.
Note that I haven't tried to read and reason about every minute change
here as there are just too many. But I'm working on a follow-up PR that
adds testing that should be significantly easier te review.
Assisted-by: Claude Code
We were passing a SpecificInterfaceId which just makes code have to do a
lookup to get the actual SpecificInterface. The caller already has the
SpecificInterface, so plumb that around.
SpecificInterfaceId really only exists when we need to stick a
SpecificInterface into an instruction as an operand.
Fixes#7731, at least under `--share-cpp-ast` which is expected to be
the future direction.
When --share-cpp-ast is enabled and any compilation unit has C++
imports, include all compilation units in the shared CppDomain inputs
and assign the domain to every unit. In ImportCpp, when a unit has no
direct C++ imports but is covered by a shared CppDomain, initialize
its C++ AST context and import namespace. This ensures units without
direct C++ imports have access to the C++ AST and code generator when
instantiating generics or referencing declarations from units that do.
Assisted-by: Antigravity with Gemini