* Split to 4 files: `function_param_int16`, `function_param_int32`,
`function_param_unsupported.carbon`, `function_return`.
* Add // === headlines on group of file shards.
* Rename shard files to shorten them and make them more consistent given
the file context.
* Deduplicate identical .h files and group their tests.
Potential future improvements:
* Split further. For example, pointers and references might be somewhat
separate from int primitives.
* Remove `import_` shard file prefix, as it repeats itself, but leaving
for now as it makes it more explicit.
Part of #5263
I think this would probably have prevented the missed include in #5469
-- it would've just failed completely with a "missing prelude"
diagnostic.
Also note this excludes the included IR from output, because it's
probably low-value to print.
Instead of wrapping `ClassStart` with `StartOnly` or `StartWithEnd`,
provide both `ClassStart` and `ClassStartOnly`. Use inheritance to share
the fields. Initializing the subclass as an aggregate is still possible,
but requires an extra set of curlies.
This flattens the switch in TypeStructureBuilder::Build to a single
level.
Since this puts 19 elements in the Any variant, we need to extend the
CARBON_KIND_SWITCH support to more than 12 elements, so we bump it up to
24.
This was suggested by @jonmeow here:
https://github.com/carbon-language/carbon-lang/pull/5430#discussion_r2078444685
Teach CARBON_KIND_SWITCH to handle mutable lvalues and rvalues, and
CARBON_KIND to forward along rvalues so that it's possible to write
`case CARBON_KIND(const T& t)`, `case CARBON_KIND(T& t)`, and `case
CARBON_KIND(T&& t)`, depending on the type that was passed to
CARBON_KIND_SWITCH.
Replace all uses of VariantMatch with their equivalent of a switch using
CARBON_KIND_SWITCH, and remove the VariantMatch helper from the
codebase.
We were importing all impls as non-final, since we forgot to set the new
field when constructing the imported Impl. Adds a test that fails before
this PR, since the imported Impl is treated as non-final.
We use CARBON_KIND_SWITCH for handling the output of TypeIterator in the
TypeStructureBuilder
Here is how the errors look when it is misused:
- If you don't cover ever type in the variant with a case
```
Enumeration value 'VariantTypeT1NotHandledInSwitch' not handled in
switch
```
Is attached to the CARBON_KIND_SWITCH() usage, the `T1` being a 0-based
index into the std::variant's type list, indicating which type was
missed.
- If you have a case for a type that is not in the variant
```
In template: constraints not satisfied for class template
'ValidCaseType' [with T = char]
... bunch of instantiation stuff ...
kind_switch.h(124, 12): Because 'char' does not satisfy
'TypeFoundInVariant'
```
Where `char` was the type I put in the `CARBON_KIND` macro, which was
not in the variant.
- If you have too many types in your variant (currently > 12)
```
In template: static assertion failed due to requirement 'sizeof...(Ts)
<= 12': CARBON_KIND_SWITCH supports std::variant with up to 12 types.
Add more if needed.
```
Is attached to the CARBON_KIND_SWITCH() usage.
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
If there will be no in-scope instructions printed, have
`FormatScopeIfUsed` skip the relevant scope.
Note, this only affects constants and imports. It's not changing the
file scope, which is usually printed when empty.
The version of clangd/clang-tidy on developer machines has slowly
diverged from the one on the CI builders, which is causing a slowly
increasing amount of pain as clang-tidy CI runs fail (incorrectly) over
things that a newer clangd/clang-tidy was perfectly fine with locally.
This bumps the Clang version used in the ubuntu builders to 19, which is
the most recent in Debian stable.
We use https://apt.llvm.org instead of LLVM's GitHub releases
(https://github.com/llvm/llvm-project/releases) as the former more
reliably has packages for newer Clang/LLVM versions on x64. The
community-build releases binaries on LLVM's GitHub have stopped
including Ubuntu packages that match the GitHub x64 Ubuntu workers for
some time (for at least the 18 and 19 releases).
By moving to apt.llvm.org packages we only download and install the
headers and libraries needed for development, rather than every output
of building llvm, which is much faster and saves lots of disk space. We
also remove the system installations of other versions of clang/llvm so
we should end up using negative disk space. We can no longer easily
cache the installation but apt.llvm.org is a reliable end point.
We bump the ubuntu image version for the github workers to 24.04, as
apt.llvm.org has stopped building images for 22.10 in 2022 at its end of
life.
The `pre_commit` workflow disabled sudo unlike the other workflows that
install Clang/LLVM, including the `clang-tidy` workflow (which is also
run on `pull_request`). We bring it into alignment with the other
workflows so that we can install the llvm packages. And we lock its
ubuntu image to 24.04 so that it can be moved in lockstep with the other
workflows that depend on Clang/LLVM.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Right now, a lot of tests have started setting `--no-dump-sem-ir`. My
thought is that we can look at:
1. Put ranges in a bunch more files.
2. Shift more towards `--dump-sem-ir-ranges=only` instead of
`--no-dump-sem-ir`, because it allows mixing fail-tests with no IR
alongside tests that contain IR.
3. Evaluate switching the default to `--dump-sem-ir-ranges=only`, and
instead set `--dump-sem-ir-ranges=if-present` only in files that want to
typically show all IR (particularly import-related tests, where ranges
don't work well).
In real-world use, my thought is also that it'd be helpful to be able to
add the dump range comments to files, see the output (i.e., the default
behavior of `if-present`) but then also be able to pass `ignore` in
order to see the full IR without modifying the file (possibly also
useful in tests). That model is why I went for tri-state handling.
Note `only` can also have an interesting side-effect. Because core files
(including min_prelude versions) typically won't have ranges, they'd be
implicitly excluded.
Such impls will never be used, so they should not exist. And test that a
final impl partially overlapping a non-final impl is accepted.
There is a question about a final impl partially overlapping a final
impl that is part of
https://github.com/carbon-language/carbon-lang/pull/5337
Deduction can import stuff which can invalidate all of the value stores.
Refactor out the code that diagnoses unused generic bindings, and scope
the reference into the ImplStore so it can't be used after. Fetch the
impl from the store again when setting the witness to error afterward if
needed.
No functionality change right now: we reject thunks where the signature
has no return type and the callee has a return type. But discarding the
expression is still the right thing to do.
With this enabled, entities that live in value stores are poisoned
whenever any action is taken that might invalidate pointers and
references to those options -- in particular, adding another item to
that value store, or attempting to load any entity from an import IR.
Subsequent uses of those pointers or references then trigger an ASan
failure.
This detects latent bugs where the pointer or reference to the entity
would become stale if we got unlucky about when the value store
reallocates, even in cases where the reallocation didn't actually
happen.
This is not enabled by default: it finds a lot of latent bugs, so our
tests don't pass with this option. This PR also includes fixes for a few
of those bugs.
---------
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com>
If a function contains instructions whose locations are in another file,
skip providing debug locations for those instructions rather than
CHECK-failing.
This happens when emitting a thunk where the signature is declared in
one file and the call target is in another file: some parts of the thunk
use the original signature as their locations, whereas other parts of it
use the location of the call target.
Non-type values were being represented as "Concrete" in the type
structure's shape, which was incorrect when the value was symbolic. Now
they are represented as "Concrete" or "Symbolic" as appropriate.
When concrete, a matching concrete ConstantId is stored in the type
structure's concrete values, so that they will be used for equality
comparison. And then impl matching comparison is taught to look in the
concrete values for mismatches when it finds a "Concrete" or
"ConcreteOpenParen" shape on both sides of the comparison, and return
false if they are not equal. This reduces the number of candidate impls
selected in impl lookup, which avoids doing type deduction against impls
that won't match anyway due to the concrete values in the query not
matching concrete values in the impl. As a result we see fewer specifics
being generated in the semir.
The TypeStructureBuilder recursively iterates through a type, interface,
or facet value and constructs a type structure that includes each
concrete and symbolic value found. This iteration is a more general
thing that can be useful elsewhere. For instance, to get just the root
type out of a general type, it is the result of the first iteration
step.
We abstract out the iteration logic into a SemIR::TypeIterator to create
a clear boundary between the work of iterating and the work of building
the TypeStructure from it.
Along the way this pointed out some issues in the TypeStructureBuilder
where it could have ambiguity between types that include non-type
values. So we add some tests for these cases, and they now pass. There
are also TODOs left behind, as concrete TypeStructure is overly specific
right now in order to keep these tests passing, which means that the
concrete elements can't be used for impl lookup matching yet. Only the
shape of concrete vs symbolic is used for now, and then type deduction
is used to compare the actual concrete types, which could be skipped
when the concrete values could be compared directly and reject an impl
for not matching.
#5445 updates to bazel 8.2.1, this does more updates (including to
buildifier, which does autofixes like the `sh_test` loads in the other
PR).
Note I'm using the latest available clang-format wheel. That's not
really something I expect people to have installed, but should mostly be
consistent. I'm specifically skipping clang-format 18 because it had
some broad regressions, and 19 got really confused by a `requires` on a
trailing return. Using the latest seemed probably okay since most people
won't see the difference. Do note that trailing returns in macros,
https://github.com/llvm/llvm-project/issues/47664, seems to be cropping
up again as an issue.
- Updates incompatible flags.
- `rules_flex` is no longer used, so enable its flag.
- Fixes `sh_test` deps for
`--incompatible_disable_autoloads_in_main_repo`
- Broadens the exception for `rules_cc` and `bazel_tools` due to changes
to runfiles deps; trying to avoid minutiae that shouldn't affect the
decision.
The current behavior hits UBSAN and ASAN issues.
Note, `RequiresNullTerminator` is already set to `false` in
`source_buffer.cpp`; setting it in `compile_helper.cpp` is making things
more consistent. The related logic is an [assert
fail](https://github.com/llvm/llvm-project/blob/main/llvm/lib/Support/MemoryBuffer.cpp#L52).
This was fuzzer-discovered.
Instead of crashing when run outside of `bazel`, make the toolchain's
`file_test` binary work properly when no test-specific environment
variables are set. This makes it a lot easier to run `file_test` under a
debugger.
There are two main changes here:
- Don't crash if `$TEST_TMPDIR` is unset. Instead, fall back to LLVM's
temporary directory (typically `$TMPDIR`). We already did this in some
places in tests. We now do it in more places.
- Don't fall back to a target label of `<target>` in the reproduction
commands if `$TEST_TARGET` is unset, because this causes all the tests
to fail because their output doesn't match the expected output due to a
differing bazel run command. Instead explicitly specify the target from
the `FileTestBase`-derived class.
Infrastructure for this has been added generally, but only rolled out to
the toolchain `file_test` binary for now.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Instead of building the definition of a thunk immediately when we
generate the thunk declaration, wait until we reach the `}` of the
outermost class, interface, etc. -- at the same time when we would parse
the definition of the thunk if it were defined inline.
This fixes issues where we fail to define the thunk because it requires
an enclosing class to be complete, or its definition depends on
something declared later in the enclosing class.
Make the representation of a suspended function scope, and its
constituent suspended components, be move-only, and switch to passing it
around by rvalue reference instead of by value because it's expensive
both to move and especially to copy.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
This is fixing a deduce crash, with regression tests added in
binding_pattern.carbon. In `needs_substitution`, it adds the
`BindSymbolicName` with the compile time bind index corresponding to the
wrong generic scope, which causes a bad result. It appears
`needs_substitution` logic is no longer needed (per zygoloid,
`CheckDeductionIsComplete` handles related issues) so can be removed.
This changes the order of IR in use_assoc_const.carbon but the result
appears equivalent to me.
This was a fuzzer-found crash.
```
class C {
impl as I;
}
```
is redeclared
```
impl C.(as I)
```
for purposes of `match_first`/`impl_priority` blocks and definitions.
---------
Co-authored-by: Josh L <josh11b@users.noreply.github.com>
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
`IRBuilderBase::SetInsertPoint` weirdly replaces our debug location with
one copied from the new insertion point, so undo its damage after
calling it.
Also included: a couple of cleanups I made while tracking this down.
Ran into this trying to debug #5428
Before:
```
lex.cpp:793:21: runtime error: null pointer passed as argument 1, which is declared to never be null
string.h:90:51: note: nonnull attribute specified here
SUMMARY: UndefinedBehaviorSanitizer: undefined-behavior lex.cpp:793:21
```
After:
```
lex.cpp:793:21: runtime error: null pointer passed as argument 1, which is declared to never be null
string.h:90:51: note: nonnull attribute specified here
#0 0x5572db22a680 in Carbon::Lex::Lexer::MakeLines(llvm::StringRef) /proc/self/cwd/toolchain/lex/lex.cpp:793:14
#1 0x5572db227ee4 in Lex /proc/self/cwd/toolchain/lex/lex.cpp:738:3
(etc)
SUMMARY: UndefinedBehaviorSanitizer: undefined-behavior lex.cpp:793:21
```
Interestingly, even though this is labelled as UB, it's using the
ASAN_SYMBOLIZER_PATH (specifically not LLVM_SYMBOLIZER_PATH). But
canonically UBSAN_SYMBOLIZER_PATH may also be used per
https://github.com/llvm/llvm-project/blob/main/compiler-rt/lib/ubsan/ubsan_flags.cpp#L53,
so I'm adding it to the list out of an excess of caution.
Also doing some small related cleanup:
- Removing `ASAN_SYMBOLIZER_PATH` from `--test_env` because it should
now be getting overridden by these settings (also was a little
inconsistent in that `LLVM_SYMBOLIZER_PATH` was not included).
- Improving the environment construction and documentation.
The canonical location of the instruction may be an entirely different
instruction, which the instruction in question was not imported from. In
particular, we shouldn't assume that we can use the constant value of an
instruction that the *location* of an imported instruction refers to as
the constant value of the imported instruction.
The only time we should be looking at the `ImportIRInstId` for a `LocId`
is when determining its location in some other file.
Fixes a crash when importing thunks (which can contain instructions
whose location points to an instruction in a differnt IR).
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
When two non-final impls have the same type structure (neither is a
specialization of the other), it is an error unless they are within a
`match_first` block. For now, we don't have `match_first` implemented,
so it's always an error.
---------
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
A final impl must be written in the same file as the root self type or
the interface. This provides tests that should fail but don't yet for
writing a final impl in a third file that defines neither, as well as a
blanket final impl over an interface outside the file that defines the
interface.
Fixes a compile failure with the new version, essentially:
```
external/+llvm_project+llvm-project/llvm/include/llvm/Support/FormatVariadicDetails.h:157:1: error: implicit instantiation of undefined template 'llvm::support::detail::missing_format_adapter<clang::LookupResultKind>'
```
Right now we construct `tree_and_subtrees_getters` a couple different
ways, it's just not obvious because one's abstracted in `check`. But
also, when formatting IR, we'll repeatedly do the `IncludeInDumps`
string check, which felt odd to me since it only needs to be calculated
once per IR.
This also shifts `CheckIRId` selection a little earlier, and in doing so
makes `CheckParseTrees` accept a sparse `units` argument. I actually
think this is a positive: it makes `CheckIRId` a little more stable
across possible command lines, when file loading fails (which is the
only time that a file will have a `CompilationUnit` but not a
`Check::Unit`).
Trying to build on the shared issue between these, I'm adding a
`MultiUnitCache` to store the calculated arrays. For the subtree
getters, this is very minor and avoids at most one incremental array
construction (moving logic out of `CompileSubcommand::Run` might be the
bigger benefit). For `include_in_dumps`, when dumping SemIR, this is
changing a calculation run once per entity (in each IR) to be calculated
once per IR (globally), i.e. O(M*N) -> O(N).
Note this seems to be marginal for performance of file_test:
- Before: Stats over 10 runs: max = 5.3s, min = 4.7s, avg = 4.9s, dev =
0.2s
- After: Stats over 10 runs: max = 4.9s, min = 4.7s, avg = 4.8s, dev =
0.1s
I was mainly thinking about this in the context of dumping SemIR ranges.
There, the impact may actually decrease because a range won't do any
cross-IR printing. But, I'm expecting to add another layer for whether
we're printing IR for a file, and that made the `should_format_entity`
callback stick out for me.
When we fill the witness table with errors, set the witness id to an
error too, which signals to impl lookups to not use the impl.
Make the use of the `Impl` from the store more consistent once it's been
added to the store (or known to be there already).