Commit Graph
2371 Commits
Author SHA1 Message Date
Jon Ross-Perkins ee467900ed Use llvm:: for specializations instead of 'namespace llvm' (#3682)
We already do this with things like llvm::DenseMapInfo, I don't know why
I was doing this with format_provider. But this should be more
consistent, and slightly better for not entering another library's
namespace.
2024-02-01 22:58:04 +00:00
josh11bandJon Ross-Perkins 03bf22e55e Parse tree for impl that is better for check stage (#3678)
Use two different nodes for "<type> followed by `as`" and "<type>
omitted before `as`, use `self`", so it is easier to determine which
case. Later the second case will push the type id for `self` onto the
node stack, making the two paths more similar.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-02-01 22:23:21 +00:00
Chandler Carruth efde1497c4 Remove the link to a GSoC organization (for now). (#3680)
We had some amazing GSoC participants last year, but because Carbon is
still pretty small, we ended up stetched a bit too much to be
sustainable. And this year, we're trying to have an even narrower focus
on the toolchain.

Between these aspects, we sadly don't have the bandwidth to run an
effective GSoC project this year. We think it's really important that we
can give mentees an excellent experience on the project, and don't want
to overcommit ourselves, or worse, let them down.

We're really hopeful to be back when the project is a bit larger though,
and we have good bandwidth to host folks.
2024-02-01 22:08:20 +00:00
Chandler Carruth a57abdfd45 Update Abseil and RE2 to latest releases. (#3676) 2024-02-01 17:51:54 +00:00
Jon Ross-Perkins 134766d50c Refactor File constructors to share logic. (#3675)
I was considering this again while working on #3674, it would be helpful
to share initialization for both regular and builtin File constructors.
2024-02-01 03:32:57 +00:00
Richard Smith a1655b6858 Tuples and tuple indexing (#3646)
Add support for extracting elements of a tuple by their numerical index.

Also formally add the well-established basic syntactic and semantic
rules for
tuples, for which we have had leads issues but no proposal, into the
design.
2024-01-31 23:42:11 +00:00
Richard Smith 9e7a17b1a1 Scaffolding for checking impls. (#3672)
Consume the components of the `impl` declaration, and set up scopes for
the child elements. We don't yet build a representation for the impl
itself.

Also, add an interface type value. This is necessary so that we have a
value for the expression on the right-hand side of `as` in an `impl`.
2024-01-31 22:18:29 +00:00
Richard Smith 44fca1669a Keep parameters in scope throughout the entity that they parameterize. (#3671)
Previously, we created scopes for implicit parameter lists and tuple
patterns, but that meant that bindings went out of scope too soon. We
now keep them in scope until the end of the enclosing declaration. This
is accomplished by pushing a scope for parameters when we handle a name
that might have them, and then popping the scope again if it turns out
that there were no parameters.

For a case such as:

```carbon
fn A(T:! type).B(U:! type).F(x: T, y: U) {
  var z: T;
}
```

... we now have the following scopes in the stack:

-   A parameter scope containing `T`.
-   A class scope for `A(T:! type)`.
-   A parameter scope containing `U`.
-   A class scope for `A(T:! type).B(U:! type)`.
-   A parameter scope containing `x: T` and `y: U`.
-   A function body scope containing `z: T`.

The innermost scope when check processes a declaration of a function,
class, or similar is now often a parameter scope rather than the
enclosing scope in which the class or function is declared, so the
target scope is now passed explicitly into the modifier checking code
that wants to inspect that enclosing scope.
2024-01-31 21:40:45 +00:00
Chandler CarruthandJon Ross-Perkins 5e88fe72c9 Switch remaining relational operator overloads to spaceship. (#3668)
This doesn't matter too much as these were expanded by the LLVM iterator
facade, but eventually this should enable that facade to be a little
less fancy and should also simplify the dispatch to directly use the
three-way comparison.

One case is a bit subtle and didn't have a comment so I added one to
explain a bit what is going on there.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-01-31 03:29:02 +00:00
Chandler CarruthandJon Ross-Perkins 1529538ad2 Use concepts and better comparisons from C++20 in IndexBase. (#3666)
This is one of the nicer improvements to boiler-plate as we benefit from
concepts simplifying the constraints on the templates, defaulted `!=`
finding `==`, and the `<=>` operator for a single definition of all
relational operators.

I've checked and the `!=` usages do seem to be effective at finding the
`==` definitions, but we still need both the RHS- and LHS-based
overloads of `==` as far as I can tell.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-01-31 02:47:34 +00:00
Chandler Carruth 2e236759ca Switch //common to use C++20 concepts. (#3665)
This removes the use of `enable_if` and tries to adopt concepts instead
of type traits when available.

The `ostream.h` change is a bit subtle as it adds a restriction not
previously in place -- that the stream is *contvertible* to
`std::ostream` as well as having it as a base class. This seems to match
the intent of the code.

The `hashing.h` code adds an implementation detail concept, and so I've
also clarified that the dispatch namespace is an internal one that isn't
part of the public API.
2024-01-30 18:14:07 +00:00
Chandler Carruth 79ac4c953a Switch migrate_cpp to use concepts. (#3667)
This converts what was essentially a concept to actually be one and uses
it.
2024-01-30 17:40:19 +00:00
Jon Ross-Perkins e583493e9f Implement some basic ImportRef handling for builtins. (#3663)
I'm basically just nudging down the path I think is right here. Adding a
small bit more support, but also more tests to capture cases that I
think will need to be verified as working.

With diagnostics like "Value of type `<function>` is not callable.",
that's because it expects a FunctionDecl but is instead finding a
ImportRefUsed. I'll need to work out the necessary support for a
callable function.
2024-01-30 17:38:38 +00:00
Richard SmithandChandler Carruth 3fa70de101 Remove some C++17 workarounds now we build in C++20 mode. (#3653)
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2024-01-30 02:19:38 +00:00
Jon Ross-Perkins 7f11012f58 CrossRefIRId -> ImportIRId (#3662)
One more (hopefully last) rename on the Import instruction renaming.

I was kind of tempted to rename to just "IRId", since the IRs aren't all
imports. However, this felt easier to read, and a better choice than
CrossRef because it's more consistent with the other ways imports exist
in code. (even if IRs aren't all imports, most use-cases are derived
from imports)

Note though that import_irs may include IRs not just from direct
imports. Beyond the builtin IR, I'm thinking that for indirect imports,
or the prelude, we may end up adding them. e.g., so that constants can
be generated for indirect imports and still correspond to a directly
known IR, and for a given IR that's indirectly imported multiple times
to be deduplicated locally. I'm not there yet, I'm just mentioning this
to help give background for naming thoughts.
2024-01-29 23:05:47 +00:00
Chandler Carruth bf02d1f4b0 Remove headers marked as unused by ClangD. (#3661)
This required adding a few headers that were found transitively before,
but not too many. This is sadly a fairly manual process of opening every
file in my IDE, but I think I got everything in `//common` and
`//toolchain`.

There are a few cases where technically we don't need `foo.h` to be
included into `foo.cpp`, but I've forced those to stay with a pragma.

I've tried to catch the places where we can cut deps in Bazel as well,
but not sure I got all of those.

I had been noticing these in other PRs and it seemed better to isolate
the change.
2024-01-29 16:15:35 +00:00
Chandler CarruthandRichard Smith 8bee5ebe83 Enable C++20 and fix infrastructure to work with it. (#3660)
C++20 mode exposes bugs in `clang-tidy` 16 that are challenging to work
around, but this should manage to do so. The patch to `bazel_clang_tidy`
has been sent upstream but no need to wait for it.

We don't actually specify a specific version of C++ in our style guide
docs, and the Google style guide has already been updated to be based on
C++20 so nothing to be done there.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-01-27 02:19:08 +00:00
josh11bandRichard Smith afd7115c0e Support determining IdKind from NodeCategory, in addition to NodeKind (#3648)
The categories `Expr`, `MemberName`, `Decl`, `Statement`, and `Modifier`
are usable since they are associated with a consistent `IdKind`. The
mapping to `IdKind` for NodeKinds that have those categories are no
longer listed explicitly, ensuring that the `NodeCategory` mapping is
the source of truth.

Also: fixes the category of the `FunctionDefinitionStart` and
`ArrayExprStart` node kinds.

Note: I've added [a section on defining constexpr constants to the
Toolchain architecture
doc](https://docs.google.com/document/d/1RRYMm42osyqhI2LyjrjockYCutQ5dOf8Abu50kTrkX0/edit?resourcekey=0-kHyqOESbOHmzZphUbtLrTw&tab=t.0#heading=h.f7682a2tpvxr).

FUTURE:

* We should switch `TuplePattern` to put an `InstId` on the `NodeStack`
instead of an `InstBlockId`, so we can handle the pattern category.
* We should make a category for names to replace uses of the `NameId`
`IdKind`.
* We should make use of these new APIs more, and propagate more-precise
types through the codebase.

QUESTION: Should I use a different approach for determining the number
of members of the `NodeKind` enum?

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-01-27 01:04:01 +00:00
Jon Ross-Perkins 8167c44a03 Merging CrossRef into ImportRefUsed, shifting builtins over. (#3659)
This is a bit of a cleanup; I probably should've just renamed CrossRef
instead of adding ImportRefUsed.

Adding `is_builtin` to InstId is more about providing a standard API for
the check, which I expect to add a little more of.

Shifts import tests to validate that the BuildValueRepr CHECK isn't
accidentally hit.
2024-01-27 00:36:58 +00:00
Chandler CarruthandJon Ross-Perkins 13de9e9d06 Fix outstanding clang-tidy errors. (#3654)
Recent runs of `clang-tidy` for me started showing more errors, and this
is a collection of changes to address them.

First, I've systematically applied the disabling tag to all C++ rules
under //explorer/... with `buildozer` so we don't spend time analyzing
this code or reporting errors from it. Not sure this was strictly
necessary, but it seemed like a nice consistency improvement.

Next, I disabled a buggy check for missing `default` cases in
`switch`es. It seems to get confused by the fancy conversions in our
`enum_base.h`. We don't miss much with this as the Clang compiler
warnings for `switch` catch most of our actual bugs. I also removed the
local disabling of this now that it is turned off centrally.

Lastly, I added error checking to two file descriptor manipulating calls
in the `file_test` infrastructure.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-01-26 18:20:47 +00:00
Jon Ross-Perkins c1a894fca7 Start adding ImportRefUsed, removing older copy behavior. (#3657)
Builds on #3656.

Under the prior "lazy" model, we had been planning to copy instructions.
Under the current "unused+used" model, we're restricting that to
constants. I'm trimming back some of the ResolveIfImportRefUnused logic
because it was more appropriate for the former model.

Right now, I'm adding ImportRefUsed direct creation in import.cpp for
namespaces. I think I'll need to do something similar in
RseolveIfImportRefUnused... I'm still trying to think about how to
manage type information there (which needs to come in as an import
reference itself, and probably have some amount of deduplication before
forming a TypeId). So ImportRefUnused lacks a type because I'm hesitant
to aggressively load it, whereas ImportRefUsed should *always* have a
type but it's just an error while I think things through.

For reference, ImportRefUnused and ImportRefUsed are mainly split in
order to track the boolean "used" without making fundamental
modifications to Inst for bit packing (this effectively instead packs a
bit into InstKind). AnyImportRef currently excludes the type because
it's mainly for diagnostic printing at the moment. It could end up with
a TypeId that would end up Invalid for ImportRefUnused, though it could
also be that the TypeId is only accessed when using ImportRefUsed
explicitly.
2024-01-26 17:26:15 +00:00
Jon Ross-Perkins adad286b74 Refactor LazyImportRef into ImportRefUnused. (#3656)
This makes some changes to the formatter so that ImportRefUnused and
ImportRefUsed will both be labeled as "import_ref" with an "unused ->
used" argument change in textual IR, but is otherwise not changing
logic.
2024-01-26 03:10:00 +00:00
Richard Smith 7b933a1126 Fix crash attempting to convert a struct to an invalid class. (#3658)
Found by fuzzer.
2024-01-26 03:06:02 +00:00
Jon Ross-Perkins 0a610e00c9 Fix missing abi requirement (#3655)
#3649 added a requirement for libc++abi-dev but missed updating the
README to mention it.
2024-01-25 18:37:40 +00:00
Jon Ross-Perkins f4a741903f Add import support for remaining decl types. (#3651)
I'd excluded these initially just because I was thinking towards copies,
but under the current model I'm trying to catch all the decl types just
for consistency. Note references will still be a TODO error
(LazyImportRef is already tested for this, it just didn't feel necessary
to add individual tests while I try to sort out behavior).

Fixes an oversight where declarations in an entity's scope were being
added to the list of exports.

Note I'm trimming some Import API arguments as now-unused.
2024-01-25 18:19:22 +00:00
Jon Ross-Perkins 1f764c8cf1 Implement merging of namespace declarations. (#3647)
Per discussion with zygoloid, namespace declarations will merge both
when repeated in a given file, and across files. This echoes how forward
declarations of other entities are allowed to repeat. When merging with
an imported namespace, fill in the parse node so that future diagnostics
point at the declaration in the same file rather than a declaration in a
different file, just for locality.

Although it might be desirable to issue a diagnostic when a namespace
declaration is repeated within a given file, that's similarly true in
other cases, but may be more desirable as a tidy-style issue rather than
preventing code from compiling. Allowing the repetition also makes the
import versus non-import cases more consistent: the namespace
declarations merge regardless of the source.
2024-01-25 07:29:40 +00:00
Jon Ross-Perkins d25bae09d1 Clean up some semir uses that can use context accessors. (#3650) 2024-01-25 01:23:36 +00:00
Chandler Carruthandjosh11b da7533ae7f Update project to require Clang 16 or newer. (#3649)
This should also unblock our switch to C++20 and other improvements.

There are three core parts of the change --

1) Updating our infrastructure to fetch and find Clang-16.
2) Updating our documentation to reflect this and help folks with any
   system issues they encounter.

The infrastructure change is unfortunately tricky. We can't get Clang 16
easily on GitHub's runner images, and in the past we've had persistent
problems with flakiness when our actions download this much during their
runs. Due to the flakiness, we've previously removed all downloading of
dependencies outside of Bazel itself, and added retry loops around Bazel
specifically to overcome flaky downloads.

This change tries to address these problems by populating the Clang and
LLVM toolchain in a place that we can then cache using the built-in
GitHub action caching infrastructure. This seems like by far the least
likely to flake way of downloading extra things into our runs. And since
these are relatively slow moving dependencies, we should populate this
cache very, very rarely.

For Linux, this downloads the binary release artifact from GitHub,
prunes out large parts of it that we don't need, and then caches this as
a local toolchain. This proves both small and fast.

For macOS, this uses a trick to cache the destination of Homebrew
installs. It unfortunately caches the *entire* Homebrew installation
though, and so it also goes to some lengths to prune and minimize how
much is installed from Homebrew. The result is "only" a 2gb cache image.
Because of the size and slower download and filesystem, the macOS runs
see a 1 - 2 minute slowdown.

We might extend the Linux infrastructure here usefully if we want to
test multiple LLVM versions. We might also extend the macOS version to
get a cheaper way to prune parts of the system and free up disk space,
or to cache other Homebrew installed tools if needed.

Last but not least, this brought to the forefront an issue with our C++
toolchain integration which relied on a specific CMake build option
being set in the LLVM toolchain install. This option isn't used in the
official release artifacts. Instead, switch to a more robust approach to
linking libc++abi statically that shouldn't have these problems.

Beyond the infrastructure changes, this also updates the documentation
to reflect requiring Clang 16 or newer, and adds some extra tips for
folks that are missing this.

The documentation is also updated to address a problem with getting the
right libc++abi files installed to support the more robust linking
strategy. This may reduce the problems we've seen in the past around
libc++abi and linking on other Linux distros as well.

---------

Co-authored-by: josh11b <josh11b@users.noreply.github.com>
2024-01-25 01:23:09 +00:00
Richard Smith 439a644960 Propagate the phase of a type from its constituent types. (#3645)
To avoid bouncing through `constant_values()` to determine whether a
type is symbolic or template, store the `ConstantId` on the `TypeInfo`
not just the `InstId`.

In addition to propagating the symbolic / template phase, this also
propagates whether a type contains an error, resulting in our no longer
producing types such as `<error>*` -- these now evaluate to simply
`<error>`. While this makes our types less precise after an error, it
also removes some follow-on diagnostics, so it seems to be an
improvement on the whole.
2024-01-25 01:00:37 +00:00
Karthik Prakash af6436c20c Add filetype highlighting to block string literals (#3642)
This PR adds filetype highlighting support to block string literals in
the TextMate grammar.

After adding the regex in the `begin` of the block string literal,
`beginCaptures` is used to access the first capturing group of the regex
(which captures the filetype) and apply the `constant.character.escape`
rule.


![image](https://github.com/carbon-language/carbon-lang/assets/116057817/edabc820-ed0f-4f5f-a180-f3ac90045b0a)

Closes #3641
2024-01-24 18:56:25 +00:00
Jon Ross-Perkins 91f0c23124 Provide diagnostic locations for imported namespaces. (#3640)
Building on #3636 which handles the general import case, add special
casing for namespaces. Namespaces can be combined cross-IR, so it's a
little more complex.

The implementation adds import_id to the Namespace instruction as a
reference to find the original using the normal structure. This is
achieved by moving the name_id to NameScope to free up space.

---

I considered a few alternatives...

I considered adding import_id to NameScope, but:

1. It's more consistent with things such as Function or Class that
provide name_id on the info object rather than the instruction.
2. I thought it more likely that there would be more NameScope cases
that might want a name_id rather than the import_id, since ImportRef
will typically be used.

I considered putting the import source (cross-ref IR id + inst id) on
the NameScope versus a separate ImportRef, which seems like the
strongest argument towards the NameScope approach because it removes an
instruction. That just felt inconsistent though, and the overhead of
instruction-per-imported-namespace should be low (theoretically few
namespaces should be used). Plus I feel a bit odd adding two
generally-unused ids to NameScope.

A specialized Namespace structure could also have been created to store
the import_id, but that would add an indirection to the NameScope.
2024-01-24 18:00:32 +00:00
czapiga ef842736f0 Fix block string literals highlighting (#3635)
Replace block string literals quotes.

Highlighting after change:

![image](https://github.com/carbon-language/carbon-lang/assets/24532774/4c55c021-b8ce-4933-9742-81a8eee6ffe1)


Closes #3634
2024-01-24 09:44:18 +00:00
Richard Smith a1f1c7438f Improve source locations for some diagnostics. (#3644) 2024-01-24 03:37:00 +00:00
Jon Ross-Perkins 2a44cedb8f Provide diagnostic locations for imports. (#3636)
Note namespaces aren't handled well, because they can be generated from
merging imports. That's on my internal TODO list.
2024-01-23 22:37:32 +00:00
Richard Smith 099bd5ce26 Produce a more descriptive error if an array type's bound is too large. (#3638) 2024-01-23 21:01:34 +00:00
kshokhin 65e95942de Choice parsing in toolchain (#3574)
Implement Choice parsing in toolchain according to
[design](https://github.com/carbon-language/carbon-lang/blob/trunk/docs/design/sum_types.md).
2024-01-23 18:22:27 +00:00
Richard Smith afd194de9d Runtime : name bindings are not constants. (#3639)
Do not create runtime name bindings for `FieldDecl`s even though they're
declared with `:`, so that we can still constant-evaluate references to
fields.
2024-01-23 17:02:08 +00:00
Richard Smith e2f21c0052 Check for a constant rather than a literal in indexing. (#3637)
In array indexing, move the check for an out-of-bounds index into the
constant evaluation logic, so that we will also benefit from it when
constant evaluating a compile-time function.

In tuple indexing, require a template constant index instead of an
integer literal.
2024-01-23 16:56:30 +00:00
Jon Ross-Perkins 3822c1b526 Split github_tools into its own bazel repo. (#3632)
The pip dependencies in github_tools are the reason the
MODULE.bazel.lock is platform-dependent. Following complaints about the
platform-dependence, split apart github_tools from the rest of the bazel
repo and make it not track the lockfile: while there's an incremental
safety risk due to not tracking checksums, it's unlikely the tools there
would ever be part of a Carbon release process. If we eventually add
Python tools that need pip to the release, it might be desirable to go
back to re-unify the bazezl repos.

This does make running pr_comments incrementally more inconvenient
because a "bazel run" needs to be run from the github_tools subdir.

As a consequence of separating the dependency, this means tests will not
be continuously run in github_tools. They're now a separate repo, and we
cannot add a dependency without restoring the platform-dependent issue.
I think pr_comments is sufficiently low value and unchanging that it is
not worth building separate CI for it.

Cleans up some legacy references to third_party/llvm-project, since now
github_tools needs to be added to the main bazelignore.
2024-01-22 21:34:02 +00:00
Richard Smith de2316bd0c Fix crash when and or or appear outside a function. (#3633)
Found by fuzzer.
2024-01-22 19:32:59 +00:00
czapiga 507e47d732 Fix syntax highlighting in VSCode (#3630). (#3631)
Add missing escape backslash to `string_escapes` matching string.

Closes #3630

VSCode highlighting working after change:

![image](https://github.com/carbon-language/carbon-lang/assets/24532774/92be2aa3-5f12-4430-9b9e-169edfa1da4a)
2024-01-20 23:16:30 +00:00
Chandler Carruth ebbdc11877 Switch to a "better" multiplicative hash constant. (#3629)
Testing this hash function with representative hash table
implementations showed significant differences in quality between
different multiplicative hashing constants. The constants used and
documented were OK but had clear limitations when merely using a single
64-bit multiplication. However, a search uncovered (partly by luck)
a constant that has empirically been shown to be both significantly
better than other constants and generally not have problematic
weaknesses. There are still plenty of collisions for string keys and
heavily loaded hash tables of course, but no examples of severe outlier
collision rates as observed with all other constants we have tried.

There is a slightly longer comment explaining some of this context and
the other constants we have tried in the PR as well.

To this day, we still don't fully understand why the constant used here
behaves so much better than other constants we have tried, including all
of those we've found in other hashing algorithms.
2024-01-20 22:14:19 +00:00
Chandler Carruth 5f62cf752d Fix an oversight that dropped the buffer. (#3628)
This lost any seed or prior hashing done, which isn't good. I've added
some basic testing that would have caught this immediately.
2024-01-20 21:09:17 +00:00
Chandler Carruth 7e9760d9e4 Simplify the index & tag extraction API for hash codes. (#3627)
The fancier API ended up not being helpful and making it harder to
optimize a hash table implemented on top of this.
2024-01-20 20:55:15 +00:00
Chandler Carruth 37b725a927 Update lock file for another platform. (#3626) 2024-01-20 05:15:56 +00:00
Richard Smith aba3c47467 Constant evaluation for aggregate access. (#3625)
Also, form `addr_of error` instead of `addr_of operand` when `operand`
is not a reference expression, so that we don't try to constant-evaluate
a meaningless expression.

Also, expect a constant for an array bound rather than specifically an
integer literal. This allows constant evaluation results to be more
easily tested by inspecting array bounds.
2024-01-20 01:44:13 +00:00
Richard Smith 87ecb34f6b Constant evaluation support for initializing expressions. (#3624)
The constant value we associate with an initializing representation is
the object representation that the initializing expression will store to
its destination.

Also include the type in the profile of an instruction. This is now
necessary for array values, which are represented as tuple_value
instructions with array type, to avoid instructions with different types
being merged by constant canonicalization.
2024-01-19 23:28:53 +00:00
Rayyan Kandjonmeow a1d03b1f56 Fixed var keyword in Carbon code example. Incorrect keyword color. (#3600)
This PR is changing the color of the var keyword in the `PrintTotalArea`
source example which is referenced in the README. The original color was
white and changed to `#FF7B72`.

---------

Co-authored-by: jonmeow <jperkins@google.com>
2024-01-19 22:30:23 +00:00
Jon Ross-Perkins a80d13ec6c Update bazel and lock to 7.0.1 (#3621) 2024-01-19 20:49:26 +00:00
Jon Ross-Perkins 9be99cad4c Don't take a NodeId argument for insts that have no parse node. (#3623)
This is low impact because we typically support a parse node, but I was
hoping to reduce ambiguity in the cases that don't.
2024-01-19 19:59:19 +00:00