Commit Graph
3255 Commits
Author SHA1 Message Date
Jon Ross-PerkinsandChandler Carruth 08f24551ec Add bit packing to NodeImpl (#4651)
Just a small packing optimization. We currently have 222 `NodeKinds`, so
this reduces us to just 30ish more we can add without needing to pack
more. However, if we did, there would be a couple options for bringing
the count down by reusing `NodeKinds` and disambiguating based on the
token kind (the 29 infix operators as an example). Or we could just undo
this.

I'm expecting this to yield a small improvement. I'll see if I can get
better numbers since my machine's not really reliable, but here are some
basic values.

Also suggesting to draw the use of `::RawEnumType` for `TokenKind`,
since bit packing appears to work without it. Hoping the `static_assert`
is easier for people to understand the size of the field.

With the change:

```
----------------------------------------------------------------------------------------------------------------------------
Benchmark                                                 Time             CPU   Iterations      Bytes      Lines     Tokens
----------------------------------------------------------------------------------------------------------------------------
BM_CompileAPIFileDenseDecls<Phase::Parse>/256         50399 ns        50359 ns        14336 104.588M/s 3.87217M/s 21.8629M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024       237823 ns       237629 ns         3072 136.721M/s 4.11986M/s 24.2058M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096       997645 ns       996771 ns          768 142.343M/s 4.04105M/s 23.9363M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384     4020308 ns      4018319 ns          192 152.041M/s 4.05966M/s 24.0874M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    16691390 ns     16683058 ns           48 151.317M/s 3.92374M/s 23.2936M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144   75265735 ns     75233476 ns            8 135.842M/s 3.48421M/s 20.6862M/s
```

Without the change:
```
----------------------------------------------------------------------------------------------------------------------------
Benchmark                                                 Time             CPU   Iterations      Bytes      Lines     Tokens
----------------------------------------------------------------------------------------------------------------------------
BM_CompileAPIFileDenseDecls<Phase::Parse>/256         51515 ns        51480 ns        13312 102.312M/s 3.78789M/s  21.387M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024       241040 ns       240900 ns         3072 134.865M/s 4.06392M/s 23.8771M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096       985593 ns       984657 ns          768 144.094M/s 4.09077M/s 24.2308M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384     4109327 ns      4105496 ns          192 148.813M/s 3.97345M/s  23.576M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    17459655 ns     17446006 ns           48   144.7M/s 3.75215M/s  22.275M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144   80802815 ns     80737489 ns            8 126.581M/s 3.24668M/s  19.276M/s
```

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2024-12-17 00:58:54 +00:00
Richard Smith c04d62a7d1 Ensure that all allocas are created in the entry block. (#4685)
Non-entry-block allocas will allocate new stack memory each time they're
reached, resulting in leaking stack memory over time for allocas in a
loop. Move all such allocas to the entry block instead, and use an LLVM
intrinsic to mark when the lifetime of the variable actually begins.
2024-12-17 00:55:06 +00:00
Richard Smith 3645143e27 Add solutions for advent of code 2024 day 1 to examples/. (#4673)
In order to support these examples, this adds two new builtins to the
toolchain: `print.char` and `read.char`, which map to the libc functions
`putchar` and `getchar`.
2024-12-17 00:47:51 +00:00
josh11bandJosh L b25117b508 Do not resolve the declaration when forming a specific for use in an eval block (#4692)
When substituting into a generic in order to form a generic eval block,
we form `SpecificId`s to track the list of arguments that should
eventually be used to form a specific referenced by the eval block.
Values within that specific are not needed and won't ever be used, so
it's safe to skip forming them in the first place.

Co-authored-by: Josh L <josh11b@users.noreply.github.com>
2024-12-17 00:39:48 +00:00
Jon Ross-Perkins 76055de063 Shuffle around yaml formatting in .clang-tidy (#4690)
I was looking at this again, considering how best to add new checks, and
realized we could just change the format and probably get better deltas
in the future.

This change should just be formatting, with no functional impact.
2024-12-16 23:59:06 +00:00
Richard Smith a10c79569e Model Core.Int as a class type (#4644)
Instead of treating `Core.Int` as the toolchain's builtin `IntType`,
model it as a class that adapts the builtin type. This aligns us better
with the intended language model, gives an associated library for
`impl`s involving `Core.Int` to live within, and opens the door adding
member functions to `Core.Int` if we decide that is desirable.
Remarkably it also seems to make the formatted SemIR a little smaller,
because a call to a generic class generates less IR than a call to a
function.
2024-12-16 22:19:23 +00:00
Jon Ross-PerkinsandRichard Smith f922988c8c Update the vscode language server setup (#4663)
Switches from js to ts, and starts bundling files in order to produce a
better package for deployment. Fixes the README.md to be a more
appropriate front page, moving dev content to development.md. Makes the
path to `carbon` configurable so that it's more stable than just running
in `bazel-bin`.

This is built using suggestions from samples at
https://github.com/microsoft/vscode-extension-samples/tree/main/lsp-sample
and
https://github.com/microsoft/vscode-extension-samples/tree/main/esbuild-sample.
Note the esbuild in particular comes from complaints from `vsce` to use
an option from
https://code.visualstudio.com/api/working-with-extensions/bundling-extension,
and esbuild is just the first option detailed there (I have no real
opinion on options).

I'm bumping the version, and will do a release after merging.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-12-16 16:05:17 +00:00
Richard Smith 7b45a28a82 Fix lowering of array indexing with an int literal. (#4686)
Such indexing operations are created by array initialization. Since we
switched integer literals to be of type IntLiteral we've been attempting
to index arrays with the (empty) representation of an IntLiteral rather
than with an actual integer value.
v0.0.0-0.nightly.2024.12.16 v0.0.0-0.nightly.2024.12.15
2024-12-14 04:47:32 +00:00
Jon Ross-Perkins 1d5d4617ca Fix mem usage tracking of semir (#4684)
The call got misplaced during refactoring.
v0.0.0-0.nightly.2024.12.14
2024-12-14 01:13:34 +00:00
Dana Jansens 18d99350a9 Add a --remote switch to new_proposal.py (#4681)
If the user's fork is not named 'origin' then the script will fail and
needs to know the user's remote name.

Fixes #1899
2024-12-13 15:46:27 +00:00
Dana Jansens c7ae2a7b18 Avoid printing enums as characters (#4676)
Given code like the following:
```
auto kind = ConversionTarget::Kind{0};
CARBON_CHECK(!loc_id.is_valid(), "hello {0} world", kind);
```

Currently we would print 'hello <the next line>', as the check string
would be treated as terminating at the '{0}', so it does not print the
rest of the string or a newline. This is because ConversionTarget::Kind
is an enum with underlying type `int8_t` which is a char, and
llvm::formatv does not look if the type is an enum and treat is
specially. So it prints it as a char rather than a number, which in this
case is a nul terminator.

With this change, the '{0}' value will be converted to a larger integer
before being passed through to llvm::formatv so that char-sized enums
will print as a number, and the result is that we will print 'hello 0
world\n' as the developer intended.
2024-12-13 14:48:05 +00:00
Jon Ross-Perkins aee098b8e2 Clean up missing library in test (#4678)
Noted in #4677
v0.0.0-0.nightly.2024.12.13
2024-12-13 00:34:57 +00:00
Jon Ross-Perkins 55c257bc93 Use a filename without a line number as a cue for autoupdate. (#4677)
This is so that diagnostics which lack a location get split file
clustering.
2024-12-13 00:05:33 +00:00
Richard Smith e71fd07dc6 Support stringifying tuple values. (#4664) 2024-12-12 23:23:44 +00:00
Richard Smith 0d835699e3 Import support for array types. (#4675) 2024-12-12 23:09:28 +00:00
Richard Smith 3e0fdd04eb Lower global variables as global definitions, not global declarations. (#4674) 2024-12-12 22:23:00 +00:00
Boaz Brickner 9ea1534535 Allow defining .h files in tests without trying to compile them as Carbon files (#4667)
This would be used to test interop with C++.
#4666
2024-12-12 22:19:24 +00:00
Jon Ross-PerkinsandDana Jansens 3ce0df67bb Add Dump functions to Check, Parse, and Lex (#4669)
- Provide `Check::Dump(context, arg)` and similar.
- gdb and lldb should do contextual lookup, and `call Dump(*this,
Lex::TokenIndex::Invalid)` has been tested with gdb.
- Since this is only for debug, keeps the functions fully separated from
code.
- Uses alwayslink to ensure objects are correctly linked, even though
there are no calls.
- `-Wno-missing-prototypes` is needed when we don't have forward
declarations.
- Code is not linked in opt builds, using `#ifndef NDEBUG`.
- This probably could be doing something in BUILD files with a
`select()`, but the `#ifndef` seemed easier.

This is based on #4620, but uses free functions instead of member
functions.

Co-authored-by: Dana Jansens <danakj@orodu.net>

---------

Co-authored-by: danakj <danakj@orodu.net>
2024-12-12 20:51:02 +00:00
Richard Smith 79ba184dab Provide a location for monomorphization failures resulting from TryToCompleteType. (#4670)
Almost all callers actually could never fail and nearly all of those
already `CHECK`-failed on failure. Add a new overload for that case, and
add a location parameter for the one remaining call.
v0.0.0-0.nightly.2024.12.12
2024-12-12 00:44:07 +00:00
Richard Smith 758b6c42ba Produce a note indicating where the specific was used from if monomorphization fails. (#4662)
Also fix a bug in `Context::GetClassType` that previously tried to
complete the class type before returning it. That's not correct --
`GetCompleteTypeImpl` is only appropriate for cases where the type can
trivially be completed and completing it can't fail -- and led to
infinite recursion with this change because we would call `GetClassType`
when producing a diagnostic if completing that class type failed.
2024-12-11 22:34:10 +00:00
Richard SmithandJon Ross-Perkins 042ac39426 Pacify CHECK failure on invalid code. (#4665)
Found by fuzzer.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-12-11 18:53:58 +00:00
Jon Ross-Perkins 61c0a8b676 Make more use of llvm STLExtras (#4668)
This is essentially the result of looking at `.begin()` uses. We also
frequently do `std::shuffle`, but unfortunately STLExtras doesn't
provide a wrapper for that.
2024-12-11 18:16:38 +00:00
Richard Smith 47285b6207 Include a fully-qualified name when stringifying types. (#4657)
For example, format the `ImplicitAs` interface as `Core.ImplicitAs`
rather than simply `ImplicitAs`.

When importing an entity in a namespace, also import a declaration of
the enclosing namespace if necessary so that we can determine its name.
2024-12-11 16:33:53 +00:00
8e8d570571 Proposal: Variadics (#2240)
Proposes a set of core features for declaring and implementing generic
variadic
functions.

A "pack expansion" is a syntactic unit beginning with `...`, which is a
kind of
compile-time loop over sequences called "packs". Packs are initialized
and
referred to using "pack bindings", which are marked with the `each`
keyword at
the point of declaration and the point of use.

The syntax and behavior of a pack expansion depends on its context, and
in some
cases by a keyword following the `...`:

- In a tuple literal expression (such as a function call argument list),
`...`
iteratively evaluates its operand expression, and treats the values as
    successive elements of the tuple.
- `...and` and `...or` iteratively evaluate a boolean expression,
combining
the values using `and` and `or`, and ending the loop early if the
underlying
    operator short-circuits.
-   In a statement context, `...` iteratively executes a statement.
- In a tuple literal pattern (such as a function parameter list), `...`
iteratively matches the elements of the scrutinee tuple. In conjunction
with
    pack bindings, this enables functions to take an arbitrary number of
    arguments.

---------

Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com>
v0.0.0-0.nightly.2024.12.11
2024-12-11 01:58:40 +00:00
Richard Smith 14724a5c9a Switch from recommending a local workspace extension to recommending our published extension. (#4661) 2024-12-10 22:39:21 +00:00
Jon Ross-Perkins e98995c936 Update the vscode extension for publishing. (#4660)
A few initial fixes just so that publishing works.


https://marketplace.visualstudio.com/items?itemName=carbon-lang.carbon-vscode
2024-12-10 22:03:48 +00:00
Jon Ross-Perkins 87b3671330 Refactor single-unit checking out of check.cpp (#4649)
This is primarily moving code around, to try to create a logical split
of the code in check.cpp, makingthe API boundaries clearer.

There's one small, deliberate logic change around false returns from
`HandleParseNode`, where before there was a `CARBON_CHECK` instantiated
by the `#define` (per `NodeKind`), and now it's outside the `#define`
(done mainly because the message didn't keep up with the `Handle##Name`
-> `HandleParseNode` rename).
2024-12-10 21:00:51 +00:00
Richard Smith 92201ceb10 Rename various TryToCompleteType functions to better describe what they do. (#4658)
As requested in review of #4652.
2024-12-10 20:56:37 +00:00
Boaz Brickner fe8b42148f Mark some //common, //toolchain/driver, //‎toolchain/install tests as small per 'Test execution time' warning (#4646)
These tests only take between 0.1s and 1.4s.
2024-12-10 20:55:58 +00:00
Richard Smith d81ed4b58f Rename mutable accessor in InstBlock store. (#4659)
Mutating a block is a strange and rare operation and shouldn't have an
innocuous name like `Get`.
2024-12-10 20:23:24 +00:00
Richard Smith eabe9f117a Track complete types required by a generic. (#4652)
When a generic requires a symbolic type to be complete, add a new
`require_complete_type` instruction to the generic eval block. During
monomorphization of such an instruction, require that type to be
complete.
2024-12-10 03:00:28 +00:00
Richard SmithandGeoff Romer e75ef34591 Improve diagnostics for missing qualified names. (#4638)
Mention the scope in which the name wasn't found.

---------

Co-authored-by: Geoff Romer <gromer@google.com>
v0.0.0-0.nightly.2024.12.10
2024-12-09 23:48:46 +00:00
Dana Jansens 4e3b8c9775 Disable modernize-use-trailing-return-type on MATCHER_P (#4656)
clang-tidy gives a false positive on the use of the MATCHER_P macro.
2024-12-09 21:13:40 +00:00
josh11bandJosh L bd0f620583 Add more tests of indexing a tuple with a non-literal (#4650)
Co-authored-by: Josh L <josh11b@users.noreply.github.com>
v0.0.0-0.nightly.2024.12.09 v0.0.0-0.nightly.2024.12.08 v0.0.0-0.nightly.2024.12.07
2024-12-06 23:22:21 +00:00
Boaz Bricknerandjonmeow daba2c72cf [NFC] Convert NameScope from struct to class (#4623)
This is a preparation change for adding name poisoning support
(https://github.com/carbon-language/carbon-lang/issues/4622), which is
expected to require more elaborate logic around NameScope since a name
can be not defined yet, defined, or poisoned.

The API separates looking up a name from getting the full entry since we
have cases where the entries are invalidated between the time we're
looking for the name and when we access (and sometimes modify) the
entry.

This change has the following benefits:
* `names` and `name_map` are internal to `NameScope` and are guaranteed
to match.
* `extended_scopes` and `import_ir_scopes` can not be manipulated (only
new scopes can be added).
* `inst_id`, `name_id` and `parent_scope_id` are constants.
* `has_error` can only be mutated from false to true.

---------

Co-authored-by: jonmeow <jperkins@google.com>
2024-12-06 21:50:50 +00:00
Jon Ross-Perkins e7a86b03c6 Remove offsets from InstId formatting, trying to name more (#4645)
The offsets were originally added to deal with churn from builtins in
the raw semir. In textual semir, we mostly see instruction IDs for
imports, and builtins have also settled down more.

On imports, where possible, use the `EntityNameId` for an import instead
of printing an instruction. Next, show the source location if we have a
node. Only show the instruction if there's no location.

This also exposes `Parse::Tree` and `TokenizedBuffer`, so that we can
pass a `SemIR::File` without the component parts. In particular this
allows us to get the `TokenizedBuffer` for import IRs without
substantial structural modifications. We may want to make these optional
for serialized `SemIR` later, but the nodes/tokens contain source
location, which we'd need for debug information -- so it's not clear how
much we can really make them optional without substantial information
loss.

Reduce arguments to just `File` in a few spots, as a result of the
accompanying `TokenizedBuffer` and `Parse::Tree`. Also updates style to
pass around `const File*` where the reference is maintained, instead of
`const File&`.

I was considering keeping a direct reference to the tree and tokens on
`Context`, but initially my thought was it wouldn't make much
difference. I can re-add those if desired, just as direct caching of the
`File` fields.
2024-12-06 21:17:24 +00:00
David Blaikie a2c939a2a2 Add an instruction for vptr initialization (#4633)
This fixes a crash in lowering, at least (though only initializes the
vptr to
null for now) - certainly open to naming feedback on the instruction, or
the
exact semantics (we could have a global vptr instruction that's
referenced from
the existing instructions for reading globals, for instance).

I guess we'll want one type parameter for the vptr_init instruction,
which is
the type that this is a vptr for? (can do that here or in a follow-on
patch)
2024-12-06 18:51:36 +00:00
David Blaikie 5719687438 Fix crash in lowering use of a global variable (#4631)
Actually presenting two options in this one review - if you look at the
specific
commits in this PR, the first commit represents my first attempt - and
if you
look at the overall PR change for the second attempt.

But I'm totally open to completely different approaches/ideas - these
were just
my rough guesses.
2024-12-06 17:34:29 +00:00
Richard Smith cd1ecf1297 When a builtin function expects type T also allow an adapter for T. (#4643)
Extends the set of function signatures that support being given a
builtin definition to include cases where a parameter or return type is
an adapter for a supported type. For example, if we can give a builtin
definition to `Add(a: i32, b: i32) -> i32`, then we can also give a
builtin definition to `Add(a: MyI32, b: MyI32) - >MyI32` where `MyI32`
adapts `i32`.

This is a prerequisite for changing `Core.Int` to be a class type that
adapts the builtin int type.
2024-12-06 03:13:05 +00:00
David Blaikie 2bb520718d Enable gdb_index unconditionally for gdb usage (#4642)
Also make `-gsimple-template-names` lldb-only due to it tripping up gdb
in some cases (I came across it breaking SmallVector pretty printing
where gdb wouldn't associate a simplified type named declaration with a
simplified type named definition in another translation unit - seems gdb
can associate a type decl/def when it sees both (if you step into both
translation units or otherwise trigger gdb loading/parsing them) but it
doesn't seem able to /search/ for the type definition). Filed
https://sourceware.org/bugzilla/show_bug.cgi?id=32421 for this.

A couple of other things I'd like to do, but don't know how:
* It might be nice to allow opting into or out of gdb_index (with the
  default being 'on' for gdb_flags, but you could opt out). But doesn't
  seem super important.
* We should turn off fission by default, it seems - bazel has trouble
  making the .dwo files available at the same path as is in the binary
  especially on partial rebuilds. (not sure if we can do that, I guess
  we can make fission a no-op/doesn't add any flags, even if we can't
  change the fission default in bazel itself)
* can we have gdb_flags imply/disable lldb_flags? (so you can use
  --features=gdb_flags without always having to add
  --features=-lldb_flags)
v0.0.0-0.nightly.2024.12.06
2024-12-05 23:48:12 +00:00
Jon Ross-Perkins 3ee41222e0 Add video from LLVM Dev meeting (#4641) 2024-12-05 22:34:53 +00:00
Richard Smith ead09da2d5 Remove some assumptions that the object representation for a type is that type itself. (#4640)
This slightly improves the handling of adapters, by making them copyable
in some cases when their adapted type is copyable.

Refactor some of the repeated checks for properties of value
representations.

In passing, move all the adapter tests in check/testdata/class to a
subdirectory since we now have quite a few of them.
2024-12-05 22:26:10 +00:00
Jon Ross-Perkins 1cba3328f7 Finish removing BuiltinInstKind (#4637) 2024-12-05 22:07:51 +00:00
Richard Smith e4eeacabe7 Convert unsupported qualifier test to be no_prelude. (#4639)
Stop relying on `i32` being a builtin type.
2024-12-05 20:31:43 +00:00
Jon Ross-Perkins 27275e6729 Change the IdBase operator== to fix reversed operator warnings (#4636)
I believe our flags enable the warning by default, it's just that it
doesn't catch this in clang-16 (maybe more; I reproduced with clang-18
and didn't keep digging). For example:

```
toolchain/parse/tree_test.cpp:86:28: error: ISO C++20 considers use of overloaded operator '==' (with operand types 'value_type' (aka 'Carbon::Parse::NodeIdInCategory<Carbon::Parse::NodeCategory::Decl>') and 'AnyDeclId' (aka 'NodeIdInCategory<NodeCategory::Decl>')) to be ambiguous despite there being a unique best viable function [-Werror,-Wambiguous-reversed-operator]
   86 |   EXPECT_TRUE(*any_decl_id == any_decl_id2);
      |               ~~~~~~~~~~~~ ^  ~~~~~~~~~~~~
```

The different `operator==` approach works except for with
`Parse::NodeId::Invalid`, which seems easy to replace with a
`.is_valid()` check.
2024-12-05 19:37:32 +00:00
Boaz Brickner 1409666e6a In indirect_import_member test, make the alias avoid name poisoning (#4635)
This is a preparation change for introducing name poisoning
(https://github.com/carbon-language/carbon-lang/issues/4622), which
would have broken this test.
2024-12-05 18:24:56 +00:00
Jon Ross-Perkins efab39cbd9 Remove InstId::Builtin members (#4632)
- `InstId::Builtin<Inst>` -> `<Inst>::SingletonInstId`
- `InstId::PackageNamespace` -> `Namespace::PackageInstId`
2024-12-05 18:13:46 +00:00
Richard Smith d79d9e0884 Factor out common work of determining how inty a type is. (#4634) v0.0.0-0.nightly.2024.12.05 2024-12-05 01:49:52 +00:00
Richard Smith 1b13125d19 Avoid relying on an implicit conversion in builtin lowering test. (#4630)
This test is trying to explicitly test builtin functions, so use an
explicit call to a builtin function for the conversion too.
2024-12-05 01:33:42 +00:00
Richard Smith 9ed65775fb Add missing library declarations to test. (#4629) 2024-12-05 01:32:11 +00:00