Commit Graph
1897 Commits
Author SHA1 Message Date
Jon Ross-Perkins 2425e28e3a Collapse names into VarStorage (#3116)
This removes BindName, putting name information directly on VarStorage.
As a side-effect of updating semantics_ir_test for this change, I also
noted that function bodies were being generated as invalid YAML so am
fixing that (just `{}` to `[]` bracketing, otherwise the test wouldn't
work anymore).

Because names are now available, I've updated lowering to use them for
vars.

In the SemIR formatter, the name is now repeated because it's a
parameter to VarStorage. I believe this is just default behavior, and
we'd have to special-case VarStorage to remove it because it's automatic
argument printing in action. On the balance, it felt like letting it
print was reasonable.

I've noted in places that the name on VarStorage is expected to be
optional, but am not adding support because I'd have no way of testing
it at present.
2023-08-22 18:55:31 +00:00
Richard Smith e05523db21 Fix a stack use after scope and a heap use after free found by fuzzing. (#3126)
There were two issues contributing to this crash:

- Primarily, the issue is that we queue up diagnostics and don't format
them into a string until we reach the end of compilation. In some code
paths in the driver, we destroyed the Semantics IR object before this
happened. But diagnostics can contain references to Semantics IR
objects, such as strings stored in the string table, which can lead to a
use after destruction bug.

This is fixed by ensuring the diagnotics consumer is flushed before
destroying any of the objects that it can refer to. The current approach
to this is not especially clean, unfortunately, but this requires
fighting C++ as this isn't the order in which it wants to destroy
things.

- This issue was obscured by the Semantics IR's string table holding a
reference to whatever underlying storage it was given rather than its
own string storage, so sometimes it would hold a reference to a string
from the source file, and sometimes a string from the tokenized buffer's
string table. The diagnostics were always flushed before the source file
was destroyed, but not before the tokenized buffer was destroyed. So to
see the issue, you'd need to have a string literal with certain contents
followed by an identifier with a name that matched those contents.

The crash is made more reliable by holding references to the Semantics
IR's string map in its string table, rather than references to someone
else's strings. This also fixes a latent bug where passing a string
temporary to SemanticsIR::AddString would store a dangling reference in
the string table. Incidentally, AddString is also changed to perform
only one hash table lookup rather than two for each added string.
2023-08-22 00:51:13 +00:00
Richard Smith d74b8f0497 Fix diagnostic messages that don't end in a period. (#3127) 2023-08-21 23:59:18 +00:00
Richard Smith d3ac21195c Flush diagnostics before printing semantics IR rather than after. (#3128)
This matches what we do for all other dump output.

No test: this is just changing the behavior of a dump mode, and is
really awkward to exercise without adding back in something like a lit
test to observe the behavior when stdout and stderr go to the same
place.
2023-08-21 23:58:37 +00:00
maan2003 7c891fdacd Language Server (#3112)
Add a language server for carbon as part of GSoC.

This currently does code outline using toolchain parser.

See development steps in utils/vscode/README.md for running and using
language server.
2023-08-21 18:39:31 +00:00
Jon Ross-Perkins 4ae0fa6f86 Adjust handling of cases where conditions are missing. (#3119)
In #3064, code was changed to look at a future token. This is an issue
because the parser is set up to enforce that tokens aren't used without
being consumed. That's part of #3118; related validation fails. Also,
since it's not necessarily the open paren that was consumed, it could be
a different opening symbol, which the closing symbol handling doesn't
check.

Under this approach, it's tracked whether an open paren was consumed,
and the open paren is associated with the state. That's more aligned
with how the parser expects to be fed information.

In paren condition handling for if and while, I'm also adding some
special casing for `if {` in particular to not assume the `{` is a
struct. I just think that this will come up somewhat often and the
resulting output is better this way (an error either way). I'm not doing
similar with `for` because there's already some `var` handling there,
and I'd need a little more time to think about structure -- whereas
right now I'm just trying to fix the crashes (`if {}`, `if []`, etc).

Fixes #3118
2023-08-18 23:04:30 +00:00
Lucile Rose NihlenandChandler Carruth fe93c0225f add documentation about debugging on macOS (#3114)
Adds some notes about the required build flags for lldb to successfully
find the symbols on macOS debug builds. Also adds a recommended debugger
configuration for interactive debugging in VSCode on macOS.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-08-18 19:01:50 +00:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique bc5edc9992 Lowering for array index. (#3111)
This PR implements the lowering for array element access. Besides, this
creates a function for pointerTY to avoid code duplication.

---------

Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-08-18 16:31:51 +00:00
a2b4cabeaa Switch the toolchain to the new CLI library. (#2979)
This also tries to restructure the command line interface to the
toolchain a bit to make it start operating more like a compiler that
could be integrated into a build system rather than primarily as
a testing tool.

1) This switches form a `dump` subcommand to a `compile` subcommand
   which has "dump" actions that can be enabled within it.

2) A distinct set of compile _phases_ that match the toolchain
   structure:
   - `lex` to run the lexer
   - `parse` to run the parser
   - `check` to fully check that the code is valid
   - `lower` to lower to LLVM's IR
   - `codegen` to generate executable code

3) The codegen phase has two output formats: textual assembly and
   a binary object. These outputs can be configured, with a default for
   an object when writing to a file and more firm default for textual
   assembly when writing to stdout.

4) Select and expose the use of the LLVM host detection to compute
   a default code generation target in the driver so that the command
   line interface can reflect this. For example, the `help` output will
   include the default target.

5) The `//toolchain/codegen` library APIs have been restructured a bit
   to make the code flow a bit more naturally when implementing the new
   command line structure. No real changes to the logic though.

There are also some minor tweaks to the command line interface based on
trying to use the shortest names for things that still seem likely to be
learnable for users:

- Switched `target-triple` to just `target`: the "triple" component to
  this name is historical and can be confusing. For example, almost all
  "triple" strings have more than three components today.

- Switched to just `--output` as now the fact that it is a file can be
  configured in the documentation -- it will render as `--output=FILE`.

This also adds support for two custom output filename modes. First, when
no output is specified, we now compute one in the conventional way for
compilers by removing the file extension of the input file and replacing
it with `.o` for an object file output or `.s` for an assembly file
output. This matches the behavior of Clang and GCC for example.

Second, output to stdout is enabled with the special output file name of
`-` since it is no longer the default. This also follows the convention
of most compilers and many other command line tools to use `-` as a file
name to signify using standard in/out pipes.

There are still some rough edges here that I suspect could be improved,
but this seems like a good start of switching over to a complete
argument parser.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Lucile Rose Nihlen <luci.the.rose@gmail.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-18 01:06:09 +00:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique 14f489dff1 Lowering bug fix for tuple indexing. (#3110)
The existing tuple index lowering was crashing while accessing the
element from return value. This PR fixes the bug.

Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-08-16 20:59:56 +00:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique 6af1c435e0 Lowering for array (#3095)
Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-08-16 18:31:18 +00:00
Jon Ross-Perkins 2e290b9e73 Fix missing APSInt include. (#3107)
Works by chance right now, but should be expected to break in future
LLVM versions.

The builtin kind include is removed as unused.
2023-08-16 18:17:57 +00:00
db91db9097 Add a rich command-line argument parsing library. (#2978)
This library is designed around supporting the kinds of use cases we
expect in the toolchain and other Carbon tools. It supports subcommands,
options, and prints help.

There is still a decent chunk of work to be done to finish polishing
this, but it should give us a solid starting point.

For details about the library and a brief roadmap, see the main comment
in `command_line.h` which provides a comprehensive overview of the
library and a roadmap of the remaining work.

Porting the driver to this went fairly well, but did require some
changes. That port is separated into a follow-up PR #2979.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Lucile Rose Nihlen <luci.the.rose@gmail.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-16 17:15:55 +00:00
Jon Ross-PerkinsandRichard Smith 09242eeebd Run clang-tidy over toolchain code. (#3105)
Amongst other changes, this removes the implicit StringRef cast in
formatter code due to google-explicit-constructor. Only one line of code
relied on the implicit cast, and style says "Do not define implicit
conversions."
(https://google.github.io/styleguide/cppguide.html#Implicit_Conversions).

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-16 15:21:58 +00:00
Jon Ross-Perkins fa857e42be Rename //testing/util to base (#3104)
Renaming per #3100
2023-08-15 21:28:56 +00:00
Prabhat Sachdeva 90d2d075f2 Explorer: add blank lines between trace of different sections (#3096) 2023-08-15 21:10:12 +00:00
josh11bandRichard Smith e15bdbb36c Update generics overview (#3061)
This reflects changes from a number of approved proposals:

- #2138 : "generic" -> "checked generic", "template" -> "template
generic"
- #2360 : "type", "facet type", "facet". Note: I am not using the term
"generic type" from #2360 since that meaning conflicts with the
generally accepted meaning of "generic type" of a type with a
compile-time parameter.
- #2760 / #2770 : internal/external impl -> extending impl
- #2964 : "symbolic constant" and "template constant"

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-15 19:58:23 +00:00
Jon Ross-Perkins 357baaeef8 Rename //explorer/common to base (#3103)
Renaming per #3100
2023-08-15 19:23:23 +00:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique a6ba7827cd Semantics for array index. (#3099)
Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-08-15 18:17:42 +00:00
Jon Ross-Perkins a692fb89a3 Rename //toolchain/common to base (#3101)
Renaming per #3100
2023-08-15 17:47:39 +00:00
Jon Ross-Perkins 320a58f97e Small file_test refactorings (#3097)
Just some small refactorings stemming from #3073.

AddCheckLines -> BuildCheckLines because the two lists are now fully
separate. Adding is_blank to be more direct about behavior than the
Print call.
2023-08-14 20:33:01 +00:00
Jon Ross-Perkins 956d65e333 Remove stray semicolon (#3098) 2023-08-14 20:30:39 +00:00
Clayton GearhartandClayton Gearhart 3628e22ca2 Put python object definition and assignment on same line (#3057)
Where there was an object initialization immediately followed by
assignment I condensed it to one line, which seems to be the convention
looking at other files. I also deleted the repetition of 'of' in some
c++ comments because it was grammatically incorrect.

---------

Co-authored-by: Clayton Gearhart <claytongearhart240@gmail.com>
2023-08-14 07:38:02 +00:00
Prabhat Sachdeva da35a9806a Explorer: make pattern match trace consistent (#3093)
Makes the trace output for pattern match more consistent with the rest
of the trace by using `Match()` and `Indent()` prefix methods and
wrapping the code in the trace with backticks.
2023-08-12 00:02:41 +00:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique 9228487d46 Semantics for array type (#3087)
Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-08-11 23:06:13 +00:00
Richard Smith 90cca11ea3 Document the de facto conventions for library dependencies in explorer (#3092)
Unlike in the toolchain, we have not been preferring LLVM facilities
over standard ones. Update the documentation to describe this and
provide some rationale.
2023-08-11 19:22:58 +00:00
Prabhat Sachdeva c5667c0e10 Explorer: add trace for InstantiateType (#3041)
Adds trace output for `Interpreter::InstantiateType` method.
2023-08-11 18:05:30 +00:00
Prabhat Sachdeva 28f563f1ad Explorer: Replace StringRef with string_view in trace stream (#3090) 2023-08-11 03:17:26 +00:00
Richard Smith 883888e22c Produce label names based on the reason for the branch. (#3086)
When producing textual semantics IR, map each branch instruction back to
the construct that produced it and use that to determine a name for the
corresponding block label.
2023-08-11 01:20:29 +00:00
Prabhat Sachdeva 2a74c807d8 Explorer: Add heading and sub-heading methods to trace (#3088)
Instead of manually putting symbols before and after heading, this
introduces two methods `Heading` and `SubHeading` inside `TraceStream`.

Both methods take `llvm::StringRef` as parameter, formats the given text
as follows and adds it into the output stream.

**Heading**
```
* * * * * * * * * *  heading  * * * * * * * * * *
-------------------------------------------------
```
**Sub heading**
```
- - - - -  sub heading  - - - - -
---------------------------------
```

Note: both methods assert that tracing should be enabled.
2023-08-10 22:40:21 +00:00
Richard Smith 7b22173cba Add precedence rules for assignment operators to the precedence diagram. (#3083)
Also indicate what can appear within parentheses.

This is intended to be a clarification, not a design change. Note that
while we previously described the operand of `++` or `--` as being
simply an expression, the operand can never be anything other than the
kinds of expression the diagram now shows due to the expression category
rules added in #2006.

Fixes #3079.
2023-08-10 22:35:51 +00:00
Richard Smith 6e01b90394 Rename some semantics nodes and some IR names to make them match better. (#3085) 2023-08-10 21:48:50 +00:00
Richard Smith 9c516c40b9 Always create a StructType node for every syntactic struct type. (#3084)
Don't elide the StructType node if we've already created an equivalent
type. Its spelling and location may be interesting to diagnostics,
tooling, debug information, etc.
2023-08-10 20:30:39 +00:00
Richard SmithandChandler Carruth 6cbf280a68 Add formatted textual IR output (#3056)
Add a textual IR format to the toolchain.

The exact details of the format are somewhat arbitrary right now, and I
expect them to change as we refine the semantics IR model, but at the
moment they're somewhat directly following the current structure of the
IR.

Semantics tests currently test both the "raw" format, which shows the
details of the representation, and the textual format, which is somewhat
higher level. We may want to revisit that decision once the textual
format is a bit more stable, and test only the textual format in most of
these tests, but for now it seems prudent to keep both sets of tests.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-08-10 19:41:39 +00:00
Richard SmithandChandler Carruth fc5a9541ce Update precedence rules to match design. (#3081)
- Only allow assignment at the top level in an expression statement.
- Allow both negation and complement as subexpressions of both bitwise
  and numeric operators.
- Remove parsing support for postincrement and postdecrement.
- Add parsing support for `as` operator.
- Use the same ambient precedence for types and non-type expressions.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-08-10 18:39:37 +00:00
Prabhat Sachdeva 3cca175a57 Explorer: Add source snippets to trace output (#3082)
Adds **Source Snippets** for interpreter and type checker's trace
output.
Note: Source snippets are trace for only declarations and statements.
2023-08-10 18:29:52 +00:00
Richard Smith 5b45c2319f Only allow assignment to durable reference expressions. (#3077)
Now that we support expression categories, use them to determine whether
the left-hand side of an assignment expression is valid.
2023-08-09 19:46:34 +00:00
Prabhat Sachdeva 0bcba9d05e Explorer: Add methods to trace for line prefixes (#3076)
Defines methods into `TraceStream` for adding line prefixes. Instead of
directly using string literals.

Example usage,
```
trace_stream_->Start() << "declaring ... " << ... ;
```
will result in,
```
->> declaring ... 
```
2023-08-09 19:05:19 +00:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique a67aeb5724 Parser for array type. (#3075)
Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-08-09 19:04:44 +00:00
Prabhat Sachdeva 2f70c7d8bb Explorer: Reduce prelude path when printing source location (#3080)
The Prelude path can be very long, which can make the `SourceLocation`'s
`Print(...)` method output unnecessarily verbose. This PR modifies the
`SourceLocation`'s `Print()` method to just print the **filename** and
**line number** for prelude.
2023-08-09 17:57:07 +00:00
Richard Smith 927375f9c6 Make LLVM IR emission a bit more concise and readable. (#3078)
Provide LLVM IR names for named types, and not for unnamed types.

Use shorter names for IR values.
2023-08-09 03:09:03 +00:00
Richard Smith 2947877518 Add support for dereference operator. (#3066)
Also fix a couple of error-recovery issues exposed by the tests for this
change.
2023-08-08 21:55:32 +00:00
Richard Smith 62205763a5 Add support for & operator. (#3055)
Refactor type canonicalization so that we can reuse the same code for
building a `T*` expression and for forming the type of an `&x`
expression.

Add basic computation of expression category in order to check that we
only take the address of durable reference expressions. This is
currently computed on demand rather than being tracked as part of the
semantics node, but in most cases can be determined by looking at only a
single expression, so caching it in the node doesn't seem worthwhile
yet. This decision should be revisited if we start doing more complex
category calculations.

Also add trivial lowering support, but it doesn't work properly yet
because lowering doesn't yet take the expression category into account.
2023-08-08 20:53:50 +00:00
Richard Smith 212188a922 Prefer to put STDOUT CHECK at the end of the file. (#3073)
Allow interleaving of STDOUT and STDERR check lines. Put STDOUT lines
after the line they're attached to, and STDERR lines before. If no
STDOUT check line is attached to any line, then put them all at the end
of the file instead.

This is intended to better handle the case where stdout contains
unreplaced mentions of line numbers, and also reflects that stdout is
typically a consequence of the test rather than commentary on it, so
placing it after the test seems likely to read better.
2023-08-08 19:51:52 +00:00
Prabhat SachdevaandRichard Smith a359da37c4 Explorer: improve type checking trace output (#3039)
Updates the trace output for type checking to be more consistent like
the rest of trace output.
This PR also includes few changes in the lit tests related to trace
output.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-08 17:58:12 +00:00
Chandler Carruth 6d3a4b84cb Mark unary operators as repeating. (#3074)
Currently, they are marked as left-associative which isn't completely
obvious for unary operators and wouldn't match the fact that we have
both prefix and postfix unary operators we intend to allow to repeat
without parentheses: `***x` and `T***`.

This fixes the graph by making the left-associative marker only for
binary operators, and using a separate marker for repeating unary
operators: a diamond.

It also adds a note that Carbon currently only has left-associativity
because the right-associative operators like assignment are statements
in Carbon.

This isn't intended ta be a change of anything in Carbon's design, just
an improvement to the documentation.
2023-08-08 17:57:25 +00:00
0d1e6bd84d Values, variables, pointers, and references (#2006)
Introduce a concrete design for how Carbon values, objects, storage,
variables,
and pointers will work. This includes fleshing out the design for:

- The expression categories used in Carbon to represent values and
objects,
how they interact, and terminology that anchors on their expression
nature.
-   An expression category model for readonly, abstract values that can
    efficiently support function inputs.
- A customization system for value expression representations,
especially as
    seen on function boundaries in the calling convention.
- An expression category model for references instead of a type system
model.
-   How patterns match different expression categories.
-   How initialization works in conjunction with function returns.
- Specific pointer syntax, semantics, and library customization
mechanisms.
- A `const` type qualifier for use when the value expression category
system
    is too abstracted from the underlying objects in storage.

---------

Co-authored-by: Geoff Romer <gromer@google.com>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Co-authored-by: Adrien Leravat <Pixep@users.noreply.github.com>
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-08 07:27:44 +00:00
Geoff RomerandRichard Smith 049fbc1ee4 Handle malformed subscript expressions (#3064)
Closes #3063

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-07 23:14:20 +00:00
Chandler Carruth 7feaed312a Fix bug in the VFS construction. (#3067)
The constructor accepts a `bool` which we were setting to `true` by
converting from a heap allocated, and thus non-null, pointer. But that
heap allocation was dead and leaked, somewhat obviously. Leak checking
is disabled on our CI at the moment, but this was failing for me locally
with our default build on Linux where it uses ASan.

The `true` value also had no effect because the default argument to the
bool parameter is itself, `true`. :sigh:

It's really sad that this compiled. But it's the same thing as using a
non-null pointer in an `if`. There doesn't seem to be any `clang-tidy`
check for a `new` expression that is implicitly converted to a `bool`
type either. I've filled
https://github.com/llvm/llvm-project/issues/64461 requesting a check or
warning to catch this in the future.
2023-08-07 16:59:11 +00:00
Mrinaal Aroraandaroramrinaal ef87f8f2b2 Fix spelling error in README (#3069)
This PR resolves the spelling error in the "Contributing" section of the
README. The sentence "Contribute the language design: feedback on
design, new design proposal" has been corrected to "Contribute to the
language design: feedback on design, new design proposal". This change
enhances the clarity and accuracy of the contribution guidelines.

Closes #3068

Co-authored-by: aroramrinaal <aroramrinaal@gmail.com>
2023-08-07 16:24:44 +00:00