This handles the basics of type and value for structs. Structurally, these look like parameters and arguments (respectively) because expressions/generics may result in multiple IR nodes being generated.
Because `{}` needs to be cast to a type for storage, I'm also adding some validation that's not specific to `{}`, e.g. that `1` shouldn't be valid as a type for storage (previously, nothing errored for that).
This adds more stringification of types, particularly literals, because they come up in value errors now.
ImplicitAs is the result of me mulling whether I'm taking the right approach on type conversions. I think it needs to return a value so that if the implicit cast rewrites the value, the result is accessible to the caller. I may reorient the current TryTypeConversion logic to be more based on the ImplicitAs logic.
I'm trying to reduce the amount of per-SemanticsNodeKind boilerplate, and make mistakes (e.g., SemanticsNodeKind not matching the Make name, misplacing the type, or Get/Make type mismatches) easier to see.
I could've done this with (more) macros, but felt that the template approach was reasonable enough and likely easier to understand/debug. I'm not sure whether there's more that I could be doing with variadics to reduce the amount of factory code, but this feels good right now.
These changes should make output more stable when builtins are added to semantics. By omitting them from nodes and printing nodes as "relative to the last builtin", I should be able to add and remove builtins without automatically affecting every test. Also by printing builtin nodes as `nodeNameOfBuiltin`, it's a little easier to understand what's going on (for me, at least).
This is something we'd discussed. I added a TODO that we may want to eventually make this a map, but I remain uncertain and think it's not something that's going to really cost us if we end up switching back. In the meantime, I think this approach does offer simplicity. As discussed too, the performance overhead of a map may ultimately not be worthwhile here versus the relative memory costs.
This is starting to build out actual lowering logic, for a really simple `fn Main() -> i32 { return 0; }`
Notes for achieving this:
- In semantics, currently function names are bound separate from the signature. When emitting IR, this turns out to be inconvenient because we want to know the name when we process the declaration and the definition. This change addresses that by merging the name into the FunctionDeclaration node, which is also accessible from the definition. It removes the separate BindName. This should be the cause of all the test changes in semantics, because the IR generated changes.
- Add a "Lowering" class which I'm using to hold the llvm builder state. This class now has minimal support for the SemanticsIR generated by the above example.
- In the "Lowering", values from expressions are stored in a DenseMap. I'll keep thinking about whether there's a cleaner way to achieve this, and I'd call it a temporary solution for now. However, this is how the `0` in `return 0` gets properly associated across SemanticsIR instructions, and it'll frequently be an issue in less trivial cases.
This starts handling return types on functions, and comparing types with `return` statements.
Note, errors remain poor because the type literal is currently associated with a builtin, losing the parse_node that specified it. This means we don't have the original source location to associate with, even though it may be helpful to point at the type in source. We could point at the signature overall, but my leaning is that we wouldn't want that long-term, so TODOs for now and may want to change a little about how the parse node is tracked once things are a little further along.
This adds tracking of call information plus basic type checking. It adds a builtin for the empty tuple, mainly so that I have the basis for a default function return type.
As an aside, it also unifies printing within SemanticsIR, fixing a missing comma after callables.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Reals were mostly handled, but this PR adds storage of them. It also switches a little towards the FloatingPointType semantic from TokenizedBuffer.
While real literals like `1.0` were handled, the type literals were not. This just adds `f64`, similar to how I also only support `i32`.
The String type literal wasn't used, so I've added support in lexer and parser. Per discussion with @zygoloid String might be renamed based on the newer type literal plan, but it's still String in explorer and the design, so this is just consistent.
The builtin_types.carbon tests the three basic types that are there right now. The test is added to both parser and semantics so that it's clear what the state is in both stages.
For parameters (and in the future, arguments too; generally comma-separated lists) track two node blocks:
1. param_ir: The complete IR.
2. param_refs: Nodes within the IR that are the "root" parameter.
param_refs should allow quick counting of the # of parameters, and more efficient comparison of call args with function parameters. param_ir should be necessary to generate the actual signature.
In order to construct this, this refactors the node_block_stack into its own class, which is reused in params_stack. These carry references to the underlying SmallVector for lazy modification in order to avoid a dependency cycle with SemanticsIR (also see notes on empty node blocks below).
When finalized, the block pair is pushed onto finished_params_stack. That's because node_stack only has space for one thing, and this is two things -- so I'm essentially choosing a trade-off of adding another stack in order to avoid consuming more space in the expectation that most parse nodes have 0 or 1 things to return, and 2 will be very rare.
As factored, this currently consolidates most empty node blocks into a single canonical empty node block. This is because I think empty blocks, i.e. `()`, will be very common. In order to achieve this, SemanticsNodeBlockStack does lazy creation.
An alternative approach would have been to use 1 node block per parameter. We decided against this in order to reduce the number of vectors being created.
This switches to single list storage of SemanticsNode. The driving motivation behind this is to simplify cross-references within a given IR. Types of nodes will frequently refer to other blocks. This causes a significant increase in the number of cross-references, which can become difficult to manage (and reason about). By reducing to a single list of nodes, cross-references are only needed when crossing IR boundaries.
Because cross-references now only have 2 things to track (IR and index), they can be a regular SemanticsNode and don't need further indirection. This wasn't motivating, but feels like it reinforces the simplification.
Note this isn't being used to deduplicate nodes, at least right now. That could lead to difficult-to-update situations, but also most nodes are associated with the underlying ParseTree::Node in order to track sources for diagnostics; as a consequence, nodes representing equal text in different source locations wouldn't be the same node. There may be future opportunities here, discussed with @zygoloid, but no action is taken at present.
We may eventually want to switch the storage of NodeBlocks to have `[start, end)` ranges instead of individual numbers, but I'm leaving that alone for now.
As an aside, I noticed I was accidentally overloading the copy constructor on SemanticsIR. I've added some disambiguation on that, but am not deleting the copy constructor per style advice (even though the type should never be copied due to storage size).
codespell tries to change `CrossReference -> cross-reference` so disabling it there.
When binding a name, add it to name lookup. On NameReference nodes, use name lookup.
- Switches from "identifiers" to the more generic "strings". Not strictly necessary here, but it's the overall direction I think we've agreed upon and wanted to do it while building more support out.
- Starts doing deduplication of strings.
- On BindName, registers names with name lookup.
- Does name lookup based on the deduplicated string.
- Per discussion with zygoloid, design is intended to be constant-time lookup regardless of the number of parent scopes.
- Adds scopes so that we can track names which will be deregistered from lookup.
Initially I'd added this to lexer, this includes parser and semantics. Also adds ComparableIndexBase to unify a few common cases where <> comparisons are supported.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This is a first pass at what semantic type checking might look like. Types propagate along nodes, we use an InvalidType object when there's an error, and once there's an InvalidType we stop doing so much type checking.
This adds some RealLiteral handling in order to get type mismatches. I'm cautious about creating some real value for SemanticsIR (since the tokenized buffer version is a bit constrained), so I'm not doing that yet. But I will probably need to in order to maintain SemanticsIR having hermetic copies of its data, without a parse tree dependency.
This starts adding builtins with TypeType and IntegerLiteralType. Note, structurally that's all they are, and not directly accessible in any way.
Adds a type field to SemanticsNode. Now, IntegerLiteralType can be identified as having type=TypeType, and IntegerLiteral as type=IntegerLiteralType. The current iteration doesn't do anything for type propagation, because I wanted to avoid making this too big.
This also switches the Identifier IR to instead BindName, with some side-effects. I'd been trying to think how to provide a name for TypeType, and switching around how things worked seemed like a better approach. And while I think it's the right direction (e.g., alias should just be a BindName), I also realized I don't need to name TypeType: there's probably a keyword to refer to the builtin, so it shouldn't use regular name lookup.
As a step towards builtins, provide more blocks. The intent is that any significant scope change will become its own NodeBlock. Builtins should produce the first set of node blocks.
Note, SemanticsIR as set up here isn't handling ordering of import processing -- I haven't thought that through much beyond that we probably want some lighter-weight processing of the parse tree to achieve it. But I think the essence of loading builtins first as their own IR block is... probably right?
I'm thinking about how to handle multiple files, and I think the current IRFactory is useful as a file-focused thing. So shifting/renaming accordingly. (doing this in its own PR to make the history a little cleaner for git's move detection)
This rewrites semantics towards a more pure instruction model, in pursuit of the simple instruction-style output.
I think I can get this approach to type-check as it goes along, but obviously this change doesn't prove that yet. I'm separating it out because it's a large rewrite of the semantics structure, tossing out a lot of what was there before. But I think it does help towards several requests, like setting up a clear path for consolidating duplicate identifiers and making the node style more standardized.
I expect to need to pass multiple args to function calls, that'd probably be storing vectors of args similar to how I'm showing identifiers and integer literals stored.
This removes the semantics namespace because (a) it was getting annoying writing the `::` everywhere, and (b) I think the leaning with Carbon is to avoid namespaces (@chandlerc asked not to put SemanticsIR/SemanticsFactory in a namespace, which is the crux of the issue). But, it's still necessary to avoid name conflicts so I just prefix everything with "Semantics" (still a lot of typing, but no `::`).
This builds on semantics-ir lit support added by #2222
The googletest setup was feeling cumbersome, especially as I'm thinking about how to add more testing: I feel like I'm wrestling with the infrastructure.
The `[[ID1]]` and so on in tests is one advantage of switching: it's easier to do matching of IDs for verification. This is also more agnostic about the numbers than before, something which I'm concerned will be important as I think about builtins.
To explain my builtins thought, I think that needs to be another SemanticsIR with basically names pointing at builtin things. But this (a) creates multiple SemanticsIRs, which would confuse the current singleton approach and (b) starts creating more fluctuation for IDs, potentially impacting the numbers used (also, chandlerc's suggested pointers for some use-cases).
Overall it felt like I was heading towards a situation with googletest where writing the tests would be really difficult, and it was adding to my hesitance to write more code in the toolchain. I'm hoping this acts as a simplification.
Note, the "cp" commit has some incremental changes to googletest that I'd considered for making it easier to add matchers, but ultimately I ended up with this outcome.
This unifies `dump-tokens` and `dump-parse-tree` so that we don't keep writing basically the same code repeatedly.
I'll need to modify the output of SemanticsIR::Print more, and may soon unify printing multiple IRs this way (particularly including the builtins SemanticsIR) but this is intended to offer a starting point.
I may eventually try to unify lit.cfg.py files, but I was thinking about whether that works in various contexts we may run in and eventually decided copying the driver/testdata/lit.cfg.py file would be the easiest solution.
This is how I'm interpreting discussion:
- Basic elements are getting set to an ID.
- SetName exists to assign a name (which can then be referred to later with an identifier expression) to an ID.
- Expressions are broken down into a series of operations which operate on IDs.
So with something like the last test:
```
fn Main() { return 12 + 34; }
```
This becomes:
```
Function(%0,
{IntegerLiteral(%3, 12),
IntegerLiteral(%2, 34),
BinaryOperator(%1, +, %3, %2),
Return(%1),
})
SetName(`Main`, %0)
```
Note I'm treating blocks as fairly equal to the top of a file now, and basically eliminating boundaries between things. That's because we have discussed also supporting code like:
```
fn Foo() {
fn Bar() {}
Bar();
}
```
Here a declaration of a function is occurring inside a code block, so it felt like eliminating the difference was the best choice.
I know you'd commented on the separation of nodes to individual files before; I still think we're going to have a lot of different types of nodes, and so separating them out into individual files makes them easier to browse.
Working on toolchain semantics:
- SemanticsIR is set up as a container for the semantic tree.
- SemanticsIRFactory builds the tree, with separate transformations for each ParseNodeKind.
- ParseSubtreeConsumer is a helper for transforming a ParseTree::Node's children, managing size/nodes to prevent errors.
- The nodes subdirectory contains SemanticIR nodes.
- MetaNode is used to represent nodes which have "sub-classes": Statements, Declarations, and Expressions.
- MetaNodeBlock is used to represent nodes which exist together in a block with name lookup: Statements and Declarations (not Expressions).
This is traversing children first in order to address the RPO format of ParseTree. This means that when lists are formed, they're reversed to be in code-order (`FixReverseOrdering`).
This is still very much incomplete -- the main intent at present is to demonstrate structure.