The doc is more up-to-date than the readme right now, and I'm not ready to migrate it back for the moment, but this should at least make it clear what the status is.
This switches to single list storage of SemanticsNode. The driving motivation behind this is to simplify cross-references within a given IR. Types of nodes will frequently refer to other blocks. This causes a significant increase in the number of cross-references, which can become difficult to manage (and reason about). By reducing to a single list of nodes, cross-references are only needed when crossing IR boundaries.
Because cross-references now only have 2 things to track (IR and index), they can be a regular SemanticsNode and don't need further indirection. This wasn't motivating, but feels like it reinforces the simplification.
Note this isn't being used to deduplicate nodes, at least right now. That could lead to difficult-to-update situations, but also most nodes are associated with the underlying ParseTree::Node in order to track sources for diagnostics; as a consequence, nodes representing equal text in different source locations wouldn't be the same node. There may be future opportunities here, discussed with @zygoloid, but no action is taken at present.
We may eventually want to switch the storage of NodeBlocks to have `[start, end)` ranges instead of individual numbers, but I'm leaving that alone for now.
As an aside, I noticed I was accidentally overloading the copy constructor on SemanticsIR. I've added some disambiguation on that, but am not deleting the copy constructor per style advice (even though the type should never be copied due to storage size).
codespell tries to change `CrossReference -> cross-reference` so disabling it there.
When binding a name, add it to name lookup. On NameReference nodes, use name lookup.
- Switches from "identifiers" to the more generic "strings". Not strictly necessary here, but it's the overall direction I think we've agreed upon and wanted to do it while building more support out.
- Starts doing deduplication of strings.
- On BindName, registers names with name lookup.
- Does name lookup based on the deduplicated string.
- Per discussion with zygoloid, design is intended to be constant-time lookup regardless of the number of parent scopes.
- Adds scopes so that we can track names which will be deregistered from lookup.
Doing some sorting of functions / enums (generally speaking, I've been trying to keep these loosely lexically sorted for lack of a better ordering).
Also removes some code that seems to be dead, and a minor TODO comment fix.
A quick check suggests this isn't necessary -- probably because the implicit enum cast is used for comparisons. I think this adds a lot to the boilerplate feel of these types, so if we can remove it there's a lot less sharing to do.
This is somewhat based on the name vs Name difference, but I figured I'd split it out and just sweep up the API on the whole while looking at a different approach to #2453
This creates a bit of extra cost in adding parse nodes in that a TODO must be added to semantics, but I think the link is going to last this way long-term. In semantics, it reduces the boilerplate of the main for loop and makes it more obvious what's missing, leaving stub functions to be filled in.
This is just the declaration, without initialization. Partly breaking it out because I'm changing the placeholder builtin types.
Might also need to separate out storage of the var from the name bind.
This makes TokenizedBuffer more consistent with ParseTree and SemanticsIR, which also wrap with [] to produce a sequence value.
It also makes it possible in driver.cpp to just prefix the line with a name, so it ends up with:
var_name: [
(content)
]
Noticed this due to bracketing comments on #2443 and trying to think of better answers. With this, we can also say that the [] bracket a variable.
Printing intermediate state should be helpful to be able to examine the input when debugging later steps.
Intermediate state is hard to trace from the individual libraries since they don't know whether `dump` is going to print the state, so this moves some trace logic into the driver which is better equipped to make the decision.
I think it'd be helpful to examine what the parse tree looks like in failure cases.
Note, I'm not sure that the behavior of the recovery situations is correct; the test had been asserting that the parse tree should indicate it's error free. However, this means that the only signal the the driver that the input is invalid is that the diagnostic emitter was used. I think it may be important to have it return a non-zero error to prevent compile, or we can turn these into warnings but then "requiring a space" is wrong.
Either way, that's a concern I have with the pre-existing recovery behavior: I'm just trying to highlight it as I make this change, because now the test is really that nodes aren't tagged with has_error.
Initially I'd added this to lexer, this includes parser and semantics. Also adds ComparableIndexBase to unify a few common cases where <> comparisons are supported.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
LLVM's bazel build has changed a bit, so this updates the tree for that.
LLVM is also moving `llvm::Optional` to match the standard API, but it seemed simpler to just switch to `std::optional`.
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
- Finishes remaining "todo" parse nodes.
- Improving error recovery for invalid designators and structs, so that the parse tree still looks similar to a valid parse tree.
- Call expressions now have the thing being called as a child (of the start) instead of a sibling.
- Use of Start is replacing use of End in several parse nodes, like structs and call expressions.
- Adjusting documentation of parse node structures in an attempt to make it more consistent and understandable.
- The current state for interfaces and if/else is mostly being documented, not altered.
This works on multiple statements to make them better for the bracketing model. Stub nodes are added in more cases of invalid syntax, simply so that the semantics has reliably structured input. Comments in parse_node_kind.def now try to show the expected parse tree structure in postorder form.
This labels If, While, and For a little differently in parse nodes so that at the start of the postorder traversal, it'll already be available to semantics which structure is being processed. I need to do a little more with If in particular, but this felt like a reasonable stopping point.
While this makes significant parser changes, the changes to parser_state.def are minimal, mostly naming-related. The actual flow isn't substantively changed, just a couple minor names and the new As(If|While) state which allows distinguishing IfCondition and WhileCondition.
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
First change towards re-implementing interfaces using the new parser. I kept it small to make sure we are on the same page regarding stack states and how the parse tree should look like.
Co-authored-by: ergawy <kareem.ergawy@guardsquare.com>
This switches the Verify method to walk postorder so that we can see how much subtree_size is really used, and shift towards removing it. It also starts calling Verify.
Also, I think I'd lost the reserve/size check during Parser refactoring, so I'm putting that back in as part of Verify.
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
The intent here is to reduce use of vanilla `int32_t` without a clear indicator of what it's referencing, and to more tightly link references with the underlying types they reference into.
The type name DataIndex doesn't feel great, but I was kind of floundering for a better name.
When there's no semicolon for an invalid EmptyDeclaration, rather than producing nothing, produce an EmptyDeclaration with the original location that led to the error.
Note this removes a direct edit (the only one) of the parse tree's error state. Elsewhere it's an indirection from adding an error node.
This is a first pass at what semantic type checking might look like. Types propagate along nodes, we use an InvalidType object when there's an error, and once there's an InvalidType we stop doing so much type checking.
This adds some RealLiteral handling in order to get type mismatches. I'm cautious about creating some real value for SemanticsIR (since the tokenized buffer version is a bit constrained), so I'm not doing that yet. But I will probably need to in order to maintain SemanticsIR having hermetic copies of its data, without a parse tree dependency.
Longer term, I think this is going to be important as part of associating errors with code after the particular node has been processed, even though it isn't used here.
This is necessary in order to use a bracketing approach; parsing doesn't know the contained format until it parses the first element, which we don't want to do look-ahead for. I think the bracketing is higher value than knowing the format before adding the node.
This changes `return`, `break`, and `continue` to treat the keyword as the "start" and semicolon as the "parent", essentially bracketing the keyword.
Pragmatically this is focusing on making `return` work with only one ParseNodeKind: because `return` and `;` now bracket the expression, we can tightly determine whether the `return` has arguments without looking at subtree size. However, it's possible that `break` and `continue` may in the future take some kind of label as an argument, so the consistency seems beneficial there too.
Note this eliminates the StatementEnd ParseNodeKind, as it's obsolete with this change.
While the Parser has similar divergent states, lists of Expressions tend to be handled more like this. I'm keeping the divergent start state in order to continue support of a distinct error, but I think this organization of FunctionParameter/FunctionParameterFinish will be less surprising.
Adds remaining expression support, and switches the default to Parser2.
Note, this doesn't delete the current Parser yet. I'll just do that in its own PR.
Adds `while` and most of the expression support. Splits apart the fixity test to demonstrate more closely which bits are still failing.
```
//toolchain/parser/testdata:basics/fail_paren_match_regression.carbon.test FAILED in 0.8s
//toolchain/parser/testdata:basics/function_call.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/package.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/structs.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/tuples.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:basics/var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_colon_instead_of_in.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_missing_in.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_missing_var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/nested.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/simple.carbon.test FAILED in 0.8s
//toolchain/parser/testdata:function/definition/with_params.carbon.test FAILED in 0.8s
//toolchain/parser/testdata:operators/fixity_in_call.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:operators/fixity_in_var.carbon.test FAILED in 0.8s
```
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
The ParseTree comments say that preorder is "easier to visualize and read". The problem is, both the ParseTree and Semantics need to operate on the postorder traversal: the ParseTree during construction, and the Semantics during processing. As a consequence, understanding the postorder traversal is important, but it's also very hard to decipher when presented preorder. This PR provides a way to see the postorder, with helpful indents to show subtrees.
This retains the preorder printing as an option for people who prefer that. I'm pretty sure it'll be easier to debug tests if we can see the postorder, so I'm making that the default.
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Also does `if` support, discussed refactorings, like PopState/PushState instead of edits.
With these changes, I'm to where I can start talking about what's still missing:
```
//toolchain/parser/testdata:basics/fail_invalid_designators.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/fail_paren_match_regression.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:basics/function_call.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/package.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/structs.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/tuples.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:basics/var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_colon_instead_of_in.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/fail_missing_in.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/fail_missing_var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/nested.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/simple.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:function/definition/with_params.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:operators/associative.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fail_missing_precedence_and_or.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fail_missing_precedence_or_and.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fail_variety.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fixity.carbon.test FAILED in 0.9s
//toolchain/parser/testdata:operators/missing_precedence_not.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/postfix_unary.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:operators/prefix_unary.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:while/basic.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:while/fail_unbraced.carbon.test FAILED in 0.6s
```
This does modify a couple `if` tests to not test so much expression syntax -- that just seems like unrelated syntax.
At present, CHECK/FATAL print their own stack trace. This switches to just using std::abort for the stack trace, as well as the CHECK printing more completely.
This has a few consequences:
1) I'm now buffering the FATAL strings in order to print it later.
2) We now print the bug report message and program arguments on failure. This is part of pretty printing and was elided before.
3) We can now have pretty printing on FATAL, e.g. to show the stacks we're building in the parser.
The intent of this approach is to eliminate recursion limits as a barrier for the parser. While it may not be urgent to address, I want to avoid pouring effort into a parser approach that we don't think will be usable long-term.
Right now this is passing a minor set of tests. It's intended to be enough to show how I'm thinking about flow control for the parser. I'm manually switching back and forth because it seemed like the easiest approach that avoids duplicating tests.
This starts adding builtins with TypeType and IntegerLiteralType. Note, structurally that's all they are, and not directly accessible in any way.
Adds a type field to SemanticsNode. Now, IntegerLiteralType can be identified as having type=TypeType, and IntegerLiteral as type=IntegerLiteralType. The current iteration doesn't do anything for type propagation, because I wanted to avoid making this too big.
This also switches the Identifier IR to instead BindName, with some side-effects. I'd been trying to think how to provide a name for TypeType, and switching around how things worked seemed like a better approach. And while I think it's the right direction (e.g., alias should just be a BindName), I also realized I don't need to name TypeType: there's probably a keyword to refer to the builtin, so it shouldn't use regular name lookup.
As I'm looking at adding builtins, it feels a bit like the types of args will end up scaling close to the number of semantic kinds. Rather than having this lead to a bunch of macros, this approach drops the macros for plain code.
Note, this adds Get functions too -- I'd planned those regardless for type safety, this is just getting ahead of the issue so that the Print function is pretty clean.
As a step towards builtins, provide more blocks. The intent is that any significant scope change will become its own NodeBlock. Builtins should produce the first set of node blocks.
Note, SemanticsIR as set up here isn't handling ordering of import processing -- I haven't thought that through much beyond that we probably want some lighter-weight processing of the parse tree to achieve it. But I think the essence of loading builtins first as their own IR block is... probably right?
I'm thinking about how to handle multiple files, and I think the current IRFactory is useful as a file-focused thing. So shifting/renaming accordingly. (doing this in its own PR to make the history a little cleaner for git's move detection)