Adds remaining expression support, and switches the default to Parser2.
Note, this doesn't delete the current Parser yet. I'll just do that in its own PR.
Adds `while` and most of the expression support. Splits apart the fixity test to demonstrate more closely which bits are still failing.
```
//toolchain/parser/testdata:basics/fail_paren_match_regression.carbon.test FAILED in 0.8s
//toolchain/parser/testdata:basics/function_call.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/package.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/structs.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/tuples.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:basics/var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_colon_instead_of_in.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_missing_in.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_missing_var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/nested.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/simple.carbon.test FAILED in 0.8s
//toolchain/parser/testdata:function/definition/with_params.carbon.test FAILED in 0.8s
//toolchain/parser/testdata:operators/fixity_in_call.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:operators/fixity_in_var.carbon.test FAILED in 0.8s
```
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
The ParseTree comments say that preorder is "easier to visualize and read". The problem is, both the ParseTree and Semantics need to operate on the postorder traversal: the ParseTree during construction, and the Semantics during processing. As a consequence, understanding the postorder traversal is important, but it's also very hard to decipher when presented preorder. This PR provides a way to see the postorder, with helpful indents to show subtrees.
This retains the preorder printing as an option for people who prefer that. I'm pretty sure it'll be easier to debug tests if we can see the postorder, so I'm making that the default.
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
Also does `if` support, discussed refactorings, like PopState/PushState instead of edits.
With these changes, I'm to where I can start talking about what's still missing:
```
//toolchain/parser/testdata:basics/fail_invalid_designators.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/fail_paren_match_regression.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:basics/function_call.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/package.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/structs.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:basics/tuples.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:basics/var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/fail_colon_instead_of_in.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/fail_missing_in.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/fail_missing_var.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:for/nested.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:for/simple.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:function/definition/with_params.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:operators/associative.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fail_missing_precedence_and_or.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fail_missing_precedence_or_and.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fail_variety.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/fixity.carbon.test FAILED in 0.9s
//toolchain/parser/testdata:operators/missing_precedence_not.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:operators/postfix_unary.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:operators/prefix_unary.carbon.test FAILED in 0.7s
//toolchain/parser/testdata:while/basic.carbon.test FAILED in 0.6s
//toolchain/parser/testdata:while/fail_unbraced.carbon.test FAILED in 0.6s
```
This does modify a couple `if` tests to not test so much expression syntax -- that just seems like unrelated syntax.
At present, CHECK/FATAL print their own stack trace. This switches to just using std::abort for the stack trace, as well as the CHECK printing more completely.
This has a few consequences:
1) I'm now buffering the FATAL strings in order to print it later.
2) We now print the bug report message and program arguments on failure. This is part of pretty printing and was elided before.
3) We can now have pretty printing on FATAL, e.g. to show the stacks we're building in the parser.
The intent of this approach is to eliminate recursion limits as a barrier for the parser. While it may not be urgent to address, I want to avoid pouring effort into a parser approach that we don't think will be usable long-term.
Right now this is passing a minor set of tests. It's intended to be enough to show how I'm thinking about flow control for the parser. I'm manually switching back and forth because it seemed like the easiest approach that avoids duplicating tests.
This rewrites semantics towards a more pure instruction model, in pursuit of the simple instruction-style output.
I think I can get this approach to type-check as it goes along, but obviously this change doesn't prove that yet. I'm separating it out because it's a large rewrite of the semantics structure, tossing out a lot of what was there before. But I think it does help towards several requests, like setting up a clear path for consolidating duplicate identifiers and making the node style more standardized.
I expect to need to pass multiple args to function calls, that'd probably be storing vectors of args similar to how I'm showing identifiers and integer literals stored.
This removes the semantics namespace because (a) it was getting annoying writing the `::` everywhere, and (b) I think the leaning with Carbon is to avoid namespaces (@chandlerc asked not to put SemanticsIR/SemanticsFactory in a namespace, which is the crux of the issue). But, it's still necessary to avoid name conflicts so I just prefix everything with "Semantics" (still a lot of typing, but no `::`).
In summary - some of the changes here are focused on producing the same test results for SemanticsIR, but I think the next step will be to change the SemanticsIR structure to reduce how much is added to the traversal stack.
Switching semantics to a postorder traversal is intended to be more efficient. The traversal stack is to eliminate risk of recursion limits within the semantic analysis that could come from layered code structures. However, we need to start considering the implications for type-checking and what the ParseTree looks like, as well as copying of data here.
As we start thinking about type-checking in SemanticsIR, it's helpful for a function to know its own signature in order to perform lookup recursive calls. The challenge in the post-order walk without this change is it doesn't know it's in a function definition (or similar) until it reaches the FunctionDeclaration; this restructures so that either:
1. For a declaration, the signature is a child of FunctionDeclaration(";")
2. For a definition, the signature is a child of FunctionDefinitionStart("{") which pairs with FunctionDefinition("}"), replacing CodeBlock.
This similarly reorients CodeBlock to be CodeBlockStart("{") as the first child of CodeBlock("}"). I'm not doing that with ParameterList here just because it affects a bit more, and felt like it could be delayed.
Overall, my goal is making the postorder traversal more intuitive along scope boundaries. I think we may also not need subtree_size, so I'm avoiding use of that now.
Currently the SemanticsIRFactory implementation is less clean than I might like (there are a couple comments to this point), but I was starting to feel like a more complete rewrite would be appropriate rather than trying to clean it up further: in particular, I think the node structures are off, but changing them is significant and also changes test output; in turn it may also warrant more substantial ParseTree changes. If you prefer from a reviewer POV, I can do a more complete rewrite.
This switch is being done in order to make it easier to update tests when the parse tree structure changes. I've been spending a lot of time doing such updates as part of refactoring the parse tree, and I expect more, so that's really at the root of all the refactoring I've been doing to lit updates.
Summary:
Adds parsing support for the `package` directive as specified by the
`Code and name organization` design doc.
Co-authored-by: ergawy <kareem.ergawy@guardsquare.com>
Working on toolchain semantics:
- SemanticsIR is set up as a container for the semantic tree.
- SemanticsIRFactory builds the tree, with separate transformations for each ParseNodeKind.
- ParseSubtreeConsumer is a helper for transforming a ParseTree::Node's children, managing size/nodes to prevent errors.
- The nodes subdirectory contains SemanticIR nodes.
- MetaNode is used to represent nodes which have "sub-classes": Statements, Declarations, and Expressions.
- MetaNodeBlock is used to represent nodes which exist together in a block with name lookup: Statements and Declarations (not Expressions).
This is traversing children first in order to address the RPO format of ParseTree. This means that when lists are formed, they're reversed to be in code-order (`FixReverseOrdering`).
This is still very much incomplete -- the main intent at present is to demonstrate structure.
I'm doing this to avoid macro name conflicts, following https://google.github.io/styleguide/cppguide.html#Preprocessor_Macros: "If you do export a macro from a header, it must have a globally unique name. To achieve this, it must be named with a prefix consisting of your project's namespace name (but upper case)."
Commands run:
```
sed -i 's/\(DCHECK\|CHECK\|FATAL\|MAKE_UNIQUE_NAME\|MAKE_UNIQUE_NAME_IMPL\|RAW_EXITING_STREAM\|RETURN_IF_ERROR\|RETURN_IF_ERROR_IMPL\|ASSIGN_OR_RETURN\|ASSIGN_OR_RETURN_IMPL\|DIAGNOSTIC_KIND\|RETURN_IF_STACK_LIMITED\)(/CARBON_\1(/g' $(git ls-files *.cpp *.h *.lpp *.ypp *.def ':!third_party')
sed -i 's/#undef DIAGNOSTIC_KIND/#undef CARBON_DIAGNOSTIC_KIND/' toolchain/diagnostics/diagnostic_registry.def
```
Note this isn't *quite* everything, but it's intended to be a large pass at everything:
```
╚╡git grep '#define ' *.cpp *.h *.lpp *.ypp *.def ':!third_party' | grep -v '#define CARBON' | grep -v _H_
explorer/syntax/lexer.lpp: #define YY_USER_ACTION \
explorer/syntax/lexer.lpp: #define SIMPLE_TOKEN(name) \
explorer/syntax/lexer.lpp: #define ARG_TOKEN(name, arg) \
explorer/syntax/parse_and_lex_context.h:#define YY_DECL \
migrate_cpp/cpp_refactoring/var_decl.cpp:#define ABSTRACT_TYPE(Class, Base)
migrate_cpp/cpp_refactoring/var_decl.cpp:#define TYPE(Class, Base) \
```
We may in particular want to do a pass to clean up #ifdef guards and make them be CARBON_ rooted.
There's currently a bug with empty files, in that it initializes SourceBuffer with an invalid StringRef that results in a crash. That got me looking at the std::optional TODO, but the issue is that there are really three states:
- Buffered
- mmapped (not buffered)
- Moved out of (no longer initialized)
Technically an optional could work if we initialize the buffer on move out, indicating the mmap is gone. But the mode setup felt better to me.
And then this also adds the size check. Which is really how I started looking at this.
I was mainly looking at keywords trying to figure out what needs work and the amount to which it doesn't reflect the design confused me (including some things that we've decided not to include, and some things I'm not aware of discussion about). I figured this cleanup would at least make it somewhat clearer why things are in there.
I'm treating https://github.com/carbon-language/carbon-lang/blob/trunk/docs/design/lexical_conventions/words.md as canonical, with `_` and `xor` as presumably deliberate exceptions. Similarly avoiding symbol tokens because I assume you'll push proposals for the difference.
I dropped the `Keyword` qualifier because `is` makes `IsKeyword` a name conflict, and dropping the qualifier seemed like the more consistent solution (it doesn't do `AmpSymbol`, after all). If we need clarity I might lean towards a separate namespace to avoid naming conflicts.
The small size is for the 1m vs 5m time limit -- all these tests _should_ be fast so a lower limit seems consistent, and the 5m timeout was getting in my way when trying to debug *actual* timeouts.
The Carbon::Testing bit is for convenience -- test libraries are generally using it, it seems like the tests should too. Note this reduces the need for `using`.
This does push NodeMatchers into Carbon::Testing -- I don't think this was benefiting from having its own namespace; `using namespace` is discouraged [under style](https://google.github.io/styleguide/cppguide.html#Namespaces), we wouldn't support an equivalent in Carbon, and it feels like it's not helping to avoid name collisions. (also tidy was bugging about it, and while I could NOLINT that, this felt like the better approach)