I've been heading this route with other parts of the parser because the overhead of adding enums and then switching on them felt tedious, and odd from a performance perspective to make calls when the caller knew the value to use. My leaning is towards this approach that makes it clearer what's actually different between the modes, and allows removing BraceExpressionKindToParserState. It's a mild code size decrease.
Also cleans up some comments about related parse nodes. Currently basic and not heavily validated.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
This makes it possible to specify both deduced and regular parameters on types. It reorganizes the handling of parameter lists in order to allow more reuse of code in this approach. Both functions and types use the new DeclarationNameAndParams handling. Overall the goal here is to take advantage of commonality in structure.
Regarding destructors, the likely approach would be to use ParameterListAsDeduced directly because `destructor` is a keyword with no declaration name and no regular parameters.
We could similarly add others -- this is intended to make it easy to add more that parse essentially the same.
The functionality expected is that types will use GetDeclarationContext in order to error on certain functionality in the declaration scope loop. e.g., with how constraints and interfaces currently don't allow definitions.
I've only moved out `package` because it's only valid on the top line. It might still be good to parse it later, but with slightly different logic because it would always be an error, and the declaration context isn't quite the right framing for that.
Also unifies some errors with `fn`.
Depends on #2646
Right now, deduced parameter handling is very narrow to `self` support. This folds it into pattern handling, which should eventually be a superset of deduced parameter support, so this will avoid more duplication of logic.
Note, this subtly adds handling of multiple deduced parameters, but not generic parameters (`:!`) so it's still not quite right.
The comments in parse_node_kind.def capture the change being made here.
Before:
```
// _external_: DeclaredName
// InterfaceBodyStart
// _external_: statements
// InterfaceBodyEnd
// InterfaceDefinition
```
After:
```
// InterfaceIntroducer
// DeclaredName
// InterfaceDefinitionStart
// _external_: declarations
// InterfaceDefinition
```
Really I just want to treat introduced things consistently. `var` defines my philosophy here: it doesn't always have a `DeclaredName`, so the `VarIntroducer` _must_ be the bounding node. By being consistent with that, I believe that overall the structure becomes easier to understand (that is, there are fewer inconsistencies to understand).
This also adds InterfaceDeclaration, since I think it can be predicted we'll have that, and it's helpful for making recover consistent with HandleDeclarationError.
Similarly, I'm also trying to standardize the loop processing a little with HandleDeclarationLoop. In the current approach, InterfaceDefinitionFinish isn't a necessary state, so I'm removing it.
Per [discussion](https://discord.com/channels/655572317891461132/655578254970716160/1078427629427904563), there's a preference for having the support this enables in the parser. For example, that the parser should detect and error on a non-default interface function's definition.
However, we do need to handle nesting of declarations. We could do that by making this a stack. I think though that it'll be more efficient to keep using state_stack_, since it'll be called in places which are a limited number of steps from the actual state. That may already be in cache since we frequently look at state_stack_, so I'm uncertain that maintaining an additional stack would be a net benefit.
The goal here is to (significantly) reduce the boilerplate needed when defining classes that wrap enums, especially those managed with the `.def`-file style X-macros that are common in the toolchain.
This should also provide both better and more consistent functionality to those classes once ported over to it.
Initially, only `ParserState`, `SemanticsNodeKind`, and `SemanticsBuiltinKind` are ported as these were also the three that JonMeow ported in his original pull/2453 "option 5". This is heavily based on that version of the code.
Goals I was considering that influenced the design:
- Keep the individual enum-wrapping classes as simple and easy to read as possible. Especially important is keeping the `.def` files that are often filled with really important documentation clean and easy to maintain over time.
- Don't rely on computed `#include`s as that is an especially dark corner of the preprocessor and breaks some build systems.
- Have a really good API of the enum-wrapping class, including nice constant names for the values, easy printing, and even easy debugger-callable methods to get the name (as opposed to the integer value).
- Keep the API that users interact with in the base class as clean and easy to read as possible.
- Reduce the boiler plate for each instance of these as much as possible.
- Avoid excessive inline generated code or constants that would result in steady growth in object file sizes and linker effort doing deduplication.
These goals aren't always compatible, so we end up needing to pick a compromise between them when in tension. I think this version is a pretty good compromise.
The original version I started with already pull most of the API into a CRTP-style base class. This version pulls *all* of the common API. This is the main tool for getting consistency and avoiding duplication. However, connecting this base class to the individual enum wrappers is still difficult. Some specific changes here that try to do as much as possible there:
- Use a slightly fancier macro pattern to reduce the boilerplate of defining the raw `enum class` prior to the wrapper class.
- Use a macro to simplify naming the base class.
- Move the name table to a `.cpp` file to avoid every inclusion generating a complete copy of the strings (that the linker has to deduplicate). This is done with some care to sharply reduce the boilerplate needed in that `.cpp` file.
- Sink the name _API_ fully into the CRTP base class. This requires some significant complexity in the implementation, but all of that is hidden behind a single implementation detail macro, and the API itself is simple and readable. This also makes it much more reasonable to test the entire system a single time next to the base class.
This version also moves from constant factory functions to normal constants. This requires two batches -- first a declaration, and then a definition -- but the API result is significantly better and similar to the original option, the macro structure reduces the cost of these. Unfortunately that makes the adoption a bit noisy, but I think its worth the churn.
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
LLVM's bazel build has changed a bit, so this updates the tree for that.
LLVM is also moving `llvm::Optional` to match the standard API, but it seemed simpler to just switch to `std::optional`.
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
- Finishes remaining "todo" parse nodes.
- Improving error recovery for invalid designators and structs, so that the parse tree still looks similar to a valid parse tree.
- Call expressions now have the thing being called as a child (of the start) instead of a sibling.
- Use of Start is replacing use of End in several parse nodes, like structs and call expressions.
- Adjusting documentation of parse node structures in an attempt to make it more consistent and understandable.
- The current state for interfaces and if/else is mostly being documented, not altered.
This works on multiple statements to make them better for the bracketing model. Stub nodes are added in more cases of invalid syntax, simply so that the semantics has reliably structured input. Comments in parse_node_kind.def now try to show the expected parse tree structure in postorder form.
This labels If, While, and For a little differently in parse nodes so that at the start of the postorder traversal, it'll already be available to semantics which structure is being processed. I need to do a little more with If in particular, but this felt like a reasonable stopping point.
While this makes significant parser changes, the changes to parser_state.def are minimal, mostly naming-related. The actual flow isn't substantively changed, just a couple minor names and the new As(If|While) state which allows distinguishing IfCondition and WhileCondition.
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
This changes `return`, `break`, and `continue` to treat the keyword as the "start" and semicolon as the "parent", essentially bracketing the keyword.
Pragmatically this is focusing on making `return` work with only one ParseNodeKind: because `return` and `;` now bracket the expression, we can tightly determine whether the `return` has arguments without looking at subtree size. However, it's possible that `break` and `continue` may in the future take some kind of label as an argument, so the consistency seems beneficial there too.
Note this eliminates the StatementEnd ParseNodeKind, as it's obsolete with this change.
While the Parser has similar divergent states, lists of Expressions tend to be handled more like this. I'm keeping the divergent start state in order to continue support of a distinct error, but I think this organization of FunctionParameter/FunctionParameterFinish will be less surprising.