Start reorienting the ParseTree towards a more efficient SemanticsIR production. (#2275)

In summary - some of the changes here are focused on producing the same test results for SemanticsIR, but I think the next step will be to change the SemanticsIR structure to reduce how much is added to the traversal stack.

Switching semantics to a postorder traversal is intended to be more efficient. The traversal stack is to eliminate risk of recursion limits within the semantic analysis that could come from layered code structures. However, we need to start considering the implications for type-checking and what the ParseTree looks like, as well as copying of data here.

As we start thinking about type-checking in SemanticsIR, it's helpful for a function to know its own signature in order to perform lookup recursive calls. The challenge in the post-order walk without this change is it doesn't know it's in a function definition (or similar) until it reaches the FunctionDeclaration; this restructures so that either:

1. For a declaration, the signature is a child of FunctionDeclaration(";")
2. For a definition, the signature is a child of FunctionDefinitionStart("{") which pairs with FunctionDefinition("}"), replacing CodeBlock.

This similarly reorients CodeBlock to be CodeBlockStart("{") as the first child of CodeBlock("}"). I'm not doing that with ParameterList here just because it affects a bit more, and felt like it could be delayed.

Overall, my goal is making the postorder traversal more intuitive along scope boundaries. I think we may also not need subtree_size, so I'm avoiding use of that now.

Currently the SemanticsIRFactory implementation is less clean than I might like (there are a couple comments to this point), but I was starting to feel like a more complete rewrite would be appropriate rather than trying to clean it up further: in particular, I think the node structures are off, but changing them is significant and also changes test output; in turn it may also warrant more substantial ParseTree changes. If you prefer from a reviewer POV, I can do a more complete rewrite.
This commit is contained in:
Jon Ross-Perkins
2022-10-18 16:15:57 -07:00
committed by GitHub
parent 4d522c8e90
commit 7b48ac7258
66 changed files with 1177 additions and 1364 deletions
+105 -92
View File
@@ -95,9 +95,9 @@ auto ParseTree::Parser::Parse(TokenizedBuffer& tokens,
TokenDiagnosticEmitter& emitter) -> ParseTree {
ParseTree tree(tokens);
// We expect to have a 1:1 correspondence between tokens and tree nodes, so
// reserve the space we expect to need here to avoid allocation and copying
// overhead.
// Reserve the space we expect to need for nodes in order to avoid allocation
// and copying overhead. This should be a one-to-one correspondence in an
// error-free tree.
tree.node_impls_.reserve(tokens.size());
Parser parser(tree, tokens, emitter);
@@ -111,6 +111,9 @@ auto ParseTree::Parser::Parse(TokenizedBuffer& tokens,
parser.AddLeafNode(ParseNodeKind::FileEnd(), *parser.position_);
CARBON_CHECK(tree.has_errors() || tree.size() == tokens.size())
<< "Failed to correctly calculate size: expected " << tokens.size()
<< ", got " << tree.size();
CARBON_CHECK(tree.Verify()) << "Parse tree built but does not verify!";
return tree;
}
@@ -135,9 +138,13 @@ auto ParseTree::Parser::ConsumeIf(TokenKind kind)
}
auto ParseTree::Parser::AddLeafNode(ParseNodeKind kind,
TokenizedBuffer::Token token) -> Node {
TokenizedBuffer::Token token,
bool has_error) -> Node {
Node n(tree_.node_impls_.size());
tree_.node_impls_.push_back(NodeImpl(kind, token, /*subtree_size_arg=*/1));
if (has_error) {
MarkNodeError(n);
}
return n;
}
@@ -229,9 +236,8 @@ auto ParseTree::Parser::FindNextOf(
}
}
auto ParseTree::Parser::SkipPastLikelyEnd(TokenizedBuffer::Token skip_root,
SemiHandler on_semi)
-> llvm::Optional<Node> {
auto ParseTree::Parser::SkipPastLikelyEnd(TokenizedBuffer::Token skip_root)
-> llvm::Optional<TokenizedBuffer::Token> {
if (AtEndOfFile()) {
return llvm::None;
}
@@ -261,7 +267,7 @@ auto ParseTree::Parser::SkipPastLikelyEnd(TokenizedBuffer::Token skip_root,
// We assume that a semicolon is always intended to be the end of the
// current construct.
if (auto semi = ConsumeIf(TokenKind::Semi())) {
return on_semi(*semi);
return semi;
}
// Skip over any matching group of tokens_.
@@ -408,44 +414,53 @@ auto ParseTree::Parser::ParseFunctionSignature() -> bool {
}
}
return params.hasValue();
return params.has_value();
}
auto ParseTree::Parser::ParseCodeBlock() -> llvm::Optional<Node> {
return ParseCodeBlock(GetSubtreeStartPosition(),
ParseNodeKind::CodeBlockStart(),
ParseNodeKind::CodeBlock());
}
auto ParseTree::Parser::ParseCodeBlock(SubtreeStart subtree_start,
ParseNodeKind start_kind,
ParseNodeKind end_kind)
-> llvm::Optional<Node> {
CARBON_RETURN_IF_STACK_LIMITED(llvm::None);
llvm::Optional<TokenizedBuffer::Token> maybe_open_curly =
llvm::Optional<TokenizedBuffer::Token> open_curly =
ConsumeIf(TokenKind::OpenCurlyBrace());
if (!maybe_open_curly) {
if (!open_curly) {
// Recover by parsing a single statement.
CARBON_DIAGNOSTIC(ExpectedCodeBlock, Error, "Expected braced code block.");
emitter_.Emit(*position_, ExpectedCodeBlock);
return ParseStatement();
// Use the unexpected token for the block start and end.
TokenizedBuffer::Token recovery_start = *position_;
AddNode(start_kind, recovery_start, subtree_start, /*has_error=*/true);
ParseStatement();
return AddNode(end_kind, recovery_start, subtree_start, /*has_error=*/true);
}
TokenizedBuffer::Token open_curly = *maybe_open_curly;
auto start = GetSubtreeStartPosition();
bool has_errors = false;
AddNode(start_kind, *open_curly, subtree_start);
// Loop over all the different possibly nested elements in the code block.
bool has_error = false;
while (!NextTokenIs(TokenKind::CloseCurlyBrace())) {
if (!ParseStatement()) {
// We detected and diagnosed an error of some kind. We can trivially skip
// to the actual close curly brace from here.
// We detected and diagnosed an error of some kind. We can trivially
// skip to the actual close curly brace from here.
// TODO: It would be better to skip to the next semicolon, or the next
// token at the start of a line with the same indent as this one.
SkipTo(tokens_.GetMatchedClosingToken(open_curly));
has_errors = true;
SkipTo(tokens_.GetMatchedClosingToken(*open_curly));
has_error = true;
break;
}
}
// We always reach here having set our position in the token stream to the
// close curly brace.
AddLeafNode(ParseNodeKind::CodeBlockEnd(),
Consume(TokenKind::CloseCurlyBrace()));
return AddNode(ParseNodeKind::CodeBlock(), open_curly, start, has_errors);
return AddNode(end_kind, Consume(TokenKind::CloseCurlyBrace()), subtree_start,
/*has_error=*/has_error);
}
auto ParseTree::Parser::ParsePackageDirective() -> Node {
@@ -460,9 +475,9 @@ auto ParseTree::Parser::ParsePackageDirective() -> Node {
CARBON_RETURN_IF_STACK_LIMITED(create_error_node());
auto exit_on_parse_error = [&]() {
SkipPastLikelyEnd(package_intro_token, [&](TokenizedBuffer::Token semi) {
return AddLeafNode(ParseNodeKind::PackageEnd(), semi);
});
if (auto semi_token = SkipPastLikelyEnd(package_intro_token)) {
AddLeafNode(ParseNodeKind::PackageEnd(), *semi_token);
}
return create_error_node();
};
@@ -530,30 +545,34 @@ auto ParseTree::Parser::ParsePackageDirective() -> Node {
}
auto ParseTree::Parser::ParseFunctionDeclaration() -> Node {
TokenizedBuffer::Token function_intro_token = Consume(TokenKind::Fn());
auto start = GetSubtreeStartPosition();
TokenizedBuffer::Token function_intro_token = Consume(TokenKind::Fn());
AddLeafNode(ParseNodeKind::FunctionIntroducer(), function_intro_token);
auto add_error_function_node = [&] {
// When handling errors before the start of the definition, treat it as a
// declaration. Recover to a semicolon when it makes sense as a possible
// function end, otherwise use the fn token for the error.
auto add_error_function_node = [&](bool skip_past_likely_end) {
if (skip_past_likely_end) {
if (auto semi_token = SkipPastLikelyEnd(function_intro_token)) {
return AddNode(ParseNodeKind::FunctionDeclaration(), *semi_token, start,
/*has_error=*/true);
}
}
return AddNode(ParseNodeKind::FunctionDeclaration(), function_intro_token,
start, /*has_error=*/true);
};
CARBON_RETURN_IF_STACK_LIMITED(add_error_function_node());
CARBON_RETURN_IF_STACK_LIMITED(add_error_function_node(false));
auto handle_semi_in_error_recovery = [&](TokenizedBuffer::Token semi) {
return AddLeafNode(ParseNodeKind::DeclarationEnd(), semi);
};
auto name_n = ConsumeAndAddLeafNodeIf(TokenKind::Identifier(),
ParseNodeKind::DeclaredName());
if (!name_n) {
if (!ConsumeAndAddLeafNodeIf(TokenKind::Identifier(),
ParseNodeKind::DeclaredName())) {
CARBON_DIAGNOSTIC(ExpectedFunctionName, Error,
"Expected function name after `fn` keyword.");
emitter_.Emit(*position_, ExpectedFunctionName);
// TODO: We could change the lexer to allow us to synthesize certain
// kinds of tokens and try to "recover" here, but unclear that this is
// really useful.
SkipPastLikelyEnd(function_intro_token, handle_semi_in_error_recovery);
return add_error_function_node();
return add_error_function_node(true);
}
TokenizedBuffer::Token open_paren = *position_;
@@ -561,40 +580,39 @@ auto ParseTree::Parser::ParseFunctionDeclaration() -> Node {
CARBON_DIAGNOSTIC(ExpectedFunctionParams, Error,
"Expected `(` after function name.");
emitter_.Emit(open_paren, ExpectedFunctionParams);
SkipPastLikelyEnd(function_intro_token, handle_semi_in_error_recovery);
return add_error_function_node();
return add_error_function_node(true);
}
TokenizedBuffer::Token close_paren =
tokens_.GetMatchedClosingToken(open_paren);
if (!ParseFunctionSignature()) {
// Don't try to parse more of the function declaration, but consume a
// declaration ending semicolon if found (without going to a new line).
SkipPastLikelyEnd(function_intro_token, handle_semi_in_error_recovery);
return add_error_function_node();
return add_error_function_node(true);
}
// See if we should parse a definition which is represented as a code block.
if (NextTokenIs(TokenKind::OpenCurlyBrace())) {
if (!ParseCodeBlock()) {
return add_error_function_node();
switch (NextTokenKind()) {
case TokenKind::OpenCurlyBrace(): {
// Parse a definition which is represented as a code block.
if (auto node =
ParseCodeBlock(start, ParseNodeKind::FunctionDefinitionStart(),
ParseNodeKind::FunctionDefinition())) {
return *node;
}
return add_error_function_node(false);
}
} else if (!ConsumeAndAddLeafNodeIf(TokenKind::Semi(),
ParseNodeKind::DeclarationEnd())) {
CARBON_DIAGNOSTIC(
ExpectedFunctionBodyOrSemi, Error,
"Expected function definition or `;` after function declaration.");
emitter_.Emit(*position_, ExpectedFunctionBodyOrSemi);
if (tokens_.GetLine(*position_) == tokens_.GetLine(close_paren)) {
case TokenKind::Semi(): {
return AddNode(ParseNodeKind::FunctionDeclaration(),
Consume(TokenKind::Semi()), start);
}
default: {
CARBON_DIAGNOSTIC(
ExpectedFunctionBodyOrSemi, Error,
"Expected function definition or `;` after function declaration.");
emitter_.Emit(*position_, ExpectedFunctionBodyOrSemi);
// Only need to skip if we've not already found a new line.
SkipPastLikelyEnd(function_intro_token, handle_semi_in_error_recovery);
return add_error_function_node(tokens_.GetLine(*position_) ==
tokens_.GetLine(close_paren));
}
return add_error_function_node();
}
// Successfully parsed the function, add that node.
return AddNode(ParseNodeKind::FunctionDeclaration(), function_intro_token,
start);
}
auto ParseTree::Parser::ParseVariableDeclaration() -> Node {
@@ -625,9 +643,10 @@ auto ParseTree::Parser::ParseVariableDeclaration() -> Node {
ParseNodeKind::DeclarationEnd());
if (!semi) {
emitter_.Emit(*position_, ExpectedSemiAfterExpression);
SkipPastLikelyEnd(var_token, [&](TokenizedBuffer::Token semi) {
return AddLeafNode(ParseNodeKind::DeclarationEnd(), semi);
});
if (auto semi_token = SkipPastLikelyEnd(var_token)) {
semi = AddLeafNode(ParseNodeKind::DeclarationEnd(), *semi_token,
/*has_error=*/true);
}
}
return AddNode(ParseNodeKind::VariableDeclaration(), var_token, start,
@@ -665,12 +684,9 @@ auto ParseTree::Parser::ParseDeclaration() -> llvm::Optional<Node> {
// Skip forward past any end of a declaration we simply didn't understand so
// that we can find the start of the next declaration or the end of a scope.
if (auto found_semi_n =
SkipPastLikelyEnd(*position_, [&](TokenizedBuffer::Token semi) {
return AddLeafNode(ParseNodeKind::EmptyDeclaration(), semi);
})) {
MarkNodeError(*found_semi_n);
return *found_semi_n;
if (auto semi_token = SkipPastLikelyEnd(*position_)) {
return AddLeafNode(ParseNodeKind::EmptyDeclaration(), *semi_token,
/*has_error=*/true);
}
// Nothing, not even a semicolon found.
@@ -759,8 +775,8 @@ auto ParseTree::Parser::ParseBraceExpression() -> llvm::Optional<Node> {
}
kind = elem_kind;
// Struct type fields and value fields use the same grammar except that
// one has a `:` separator and the other has an `=` separator.
// Struct type fields and value fields use the same grammar except
// that one has a `:` separator and the other has an `=` separator.
auto equal_or_colon_token =
Consume(kind == Type ? TokenKind::Colon() : TokenKind::Equal());
auto type_or_value = ParseExpression();
@@ -877,8 +893,8 @@ auto ParseTree::Parser::ParsePostfixExpression() -> llvm::Optional<Node> {
default:
return expression;
}
// This is subject to an infinite loop if a child call fails, so monitor for
// stalling.
// This is subject to an infinite loop if a child call fails, so monitor
// for stalling.
if (last_position == position_) {
CARBON_CHECK(expression == llvm::None);
return expression;
@@ -895,8 +911,8 @@ static auto IsAssumedStartOfOperand(TokenKind kind) -> bool {
TokenKind::StringLiteral()});
}
// Determines whether the given token is considered to be the end of an operand
// according to the rules for infix operator parsing.
// Determines whether the given token is considered to be the end of an
// operand according to the rules for infix operator parsing.
static auto IsAssumedEndOfOperand(TokenKind kind) -> bool {
return kind.IsOneOf({TokenKind::CloseParen(), TokenKind::CloseCurlyBrace(),
TokenKind::CloseSquareBracket(), TokenKind::Identifier(),
@@ -904,9 +920,9 @@ static auto IsAssumedEndOfOperand(TokenKind kind) -> bool {
TokenKind::StringLiteral()});
}
// Determines whether the given token could possibly be the start of an operand.
// This is conservatively correct, and will never incorrectly return `false`,
// but can incorrectly return `true`.
// Determines whether the given token could possibly be the start of an
// operand. This is conservatively correct, and will never incorrectly return
// `false`, but can incorrectly return `true`.
static auto IsPossibleStartOfOperand(TokenKind kind) -> bool {
return !kind.IsOneOf({TokenKind::CloseParen(), TokenKind::CloseCurlyBrace(),
TokenKind::CloseSquareBracket(), TokenKind::Comma(),
@@ -1113,12 +1129,9 @@ auto ParseTree::Parser::ParseExpressionStatement() -> llvm::Optional<Node> {
emitter_.Emit(*position_, ExpectedSemiAfterExpression);
}
if (auto recovery_node =
SkipPastLikelyEnd(start_token, [&](TokenizedBuffer::Token semi) {
return AddNode(ParseNodeKind::ExpressionStatement(), semi, start,
true);
})) {
return recovery_node;
if (auto semi_token = SkipPastLikelyEnd(start_token)) {
return AddNode(ParseNodeKind::ExpressionStatement(), *semi_token, start,
/*has_error=*/true);
}
// Found junk not even followed by a `;`.
@@ -1196,11 +1209,11 @@ auto ParseTree::Parser::ParseForStatement() -> llvm::Optional<Node> {
"not implemented yet!",
TokenKind);
emitter_.Emit(*position_, ExpectedParenAfter, TokenKind::For());
// TODO: A proper recovery strategy is needed here. For now, I assume that
// all brackets are properly balanced (i.e. each open bracket has a
// TODO: A proper recovery strategy is needed here. For now, I assume
// that all brackets are properly balanced (i.e. each open bracket has a
// closing one).
// This is temporary until we come to a conclusion regarding the recovery
// tokens strategy.
// This is temporary until we come to a conclusion regarding the
// recovery tokens strategy.
return llvm::None;
}
@@ -1224,8 +1237,8 @@ auto ParseTree::Parser::ParseForStatement() -> llvm::Optional<Node> {
}
// A separator is either an `in` or a `:`. Even though `:` is incorrect,
// accidentally typing it by a C++ programmer might be a common mistake that
// warrants special handling.
// accidentally typing it by a C++ programmer might be a common mistake
// that warrants special handling.
bool separator_parsed = false;
bool in_parsed = false;