Trailing comments (#7441)

Carbon currently requires a comment to be the only non-whitespace on its
line. A `//` comment that follows other content on a line, called a
_trailing comment_, is a lexer error. This proposal removes that
restriction, allowing a comment to follow other content on a line.
Everything else about comments is unchanged: a comment still begins with
`//`, still requires whitespace after the `//`, and still runs to the
end of the line. Carbon continues to provide only line comments; no
block or intra-line comments are added.

Three observations motivate the change. First, trailing comments are
well suited to short _annotations_ attached to a specific entity or
value on a line. Second, the lexer design now makes it trivial to lex
trailing comments, and in fact requires extra logic and potentially cost
to reject them. Third, C++ code routinely uses trailing comments, so
allowing them lets Carbon carry the layout of migrated code over
directly, rather than reworking each comment to read well in a different
structure.

Implementation notes (beyond the proposal's design):

Keeping trailing comments cheap to lex required a few supporting
changes, all of which keep the cost off the lexer's hot path:

- The lexer already dispatches `//` to comment lexing wherever it
appears, so classifying a comment as trailing is a single O(1) check of
whether the `//` is the line's first non-whitespace (`start + indent`).
The hot comment path is otherwise unchanged.

- That check relies on each line's recorded indentation being its real
leading whitespace. Multi-line string literals previously recorded the
column where the literal opened for the lines they span; they now record
the true (closing-delimiter) indentation instead.

- Parser error recovery (`SkipPastLikelyEnd`) had relied on that
opening-column indentation to keep tokens following a multi-line string
literal attached to the same construct. It now reconstructs that
relationship directly by consulting the line on which the literal
opened, including when other tokens follow the closing delimiter (such
as `''' + "more"`). This is on the cold recovery path.

- `CommentData` records the trailing bit in the high bit of its length
field, keeping it at 8 bytes.

Assisted-by: Claude Code
This commit is contained in:
Chandler Carruth
2026-07-04 06:42:02 +00:00
committed by GitHub
parent 0460f6b7ba
commit f0848b1f5e
26 changed files with 1040 additions and 193 deletions
+12 -13
View File
@@ -145,18 +145,18 @@ auto Context::SkipPastLikelyEnd(Lex::TokenIndex skip_root) -> Lex::TokenIndex {
return *(position_ - 1);
}
Lex::LineIndex root_line = tokens().GetLine(skip_root);
int root_line_indent = tokens().GetIndentColumnNumber(root_line);
int root_line_indent =
tokens().GetIndentColumnNumber(tokens().GetLine(skip_root));
// We will keep scanning through tokens on the same line as the root or
// lines with greater indentation than root's line.
auto is_same_line_or_indent_greater_than_root = [&](Lex::TokenIndex t) {
Lex::LineIndex l = tokens().GetLine(t);
if (l == root_line) {
return true;
}
return tokens().GetIndentColumnNumber(l) > root_line_indent;
// We keep scanning through tokens that don't start their own line and
// through lines indented more than the root's line. Tokens that don't start
// their line include the rest of the root's line and, because comments run
// to the end of the line, tokens after a multi-line string literal's
// closing delimiter (as in `''' + "more"`), both of which continue the
// construct regardless of indentation.
auto keep_scanning = [&](Lex::TokenIndex t) {
int indent = tokens().GetIndentColumnNumber(tokens().GetLine(t));
return tokens().GetColumnNumber(t) > indent || indent > root_line_indent;
};
do {
@@ -179,8 +179,7 @@ auto Context::SkipPastLikelyEnd(Lex::TokenIndex skip_root) -> Lex::TokenIndex {
// Otherwise just step forward one token.
++position_;
} while (position_ != end_ &&
is_same_line_or_indent_greater_than_root(*position_));
} while (position_ != end_ && keep_scanning(*position_));
return *(position_ - 1);
}