Stitch together adjacent comments using the indent. (#4397)

This is improving the comment production to produce fewer distinct
comments.

At present, comment processing uses strict prefix matching. It either
expects `// ` (with a space) for valid comments, or just `//` (without a
space) for invalid comments that lacked the space.

As a consequence, the following would be three comments:

```
// Comment 1
//
//
// Comment 4
```

This is because a 3-character prefix is used for valid comments. The
prefix switches between lines 1 and 2, and again between lines 3 and 4,
each resulting in a separate comment.

For contrast, this is one comment because only a 2-character prefix is
used:

```
//Comment 1
//
//
//Comment 4
```

That's because all lines lack a suffix space.

Additionally, with SIMD 16-byte boundaries, further splits can occur if
processing needs to transition to non-SIMD.

Here, I'm trying to just address all of this by:

1. Stitching together adjacent comments. Since a lexed comment starts at
the `//` excluding the indent, the delta from the prior comment must be
precisely the indent.
2. Adding support for switching from SIMD to non-SIMD on file
boundaries.

I considered trying to have a separate `//\n` prefix for SIMD processing
of `// `, but I wasn't sure about the tradeoff of doing both at the same
time (in particular, it'd require constructing a string for the
different prefix), thus this stitch approach. This does mean multiple
passes will be required for a typical long comment structure using blank
comment lines to separate paragraphs (for performance reasons, I will
recommend engineers not write comm... nevermind).

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
This commit is contained in:
Jon Ross-Perkins
2024-10-10 22:25:41 +00:00
committed by GitHub
co-authored by Chandler Carruth
parent 9e5e33082c
commit 0db96ebc52
5 changed files with 130 additions and 12 deletions
+12
View File
@@ -357,6 +357,18 @@ auto TokenizedBuffer::GetCommentText(CommentIndex comment_index) const
return source_->text().substr(comment_data.start, comment_data.length);
}
auto TokenizedBuffer::AddComment(int32_t indent, int32_t start, int32_t end)
-> void {
if (!comments_.empty()) {
auto& comment = comments_.back();
if (comment.start + comment.length + indent == start) {
comment.length = end - comment.start;
return;
}
}
comments_.push_back({.start = start, .length = end - start});
}
auto TokenizedBuffer::CollectMemUsage(MemUsage& mem_usage,
llvm::StringRef label) const -> void {
mem_usage.Add(MemUsage::ConcatLabel(label, "allocator_"), allocator_);