Issue a diagnostic if we try to parse a source file that is too large. (#4429)

Previously in an optimized build we'd produce bogus tokens, such as
tokens with incorrect IdentifierIds, and in a debug build we would try
to CHECK-fail -- but actually wouldn't, because we're incorrectly
checking for `2 << bits` instead of `1 << bits`. I hit this while I was
trying to do some profiling and was seeing some very strange
diagnostics.

The diagnostic is pointed at the first token that is beyond the limit to
help people determine where to split their files.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
This commit is contained in:
Richard Smith
2024-10-22 23:21:35 +00:00
committed by GitHub
co-authored by Jon Ross-Perkins
parent af816cda90
commit e68e54dae4
5 changed files with 60 additions and 7 deletions
+15
View File
@@ -1107,6 +1107,21 @@ TEST_F(LexerTest, DiagnosticUnrecognizedChar) {
compile_helper_.GetTokenizedBuffer("\b", &consumer);
}
TEST_F(LexerTest, DiagnosticFileTooLarge) {
Testing::MockDiagnosticConsumer consumer;
static constexpr size_t NumLines = 10'000'000;
std::string input;
input.reserve(NumLines * 3);
for ([[maybe_unused]] int _ : llvm::seq(NumLines)) {
input += "{}\n";
}
EXPECT_CALL(consumer,
HandleDiagnostic(IsSingleDiagnostic(
DiagnosticKind::TooManyTokens, DiagnosticLevel::Error,
TokenizedBuffer::MaxTokens / 2, 1, _)));
compile_helper_.GetTokenizedBuffer(input, &consumer);
}
// Appends comment lines to the string, to create a comment block.
static auto AppendCommentLines(std::string& str, int count, llvm::StringRef tag)
-> void {