Fix uniform identifier generation for lengths over 64 (#7459)

`GetIdentifiersImpl` unconditionally sliced the 64-entry
`IdentifierLengthCounts` table even for uniform distributions, which the
API documents as having no `max_length` limit. Requesting a uniform
distribution with `max_length > 64` therefore tripped an out-of-bounds
slice assertion. Only compute the table slice on the non-uniform path,
where `max_length <= 64` is already enforced.

Add `IdentifierByteSumStableAcrossSeeds`, which exercises this path (a
uniform request up to length 200) and checks the core invariant that the
total identifier byte count is independent of the random seed across a
spread of parameters.

Assisted-by: Claude Code

---------

Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
This commit is contained in:
Chandler Carruth
2026-07-07 06:52:57 +00:00
committed by GitHub
co-authored by josh11b
parent a2890716ba
commit 383cfbb023
2 changed files with 93 additions and 4 deletions
+8 -4
View File
@@ -654,11 +654,15 @@ auto SourceGen::GetIdentifiersImpl(int number, int min_length, int max_length,
idents.reserve(number);
// First, compute the total weight of the distribution so we know how many
// identifiers we'll get each time we collect from it.
// identifiers we'll get each time we collect from it. For a uniform
// distribution every length has weight one, so the sum is simply the number
// of lengths; this also avoids indexing the bounded `IdentifierLengthCounts`
// table, which only covers lengths up to 64 and which uniform callers are
// allowed to exceed.
int num_lengths = max_length - min_length + 1;
auto length_counts =
llvm::ArrayRef(IdentifierLengthCounts).slice(min_length - 1, num_lengths);
int count_sum = uniform ? num_lengths : Sum(length_counts);
int count_sum = uniform ? num_lengths
: Sum(llvm::ArrayRef(IdentifierLengthCounts)
.slice(min_length - 1, num_lengths));
CARBON_CHECK(count_sum >= 1);
int number_rem = number % count_sum;