3 Commits
Author SHA1 Message Date
Chandler Carruth 632a81fcb6 Keep benchmark values passed to DoNotOptimize in registers (#7870)
Clang implements the `"+r,m"` constraint in `benchmark::DoNotOptimize`
by always choosing memory, so each call stores the value to the stack
and loads it back. Most of our benchmarks call it on a loop counter or
another value carried to the next iteration, which puts that store and
reload on the loop's critical path.

On an Apple M1, the cost of that round trip depends on which register
the compiler uses to address the stack slot, and unrelated code changes
move it. Changing only that register from `x29` to `sp`, with the same
address, made `BM_SetLookupHitPtr<Set<int>>` 41.5% to 49.1% slower.

Add `Carbon::Testing::DoNotOptimize` in
`//testing/base:benchmark_helpers`, and switch every benchmark to it. It
only accepts types it can keep in registers: integers, enums, and
pointers go into a `"+r"` constraint with a `"memory"` clobber, and
containers pass their `data()` pointer to that form, so the compiler
must assume the call reads and writes their contents. Any other type is
a compile error. The container form doesn't block optimizing based on a
container's size; passing the container's address does.

Benchmarks that pass a loop counter or carried value no longer measure a
store and reload per iteration, so their numbers aren't comparable with
earlier runs. For example, the hashing latency benchmarks no longer
include one in each hash's latency.

Assisted-by: Claude Code
2026-10-02 00:15:21 +00:00
Lucile Rose Nihlen 625f2ca629 precompile and cache Carbon prelude (#7432)
Refactors the link driver to automatically compile and cache the carbon
prelude for use in linking.

Implements a `carbon_library` rule for compiling the Core library
dependencies in the examples.
2026-07-16 18:02:14 +00:00
Chandler CarruthandChristopher Di Bella 361c832713 Add a prelude compilation benchmark (#7368)
Adds `toolchain/benchmarking/prelude_benchmark.cpp`, which measures the
time to compile the Core prelude and reports the most interesting SemIR
memory statistics as benchmark counters.

The prelude is compiled by checking a file that imports it: the implicit
prelude import causes the check phase to lex, parse, and check the full
set of prelude files, so this is a direct measure of prelude compilation
cost. Four input variations exercise increasing amounts of the prelude:
an empty file, a minimal single-type use, an operator-heavy file that
hits many impls, and a compact file that pulls in a wide swath of the
prelude.

Memory usage is queried directly: `Driver::set_mem_usage` takes a
`MemUsage` that a compile merges each file's usage into; the benchmark
passes one, compiles, and sums the entries by label. A compilation unit
collects into its own `MemUsage` whenever usage is dumped or a sink is
provided (decided in `SetMultiUnitCache`); after a file is done it dumps
that `MemUsage` per-file as before and, if a sink was provided, merges
into it via a new `MemUsage::Add(const MemUsage&)` overload. `MemUsage`
also exposes its entries via a public `Entry` type and an `entries()`
accessor.

Also extends `scripts/bench_runner.py` to (1) treat Mem-prefixed
counters as cost metrics (smaller is better) and (2) tolerate metrics
that aren't reported by every benchmark in a binary.

Assisted-by: Claude Code

---------

Co-authored-by: Christopher Di Bella <cjdb.ns@gmail.com>
2026-06-29 05:39:44 +00:00