mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-10-05 22:02:55 +01:00
Clang implements the `"+r,m"` constraint in `benchmark::DoNotOptimize` by always choosing memory, so each call stores the value to the stack and loads it back. Most of our benchmarks call it on a loop counter or another value carried to the next iteration, which puts that store and reload on the loop's critical path. On an Apple M1, the cost of that round trip depends on which register the compiler uses to address the stack slot, and unrelated code changes move it. Changing only that register from `x29` to `sp`, with the same address, made `BM_SetLookupHitPtr<Set<int>>` 41.5% to 49.1% slower. Add `Carbon::Testing::DoNotOptimize` in `//testing/base:benchmark_helpers`, and switch every benchmark to it. It only accepts types it can keep in registers: integers, enums, and pointers go into a `"+r"` constraint with a `"memory"` clobber, and containers pass their `data()` pointer to that form, so the compiler must assume the call reads and writes their contents. Any other type is a compile error. The container form doesn't block optimizing based on a container's size; passing the container's address does. Benchmarks that pass a loop counter or carried value no longer measure a store and reload per iteration, so their numbers aren't comparable with earlier runs. For example, the hashing latency benchmarks no longer include one in each hash's latency. Assisted-by: Claude Code