Commit Graph
75 Commits
Author SHA1 Message Date
Jon Ross-Perkins 53c257d2e2 Switch libpfm and boost.unordered to BCR versions (#6847)
Assisted-by: Google Antigravity with Gemini
2026-03-06 21:57:23 +00:00
Chandler Carruth b2ab53e49c Fix an incompatiblitiy between our YAML and ErrorOr test helpers (#6636)
The YAML test helpers didn't use the `Printable` abstraction in one
place and instead directly used `<<` with a `std::ostream`. This matches
the `require`s expression in the `error_test_helpers.h` printing logic
for `ErrorOr`, but fails to provide the necessary implementation for
`llvm::formatv` to succeed with the `Yaml::Value` type.

The main fix is to use `Printable` and to define the `Print` method in
terms of `llvm::raw_ostream`. We already have all the mapping hooks in
place to also support `std::ostream` when needed based on that
definition.

This also adds some constraints to the printing in
`error_test_helpers.h` so it is a bit less under-constrained and more
understandable when it is correctly being used. These are just tidying
though, they aren't what makes these headers work together.

I've added a test to try and make sure these test helpers compose as
well.
2026-01-21 17:41:46 +00:00
Chandler Carruth e545929386 Pivot towards relative paths for installs and runtimes (#6547)
When building in Bazel actions, notably building runtimes, using
absolute paths makes the results non-hermetic and generally less
cache-friendly.

This restructures the code to only form an absolute path as part of the
`bazel run` change of working directory. It also tries to make the API
for doing this a bit more clear by taking the `exe_path` and
transforming it internally.

To support this, this PR also generalizes the `RemovingDir` to support
relative paths. While these can be tricky -- the working directory needs
to not change while they exist -- that isn't a reason to fully exclude
them and they're useful for implementing relative-path runtimes, etc.
2026-01-01 22:20:40 +00:00
Jon Ross-Perkins 6b775b3014 Switch benchmark tests to dry_run from min_time (#6433)
Trying to work around failures such as
https://github.com/carbon-language/carbon-lang/actions/runs/19680163053/job/56371861907...
min_time is only setting the minimum number of iterations, so the
benchmark framework is validly choosing to run 1k times. dry_run should
only run 1 repetition, making this both faster and more reliable in
terms of execution time.
https://google.github.io/benchmark/user_guide.html#running-benchmarks
for flag documentation.
2025-11-25 20:15:05 +00:00
Chandler CarruthandDana Jansens 4024d300bc Add a more friendly "latch" synchronization tool (#6372)
The standard `std::latch` is very restrictive in how it can be used, and
this makes it hard to easily leverage for simple coordination between a
set of dynamically scheduled tasks, where there isn't an interesting
synchronizing "merge" or future result.

This tool makes it easy to establish a latch, hand out handles to it,
and once all are destroyed, take whatever relevant action.

Note: this is split out of a larger change that uses it. I can wait
until the use case is ready, but seemed nice to review this separately.

---------

Co-authored-by: Dana Jansens <danakj@orodu.net>
2025-11-15 02:00:41 +00:00
Jon Ross-PerkinsandRichard Smith 973d721916 Some more edits to EnumBase and EnumMaskBase (#6054)
Adds a unit test, and some smaller edits:

- Remove the `=` when defining names, in order to change `}` placement
by clang-format on uses.
- context:
https://github.com/carbon-language/carbon-lang/pull/6053#discussion_r2343423178
- I believe with `EnumBase` that keeping the `=` had been a deliberate
choice, so this PR is intended to confirm that removing it is okay.
- Delete `EnumMaskBase::name`
- context:
https://github.com/carbon-language/carbon-lang/pull/6053#discussion_r2344233707
- We can't just do nothing because `EnumBase::name` uses indexing that's
incompatible with `EnumMaskBase`.
- Some small comment cleanups.
- Tests don't need to be in the `Carbon` namespace anymore, macros work
fine in other namespaces, but it's still the right namespace.
- Documentation on `EnumBase::name` seems to be referring to a prior
structure, wherein we had a macro defining the function instead of the
`Names` array.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2025-09-12 22:59:37 +00:00
Jon Ross-Perkins 6cc5d7ed2a Add an EnumMaskBase type (#6053)
This is a bit of an experiment to see if there's a reasonable way to
write a shared enum type, rather than writing per-case wrappers for
things like `HasTypeQualifiers` or the printing. I think it's a bit
borderline complexity right now, but I'm not sure I can reduce it much
further.

This changes from things like `Internal::EnumClassName##RawEnum` to
`Internal::EnumClassName##Data::RawEnum` so that the enum entries can
have back references to bit shifts without needing to know the
containing type name. Because I'm trying to reduce duplication between
mask and non-mask enums, I did this to non-mask enums too.

This was motivated by #6035 adding another enum mask (which will grow
more entries, and is intended to switch if this is accepted), but I'm
not using that PR as a base here because I didn't want the merge
dependency.
2025-09-12 18:04:10 +00:00
Chandler Carruth 969abfe814 Follow-up fixes to filesystem code (#5949)
Tidies up extraneous move, unnecessary function style type cast, and
simplifies the temporary directory string construction. These were
noticed during another PR review.

Also corrects support for older glibc versions, including the
GNU-specific quirks of `strerror_r`. Restricts the fancier formatting
with the name of the error number to when a recent glibc is available.

Lastly, filters the benchmarks in the benchmark test down to smaller
ones to avoid test timeout flakiness.
2025-08-12 21:51:44 +00:00
Chandler CarruthandDana Jansens 42d29764c0 Introduce a custom filesystem library (#5888)
The standard filesystem API lacks significant functionality, ranging
from correct and secure creation of directories and files within them by
using `openat` and avoiding [TOCTOU] issues, to support for filesystem
locking.

[TOCTOU]: https://en.wikipedia.org/wiki/Time-of-check_to_time-of-use

The LLVM filesystem library has more functionality, but uses an API that
is increasingly diverging from the standard, and also fails to defend
against TOCTOU.

This library is designed to carefully model the Unix or POSIX filesystem
concepts of `openat` to avoid TOCTOU. However, it also tries to limit
itself to an API subset that LLVM's filesystem library has also
implemneted and so we have a strong reason to expect to be possible to
port to Windows reasonably.

This PR included several benchmarks that show that this implementation
is also faster for the majority of operations than the C++ standard
library. The only places where there is a consistent regression is in
recursively creating directories, and this is directly connected to the
approach of using `openat` as the basis. Even there, while the wall time
regresses, the cycles and instructions are significantly improved.

There are a number of operations not yet included here, I've focused on
a core set of opening, closing, creating, and removing, and then adding
those that I saw the current toolchain code using actively. I'll plan to
expand the operations as needed going forward.

A follow-up PR that I'll finish polishing and send next ports
`//toolchain/install` to consistently use this library and
`std::filesystem::path` to both exercise the library and showcase its
use. I'll be working systematically across the toolchain to converge all
the code, extending this library as needed.

For reference, benchmark results on my macOS laptop:
https://gist.github.com/chandlerc/29d1f4d465a835b8be5174a48dad2e8f

Benchmark results on a Asahi Linux M1 Mac Mini:
https://gist.github.com/chandlerc/c42d43dd6b9b91746ab314b2afa152f7

Benchmark results on a Linux server with weirdly slow FS operations:
https://gist.github.com/chandlerc/48301a7383eb3972d53351b7e35e0561

---------

Co-authored-by: Dana Jansens <danakj@orodu.net>
2025-08-12 01:38:31 +00:00
Chandler CarruthandJon Ross-Perkins b39c7c93aa Add hashtable benchmark coverage for integers with low zero bits (#5735)
These have unique challenges for our hashing scheme, and so its useful
to make sure the hash functions we use can handle them.

Some other work on Abseil's hash tables uncovered that this might be
risky and may have surfaced some improvements to reduce the impact here,
but the first step seems to try and start covering this path in the
benchmarks.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2025-06-28 00:52:58 +00:00
Jon Ross-Perkins 9855818bb8 Move PrettyStackTraceFunction to common (#5739)
I'm looking at using this as part of file_test to dump streaming,
related to #5733
2025-06-26 18:39:54 +00:00
Richard SmithandJon Ross-Perkins dc7839e893 Add a new facility GrowingRange for a range that might grow during iteration. (#5641)
Use it to replace most existing modernize-loop-convert lints with
range-based for loops. As requested in review of #5475.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2025-06-10 23:21:17 +00:00
Richard SmithandDana Jansens 6753a715f6 Avoid moving around large suspended function states in the deferred definition worklist. (#5608)
We already go to some effort to avoid moving these, but we end up still
moving them twice: once when adding to the worklist and again when
reversing a chunk of the worklist.

* To avoid a move when constructing the worklist, add an `EmplaceResult`
utility that allows the result of a function call to be emplaced into a
container.
* To avoid moves when reversing the list, stop reversing it. Instead of
reversing the list and popping tasks as we run them, we accumulate a
sequence of tasks for a deferred definition region, run them in the
order they were enqueued, then pop them all at the end. This will in
some cases increase the high-water-mark of the size of the worklist, but
not asymptotically. The same high-water-mark could be reached with the
old approach by reordering the declarations in the source file.

In passing, we no longer create `LeaveDeferredDefinitionRegion` tasks
for non-nested regions. We don't need them, because we can detect that
condition by our reaching the end of the worklist. This means that the
enter / leave region actions are now always in correspondence -- we only
create them for *nested* regions. The tasks have been renamed to convey
this.

We still move the suspended function states around if the worklist grows
to over 64 entries and gets reallocated. We could potentially address
that issue too by switching to a chunked allocation strategy as is used
by `ValueStore` and then make the tasks noncopyable, but I'm not
attempting that in this PR.

---------

Co-authored-by: Dana Jansens <danakj@orodu.net>
2025-06-10 13:37:16 +00:00
Jon Ross-Perkins 14f19b5a86 Use TypeEnum for ScopeId to refactor call structure (#5491)
I'm trying to make the offsetting a little easier to understand, and
also get a better `requires` structure on calls. The second is for an
attempt to refactor the `Formatter` API, but also changing the `InstId`
`derived_from` requires seems helpful for clarity on what's really
happening.
2025-05-20 22:16:28 +00:00
Jon Ross-Perkins dbf12eb3fc Add a SameAsOneOf helper (#5490)
Trying to make repeated `std::same_as` easier to write. Calling it
"concepts.h" because I figure we'll maybe have a couple more things like
this.

Was looking at this because I may add a couple more similar constructs.
2025-05-16 01:39:19 +00:00
Jon Ross-Perkins 2f81858a36 Switch BuildData to char arrays (#5464)
string_view was suggested at
https://github.com/carbon-language/carbon-lang/pull/5451#discussion_r2080640267,
but it turns out it's helpful to be even more hermetic for build
configuration.
2025-05-12 23:46:02 +00:00
Jon Ross-PerkinsandChandler Carruth b1004012c3 Add linkstamp support to get the target name (#5451)
This removes hardcoding of the test target name from file_test.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2025-05-12 19:01:01 +00:00
Dana Jansens 517c4d3c20 Remove VariantMatch; use CARBON_KIND_SWITCH for std::variants (#5437)
Teach CARBON_KIND_SWITCH to handle mutable lvalues and rvalues, and
CARBON_KIND to forward along rvalues so that it's possible to write
`case CARBON_KIND(const T& t)`, `case CARBON_KIND(T& t)`, and `case
CARBON_KIND(T&& t)`, depending on the type that was passed to
CARBON_KIND_SWITCH.

Replace all uses of VariantMatch with their equivalent of a switch using
CARBON_KIND_SWITCH, and remove the VariantMatch helper from the
codebase.
2025-05-12 18:56:52 +00:00
Jon Ross-Perkins 1bb5fe73f0 Update to bazel 8.2.1 (#5445)
- Updates incompatible flags.
- `rules_flex` is no longer used, so enable its flag.
- Fixes `sh_test` deps for
`--incompatible_disable_autoloads_in_main_repo`
- Broadens the exception for `rules_cc` and `bazel_tools` due to changes
to runfiles deps; trying to avoid minutiae that shouldn't affect the
decision.
2025-05-08 00:10:14 +00:00
Richard SmithandJon Ross-Perkins e060342411 Defer building thunks until the end of the enclosing definition. (#5403)
Instead of building the definition of a thunk immediately when we
generate the thunk declaration, wait until we reach the `}` of the
outermost class, interface, etc. -- at the same time when we would parse
the definition of the thunk if it were defined inline.

This fixes issues where we fail to define the thunk because it requires
an enclosing class to be complete, or its definition depends on
something declared later in the enclosing class.

Make the representation of a suspended function scope, and its
constituent suspended components, be move-only, and switch to passing it
around by rvalue reference instead of by value because it's expensive
both to move and especially to copy.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2025-05-07 22:20:39 +00:00
Dana Jansens 9a6c74f0cd Introduce FindIfOrNull() FindIfOrNone() and Contains() (#5322)
`FindIfOrNull` returns a pointer to the element in the range if it's
found, and nullptr otherwise. `FindIfOrNone` returns a copy of the
element in the range if it's found, and `T::None` (for a range of
elements of type `T`) otherwise. `Contains` returns a bool indicating
whether the element in the range is found.

These functions replace `llvm::find()` and `llvm::find_if()` when you
want a single answer back instead of an iterator. This avoids the need
to check against `end()`, allowing the return condition to be tested as
a standard bool.

We replace uses of `find()` and `find_if()` that did not require an
iterator with these new helpers.

Note that the return type of `FindIfOrNull` is a pointer since we can
not write `optional<T&>`, which must be tested for null. If the null
check is omitted, UB occurs and the resulting code may end up with an
incorrect pointer (https://crbug.com/40153300) into the range (or
elsewhere), rather than a null dereference. And this would be very
confusing to debug. Hopefully debug builds and sanitizers keep this from
being an issue we sink a bunch of time into debugging.
2025-04-18 14:17:48 +00:00
Jon Ross-Perkins 72cfaad1c7 Remove the indirect_value library (#5312)
This was used by explorer, and no longer has uses.
2025-04-15 20:16:58 +00:00
Jon Ross-Perkins 8c3fa80691 Add cc rule wrappers for cc_env (#5277)
Rules executed by bazel don't necessarily have the right environment to
find the symbolizer, which was the intent of `cc_env` setting
`LLVM_SYMBOLIZER_PATH`. So far, this has kind of been a case-by-case
fix, but every so often I'm trying to debug a crash in a test that
doesn't provide it. Rather continuing down this route, instead add
drop-in wrappers for cc rules so that it's hard to forget.

Note `bazel/cc_rules` is intended to mirror `bazel/carbon_rules` and
`bazel/cc_toolchains`, rather than `@rules_cc`.

AFAICT there isn't a great way to add this as a default for the `bazel
run` environment. It's not typically going to be set on its own,
forwarding `$PATH` would be too broad, and the [action
`env_sets`](https://bazel.build/docs/cc-toolchain-config-reference#using-action-config)
I think are not quite what we need (I think those don't include output
execution, only compilation).
2025-04-11 19:58:54 +00:00
Jon Ross-PerkinsandGeoff Romer 4c4c4a4d2c Add RawStringOstream for slightly simpler streaming to strings (#4817)
This adds a RawStringOstream. Versus TestRawOstream, which is
consolidated over to RawStringOstream, it uses a string for storage
instead of a vector, mainly to support move-to-string semantics. Versus
llvm::raw_string_ostream, it owns the string and supports pwrite (which
is needed for driver and its fd_ostream compatibility requirement).

This converts most uses of llvm::raw_string_ostream, leaving behind a
few in InstNamer that explicitly cannot own the string, such as:

```
     llvm::raw_string_ostream(name)
          << "_" << tree.tokens().GetColumnNumber(token);
```

I have this as its own library so that it can use CHECK.

Yes this doesn't save much code, but it's code we repeatedly write.

---------

Co-authored-by: Geoff Romer <gromer@google.com>
2025-01-18 01:11:44 +00:00
Boaz Brickner fe8b42148f Mark some //common, //toolchain/driver, //‎toolchain/install tests as small per 'Test execution time' warning (#4646)
These tests only take between 0.1s and 1.4s.
2024-12-10 20:55:58 +00:00
Jon Ross-Perkins 5880954041 Refactor command line errors to mirror diagnostic style (#4568)
This changes to an `Error` return to let the driver do the "error: "
prefix, except for one case with `help` that needs more work to change
(I'm not planning on picking up that TODO). It also changes
capitalization, backtick use, and a few minor punctuation things to try
to better match the diagnostic style.

This also adds `Error` matchers so that the changes to command line
testing are clearer.
2024-11-26 19:22:05 +00:00
Chandler Carruth b08fefc896 Change the test timeouts for the benchmarks to moderate. (#4570)
After some poking, it would take a more significant change to
restructure the string generation to take less time when run under ASan,
and it's not worth it at the moment.

For future reference, nearly half the time here is in building the
global data structures of random string contents, not in the actual
benchmark functions. If/when we want to improve this, we should switch
to a growing pool of random strings similar to what `SourceGen` uses.
That lets it not allocate the full size of data when just testing that
the benchmark doesn't crash.

I thought about having these benchmarks switch to use `SourceGen`, but
I'd like to keep them stand-alone if easy, and there are some important
differences that would have to be adapted around which wouldn't be
trivial. I'd rather come back in with a better generation strategy than
re-use the source code one here.
2024-11-22 16:26:54 +00:00
Chandler Carruth 637539726d Switch to our custom benchmark main. (#4567)
I'm working on speeding up the benchmark tests and noticed they weren't
using our main, seemed worth cleaning that up.
2024-11-22 00:46:09 +00:00
Jon Ross-Perkins 493d766a97 Have sh_test directly invoke benchmarks (#4552)
These tests typically take 10-20s, but I'm seeing some timeouts
[here](https://github.com/carbon-language/carbon-lang/actions/runs/11899548036/job/33158400417).
This seemed particularly suspicious due to the _absence_ of output
(copied below). That got me looking, and maybe the subprocessing tickles
a cpu bottleneck, so proposing this approach to remove the exec. Even if
this doesn't solve the flakiness, I think it's a simpler implementation.

Note I believe this is intended to work. The `sh` rules rely on shebangs
(as noted at https://bazel.build/reference/be/shell#sh_test), and are
essentially just subprocessing to the input. Note this could've also had
`args` on a `cc_test` rule, but I'd expect the same args to be passed to
`run` where instead the benchmark behavior should be default (and I'm
assuming you'd rather not have args there). Fundamentally this becomes a
symlink:

```
bazel-bin/common/map_benchmark_test -> .../execroot/_main/bazel-out/k8-fastbuild/bin/common/map_benchmark
```

Copying snippet from timeout below:

```
==================== Test output for //common:map_benchmark_test:
      /private/var/tmp/_bazel_runner/e591f63ed099023de1f206992dfce127/execroot/_main/bazel-out/darwin_arm64-fastbuild/testlogs/common/map_benchmark_test/test.log
-- Test timed out at 2024-11-18 19:32:13 UTC --
INFO: From Testing //common:map_benchmark_test:
================================================================================
```
2024-11-20 19:00:37 +00:00
4845f40dff Switch CARBON_CHECK to a format string API (#4285)
This switches `DCHECK` and `FATAL` as well.

The goal is to reduce the code size impact of these assertions so that
we can keep more of them enabled. Currently, the largest cost I see from
`CHECK` is not the actual check or the cold code itself, but actually
the failure to inline trivial functions due to the presence of the cold
code. This means that our goal isn't to reduce apparent code size in the
final binary but the LLVM IR cost assessed for these routines in the
inliner, which closely correlates with code size but is a bit different.

As discussed in #4283, experimentation shows that a single function call
with a minimal number of arguments is the lowest cost model for these.
This is easily achieved with a format-string API that internally uses
`llvm::formatv`. This PR is essentially the `CHECK` version of #4283.

However, the check macros are substantially harder to make work with
both format strings and streaming because they also take a condition.
Also, unexpectedly, I was very successful at devising a regular
expression based automated rewrite from the streaming to the format
string form with only low 10s of manual fixes. This includes compacting
strings broken up across lines, etc. Given how well that went, I've
prepared this PR which just directly switches to the format string API
and migrate everything to use it.

One nice side-effect is that the format string approach ends up greatly
simplifying the implementation here as well.

This is ... *shockingly* effective. Parsing speeds up by more than 3%
with just this change. And checking speeds up by **8%** with this change
alone:
```
BM_CompileAPIFileDenseDecls<Phase::Parse>/256      86.3µs ± 1%  82.9µs ± 1%  -3.94%  (p=0.000 n=17+19)
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024      431µs ± 1%   415µs ± 1%  -3.76%  (p=0.000 n=18+19)
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096     1.77ms ± 1%  1.71ms ± 1%  -3.18%  (p=0.000 n=18+19)
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384    7.44ms ± 1%  7.17ms ± 2%  -3.56%  (p=0.000 n=18+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    30.7ms ± 1%  29.7ms ± 1%  -3.15%  (p=0.000 n=18+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144    131ms ± 1%   127ms ± 1%  -2.81%  (p=0.000 n=18+18)
BM_CompileAPIFileDenseDecls<Phase::Check>/256       878µs ± 2%   800µs ± 1%  -8.91%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/1024     1.88ms ± 2%  1.72ms ± 1%  -8.56%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/4096     5.78ms ± 2%  5.28ms ± 1%  -8.70%  (p=0.000 n=20+18)
BM_CompileAPIFileDenseDecls<Phase::Check>/16384    21.9ms ± 1%  20.1ms ± 1%  -8.02%  (p=0.000 n=18+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/65536    90.4ms ± 2%  83.1ms ± 1%  -8.04%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/262144    381ms ± 2%   352ms ± 1%  -7.79%  (p=0.000 n=19+19)
```

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
2024-09-12 16:42:08 +00:00
Chandler Carruth 0c8ab663c9 Migrate all CARBON_VLOG to the format string variant. (#4284)
This mostly uses a hilarious set of regular expressions to mechanically
switch all but two uses, and then manually fixed the last two. There
weren't too many.

Also simplifies the `vlog` implementation now that it's all going
through a format string.

This alone has a nice impact on parse and check of about 2% and 1%
respectively. The impact on lex in my timings looks like noise (no
change in instruction count, unlike the other phases).
```
name                                               old cpu/op   new cpu/op   delta
BM_CompileAPIFileDenseDecls<Phase::Lex>/256        39.1µs ± 3%  38.1µs ± 2%  -2.42%  (p=0.000 n=20+19)
BM_CompileAPIFileDenseDecls<Phase::Lex>/1024        187µs ± 3%   183µs ± 1%  -2.30%  (p=0.000 n=20+20)
BM_CompileAPIFileDenseDecls<Phase::Lex>/4096        776µs ± 4%   756µs ± 1%  -2.62%  (p=0.000 n=20+20)
BM_CompileAPIFileDenseDecls<Phase::Lex>/16384      3.36ms ± 1%  3.33ms ± 1%  -0.90%  (p=0.000 n=18+18)
BM_CompileAPIFileDenseDecls<Phase::Lex>/65536      14.4ms ± 2%  14.2ms ± 1%  -1.41%  (p=0.000 n=20+20)
BM_CompileAPIFileDenseDecls<Phase::Lex>/262144     65.7ms ± 1%  65.2ms ± 2%  -0.86%  (p=0.002 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/256      87.5µs ± 1%  86.3µs ± 1%  -1.43%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024      438µs ± 2%   431µs ± 1%  -1.54%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096     1.81ms ± 2%  1.77ms ± 1%  -2.12%  (p=0.000 n=20+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384    7.54ms ± 1%  7.43ms ± 1%  -1.44%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    31.2ms ± 1%  30.6ms ± 1%  -2.03%  (p=0.000 n=20+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144    133ms ± 1%   130ms ± 1%  -1.85%  (p=0.000 n=20+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/256       882µs ± 1%   878µs ± 1%  -0.52%  (p=0.001 n=17+19)
BM_CompileAPIFileDenseDecls<Phase::Check>/1024     1.90ms ± 2%  1.88ms ± 1%  -1.17%  (p=0.000 n=19+19)
BM_CompileAPIFileDenseDecls<Phase::Check>/4096     5.85ms ± 2%  5.76ms ± 1%  -1.43%  (p=0.000 n=20+19)
BM_CompileAPIFileDenseDecls<Phase::Check>/16384    22.2ms ± 2%  21.9ms ± 2%  -1.20%  (p=0.000 n=20+19)
BM_CompileAPIFileDenseDecls<Phase::Check>/65536    91.2ms ± 2%  90.3ms ± 1%  -1.00%  (p=0.000 n=20+19)
BM_CompileAPIFileDenseDecls<Phase::Check>/262144    382ms ± 1%   380ms ± 1%  -0.51%  (p=0.003 n=18+19)
```
2024-09-11 12:11:23 +00:00
e48101b608 Switch CARBON_VLOG to support a format string API. (#4283)
The goal is to replace our stream operator APIs with format string APIs
that can be made to have much less impact on inlining and other
optimizations of the performance critical path through the code.

Several experiments show that the most compact representation we can
arrange for is one that calls an uninlined function and passes a minimal
number of arguments to it. It doesn't help to do any work to minimize
the arguments such as building a lambda -- the cost of extra code to
merge the arguments is likely to outweigh the benefit.

Initial experiments showed that switching a hot but uninlined function
to this new API enabled inlining and the subsequent performance
improvement.

This also adds a 'TemplateString` utility that allows using a string
literal as a template parameter. This is useful to remove the format
string itself from the arguments passed to the function by passing it as
a template argument instead.

Currently, support is left in place for both APIs because with
`CARBON_VLOG` we can detect whether or not any message was provided
expecting a format string. This should allow incrementally migrating
code to this API. I've added some test coverage in this PR, but I'll
separate out any switching of parts of the codebase over.

The goal is to eventually replace all the usages and remove the
streaming support entirely.

This PR doesn't update `CARBON_CHECK` in the same way because it is
substantially more complex to switch. I have a few experimental PRs
looking at that and will discuss how best to approach this with the
specific challenges check presents separately. But the goal is for all
of the macro-based output APIs to move to format strings rather than
streams.

---------

Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-09-11 10:38:59 +00:00
Chandler Carruth 72cb9d0d06 Refactor testing exe path and benchmark main handling. (#4216)
Consolidates both main libraries into `//testing/base`, and factors out
the exe path handling for benchmarks and unit tests into a common
library to remove duplication. Refactors how that logic is managed to be
cleaner and avoid a confusing bool that came up in code review.

Updates all the tests and benchmarks that use these. I still need to
update other benchmarks to use the same main, but I wanted to keep this
PR somewhat minimal.

This also fixes a bug noticed in passing that the compilation benchmark
didn't have the required dependency on the benchmark library itself,
just the benchmark main library.
2024-08-14 17:50:08 +00:00
a9c815c9f4 Introduce a source generator and end-to-end compile benchmarks (#4124)
The big addition here is a very, very rough and very early skeleton of a
source code generator framework. This builds upon the lexers identifier
synthesis logic, improving on its framework and wiring it up with the
most rudimentary of source file generation. This is just enough to
roughly replicate my "big API file" source code benchmarks.

The source generation works *very* hard to both vary the structure and
content of the source as much as possible while ensuring the same
*total* amount of each construct is in use, from bytes in identifiers to
line breaks, parameters, etc. This lets us generate randomly structure
inputs that should consistently take the exact same amount of total work
to compile.

The complex identifier synthesis logic from the lexer's benchmark is
moved over here and the lexer uses APIs in the source generator for
identifiers. The other source synthesis in the lexer's benchmark isn't
yet moved over, but should likely be slowly absorbed here as it can be
refactored into a more principled and re-usable form. Some bits may stay
of course if they're just too lexer-specific.

Next, this adds a simple end-to-end compile benchmark for the driver
that directly and much more clearly reproduces all the measurements I've
done manually up until now. It should also be easy to extend to more
patterns over time as we add support to the source generator to produce
those patterns.

Last but not least, I've added a tiny CLI to the source generator so
that you can generate source code manually. This is especially nice for
generating demo source code to actually run through the driver or look
at in an editor. The CLI can also generate C++ source code which lets us
do some minimal comparative benchmarking between Carbon and C++/Clang.

There are huge number of TODOs in the source generation framework. This
is going to be a large ongoing effort I suspect.

There are also a bunch of rough edges I've left to try and get this out
for review sooner. I've left TODOs for refactorings that really need to
be done here, but hoping these can maybe be follow-ups. If not, please
flag and I'll try to layer them on here.

Sample compile benchmark output, nicely showing where we are w.r.t. our
goal speeds (2x behind on lex and check, 5x on parse) at least on a
recent AMD server CPU:
```
------------------------------------------------------------------------------------------------------
Benchmark                                                 Time             CPU   Iterations      Lines
------------------------------------------------------------------------------------------------------
BM_CompileAPIFileDenseDecls<Phase::Lex>/256           29420 ns        29419 ns        22860 6.62847M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/1024         146130 ns       146128 ns         4840 6.69959M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/4096         601584 ns       601577 ns         1020 6.69573M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/16384       2547578 ns      2547313 ns          280   6.404M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/65536      10816591 ns     10816389 ns           80 6.05193M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/262144     52191320 ns     52189828 ns           20 5.02261M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/256        101706 ns       101698 ns         6900 1.91745M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024       512161 ns       512162 ns         1380  1.9115M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096      2078426 ns      2078430 ns          340   1.938M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384     8795786 ns      8795583 ns          100 1.85468M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    35073596 ns     35072973 ns           20 1.86639M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144  151100688 ns    151097370 ns           20 1.73483M/s
BM_CompileAPIFileDenseDecls<Phase::Check>/256        957059 ns       957049 ns          740 203.751k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/1024      1956134 ns      1955985 ns          360 500.515k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/4096      5797864 ns      5797417 ns          120 694.792k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/16384    21219608 ns     21217584 ns           40 768.843k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/65536    96311116 ns     96302334 ns           20 679.734k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/262144  371637963 ns    371609964 ns           20 705.387k/s
```

Lest someone think this is *bad*, the fact that we're already within 2x
of our rather audacious goals makes me quite happy. =D

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-08-13 19:37:55 +00:00
Jon Ross-PerkinsandChandler Carruth d437e4bffe Create an array stack type for a shared use-case (#4100)
Based on discussion around the region handling in generic_region_stack,
create a generic structure for the stack-of-vectors support. I also want
to add this to InstBlockStack, but that's a little more complex due to
GlobalInit, so cutting a PR here to check with review.

My work here is how I noticed #4099; I want to be sure that I'm correct
about the issue, but it's the difference between being able to use
PeekArray or not.

Note in scope_stack.h, I believe we could remove next_compile_time_index
and make it just based on elements_size(). However, I want to verify
with you before I make further changes there.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2024-07-03 16:58:47 +00:00
Chandler Carruth a8748f3e2d Key context improvements (#4095)
This injects a customization point for hashtable-specific equality
testing that the key context uses by default. While this is rarely
needed, there are LLVM types where it is necessary and it seems a good
general tool to have to avoid unnecessary complexity from custom key
contexts when a simple customization of equality is all that is
required.

This also adds a CRTP mixin for implementing a common pattern of key
contexts where the context provides translation of some key types into
another type, potentially using state. Rather than having to implement
the entire key context API, code can derive from this template and
simply provide a set of overloads for the types it wants to translate.
Any key types used which can be passed to one of those overloads will
get translated before following the same logic as the default key
context. While this updates the only usage so far of this pattern, a
subsequent PR will add several more users making the pattern worth
abstracting here.
2024-07-02 21:20:14 +00:00
Chandler CarruthandJon Ross-Perkins b51dc7f8e2 Introduce version and build info stamping. (#4054)
This adds a defined Carbon version to the Bazel build and codebase that
can be used both to implement features like version checks and to report
a meaningful version on the command line. This replaces a hard-coded
string and a TODO in the driver.

As part of this, it adds support for defining the version in Bazel, and
special build flags for overriding relevant parts such as the
pre-release marker used. The exact structure and meaning of our version
string, including the pre-release parts, is implemented here in line
with the draft proposal:

https://docs.google.com/document/d/11S5VAPe5Pm_BZPlajWrqDDVr9qc7-7tS2VshqO0wWkk/edit?resourcekey=0-2YFC9Uvl4puuDnWlr2MmYw

This also introduces a workspace status command to the repository to
extract the git commit SHA and other information when building, and the
logic to stamp that into binaries as part of the version string when
useful. The technique used leverages weak symbols with whole archive
linking to allow a link-time override of unstamped data with stamped
data in the leaf executable. This makes building with `--stamp` a
reasonable default, especially for development builds. The CI system is
explicitly opted out of this as there it has no benefit.

Last but not least, all of these are wired into the install rules so
that we build installable packages with the version number in a
conventional place in the directory and filename.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-06-18 00:33:05 +00:00
Chandler Carruth 26ead9addc Add a build of boost_unordered for benchmarking. (#4045)
This uses a header-only extraction of the Boost unordered hashtable
project to allow a trivial Bazel build and for us to benchmark against
it effectively.
2024-06-10 22:10:10 +00:00
Chandler Carruthandjosh11b 21a81bc59e Introduce custom hash table data structures. (#3940)
The hash table design is heavily based on Abseil's ["Swiss
Tables"][swiss-tables] design. It uses an array of bytes storing
metadata about each entry and an array of entries where each is a pair
of key and value. The metadata byte consists of 7-bits of hash of the
key (distinct from the bits used to index the table), and one bit
indicating the presence of a special entry -- either empty or deleted.

[swiss-tables]: https://abseil.io/about/design/swisstables

There are a large range of optimizations and other nuanced aspects of
this hash table design and implementation, a good point to understand
that context is `raw_hashtable.h` which has an overview of the design
and references to various other files for relevant details.

---------

Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
2024-06-08 01:50:02 +00:00
Chandler Carruth 582b2c04dd Allow GoogleTest tests to locate their executable path. (#3992)
This also factors out the code for doing this location from the driver
to a tiny helper library.

The motivation for this is letting tests locate data files like the
prelude or other files needed by the toolchain more easily. A subsequent
PR will use this heavily in the Clang runner and related logic for
example.
2024-05-28 23:50:57 +00:00
f5c01ee487 Add a custom main library for benchmarks. (#3872)
This largely reproduces the upstream one, but has a few advantages:

1) It initializes LLVM which can be important if we use the backends.
2) It parses commandline flags with Abseil allowing the use of Abseil
   flags in benchmarks.
3) It flips the default for tabular display of results which we use
   pretty heavily.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com>
2024-04-09 20:36:41 +00:00
Richard SmithandJon Ross-Perkins f9ce0b194d Defer parsing of method bodies until the end of a suitable enclosing scope. (#3832)
In parse, form a list of methods that are defined inline, tracking where
they start, where they end, and which other inline methods are nested
within them.

In check, when we reach an inline method body, skip it and add it to a
worklist to be processed later. We also track when we reach the start
and end of a context in which inline method bodies are deferred, so that
we know when to replay the bodies.

When suspending a function definition to be processed later, the
`DeclNameStack` entry is moved to separate storage, including popping
the corresponding scopes from the scope stack and removing the
corresponding lexical names from lexical lookup. Later, when we return
to the function and parse its definition, the `DeclNameStack` entry is
restored. The same is done when we reach the end of a nested context
that can have inline methods, so that we can reenter the nested scope
before processing its members.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-04-01 18:25:27 +00:00
Jon Ross-Perkins a3b1c433be Remove legacy repo_name settings (#3772)
I'd kept these in to separate the bazel module update from the BUILD
file changes, then forgot about it. I think all of these can be cleanly
removed now. I think it's something we should clean up for consistency
with the bazel central repository names; I think it's best to reduce
that divergence.

llvm_zlib and llvm_zstd remain because of how llvm depends on the
particular names.
2024-03-13 22:58:56 +00:00
Richard SmithandJon Ross-Perkins bf8697113a Move llvm::Initialize* calls to main. (#3449)
Per their documentation, the `llvm::Initialize*` functions are only
supposed to be called by the main program, not by a library like
toolchain/codegen. Fixes a hang due to a data race in multithreaded
autoupdate.

Add a utility class `Carbon::InitLLVM` to do the common LLVM
initialization shared by all Carbon tools, optionally including
initializing the LLVM targets. Because the LLVM targets add a lot of
binary size, only initialize them for binaries that opt in by depending
on a new target `//common:all_llvm_targets`.

Also fix `//explorer:file_test` and `//explorer:file_test.trace` to
share a binary rather than linking an identical binary twice.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-12-07 01:45:47 +00:00
f59a6cdbdd Introduce a Carbon hashing framework. (#3327)
# Overview

This is a latency-optimized hashing framework based on Abseil's and
others. At it's core it uses both a normal 64-bit multiply as well as a
64-bit multiply capturing both low and high 64-bit components of the
result and XOR-ing them together. These are the primitives used in
FxHash and Abseil respectively, although they both appear in others.

The implementation has been *substantially* optimized for short inputs
and latency over quality. As a result, this function does not remotely
pass the SMHasher quality tests. However, basic collisions are rare, and
I've included a small subset of the SMHasher collision testing directly
to make sure the quality doesn't slip too far inadvertently.

The customization framework is roughly similar to Abseil's and LLVM's
but has been simplified significantly, inspired in some respects by the
AHash API design and in others by my experience of all performance
sensitive hashing implementations needing to work at a very low level to
hit their performance targets. The abstractions are stripped down to
facilitate this.

# Details of the performance optimization

This function is 2x - 4x faster than LLVM's on small inputs, and up to
2x faster than Abseil. Significant effort has gone into optimizing short
strings in particular compared to Abseil.

Small integer and pointer hashing is also faster than Abseil's by
leveraging a lower quality 64-bit multiply in some cases inspired by
FxHash. One consequence is that this routine is particulary fast for
32-bit integers.

The short string improvements largely come from packing more of the
bytes of string into as few multiplies as possible. While this fails to
mix the bits sufficient to hit SMHasher's strict avalanche criteria and
does leave some collision windows, it provides dramatic latency
improvements. Some of these techniques come from Abseil's own bulk
hashing routine but re-applied here. Others are novel, for example using
small sizes to sample nicely uniform random data to efficiently handle
the very small number of bits of data that need to be hashed.

The other observed improvement is diligent handling of pairs and tuples
and fairly aggressively turning things into integers. Some of the
comparisons with Abseil aren't realistic as the Abseil hash table does
some of these mappings before hashing. I've done this directly in the
hash function as that seems cleaner.

For long strings, the performance is comparable or a bit better than
Abseil, and significantly better than LLVM's hash function.

Overall, for short inputs this is hoped to be the fastest hash function
that still gets "just enough" mixing for modern hash tables to perform
well.

# Details of the quality vs. latency tradeoff

A key insight is that modern hash tables don't need especially high
quality hash functions, but do benefit from something beyond the
identify function. That isn't the target of SMHasher or other quality
assessing tools and has resulted in unnecessarily aggressive hashing for
any functions actually evaluated against it. Many hash functions turn
off the high quality implementations evaluated with SMHasher for integer
or pointer keys to recover latency & performance (AHash for example),
but the same performance-oriented design applies beyond these narrow
types, for example for short strings.

However, a consequence is that there are serious limits to the quality
of the hash function. The avalanche test is failed hilariously, etc.,
but in the exact same ways as Abseil itself fails it for integer keys.
There are also real collisions spaces. For example, for 16-byte strings,
there is one 64-bit value for the first 8 bytes that will have the same
hash regardless of the other 8 bytes of the string. Some minor effort is
taken to make this pattern unlikely to be a practical problem, but it is
a clear theoretical weakness.

It also means that this hash function couldn't be further from providing
any hash-flooding DoS attack protection -- I expect it to be trivially
easy to attack in this way by a motivated adversary. Defending against
these attacks is defined as out-of-scope, in large part because even
attempts that have made a compelling effort to address these issues such
as HighwayHash have found serious limits. Instead, this takes a
principled position that any such defense should be provided entirely at
the data structure level with a strong worst-case bound rather than
through strengthening the hash function.

# Future work

A subsequent PR will introduce a hash table inspired very heavily by the
design of Abseil's "SwissTable" and using this hash function. The goal
is to provide a significant improvement to hot hash tables such as the
identifier table in the lexer of Carbon's toolchain.

# Detailed benchmark data

The benchmarks introduced are heavily inspired by the latency
benchmarking of hash functions in Abseil. I've adapted them to fit
better into Carbon's coding style and to try to have more stable results
with broader coverage of types and string sizes.

Running the benchmarks directly gives horizontal comparisons across
different hash functions. That can be hard to read, so here is *just*
the newly introduced hash function benchmark results on an AMD server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          3.11ns ± 1%
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         3.11ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      4.11ns ± 1%
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         3.12ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    4.13ns ± 1%
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         3.16ns ± 2%
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             3.16ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    4.03ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    4.04ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    4.34ns ± 2%
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        4.04ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        4.34ns ± 2%
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            4.34ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        4.33ns ± 1%
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        1.95ns ± 4%
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        1.70ns ± 3%
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       3.52ns ± 3%
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       4.46ns ± 2%
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       7.69ns ± 1%
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      14.8ns ± 1%
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      21.5ns ± 1%
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     34.6ns ± 0%
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     63.1ns ± 1%
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      118ns ± 1%
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      225ns ± 1%
```

And on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          5.28ns ± 0%
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         5.29ns ± 0%
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      7.02ns ± 0%
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         5.34ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    7.07ns ± 4%
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         5.36ns ± 2%
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             5.36ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    7.19ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    7.29ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    7.31ns ± 4%
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        7.29ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        7.31ns ± 4%
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        8.69ns ± 3%
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        2.64ns ± 2%
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        2.90ns ± 4%
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       6.14ns ± 1%
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       8.27ns ± 1%
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       13.8ns ± 0%
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      31.2ns ± 0%
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      49.9ns ± 0%
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     86.9ns ± 0%
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      163ns ± 0%
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      312ns ± 0%
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      610ns ± 0%
```

I don't have the same nice statistical multi-run error bars, but one run
from my M1 MacBook:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                             3.89 ns
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                            3.87 ns
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>         4.39 ns
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                            3.93 ns
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>       4.98 ns
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                            3.87 ns
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                                3.87 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>       4.86 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>       4.43 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>       4.41 ns
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>           4.44 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>           4.69 ns
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                         4.33 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>       4.38 ns
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>               4.34 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>           4.35 ns
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>           4.38 ns
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                           1.15 ns
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                          0.973 ns
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                          3.03 ns
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                          3.97 ns
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                          6.64 ns
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                         12.5 ns
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                         17.9 ns
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                        27.9 ns
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                        48.1 ns
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                        87.3 ns
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                         166 ns
```

And here I have internally replaced the Carbon hash function with
Abseil's hash function for "before" and then restored it in the "after"
and computed the delta for each benchmark. This basically shows the
speed-up (lower time -> lower latency -> speed-up -> good) over Abseil
on an AMD server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          4.00ns ± 1%  3.10ns ± 0%  -22.45%  (p=0.000 n=20+15)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         4.01ns ± 1%  3.10ns ± 1%  -22.64%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      6.25ns ± 1%  4.10ns ± 1%  -34.30%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         4.02ns ± 1%  3.12ns ± 1%  -22.50%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    6.25ns ± 1%  4.11ns ± 1%  -34.20%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         4.03ns ± 1%  3.14ns ± 1%  -22.17%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             5.95ns ± 1%  3.14ns ± 1%  -47.24%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    6.04ns ± 1%  4.01ns ± 1%  -33.64%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    5.96ns ± 1%  4.02ns ± 1%  -32.51%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    5.93ns ± 1%  4.30ns ± 1%  -27.56%  (p=0.000 n=20+17)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        7.97ns ± 1%  4.02ns ± 1%  -49.50%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        7.98ns ± 1%  4.32ns ± 1%  -45.88%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      4.40ns ± 2%  4.32ns ± 1%   -1.81%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    5.94ns ± 1%  4.32ns ± 1%  -27.25%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            10.0ns ± 1%   4.3ns ± 1%  -56.56%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        8.04ns ± 1%  4.32ns ± 1%  -46.29%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        7.95ns ± 1%  4.33ns ± 1%  -45.59%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        3.28ns ± 3%  1.93ns ± 4%  -41.19%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        3.05ns ± 3%  1.69ns ± 4%  -44.52%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       5.88ns ± 2%  3.50ns ± 3%  -40.42%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       8.92ns ± 1%  4.44ns ± 2%  -50.22%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       12.0ns ± 1%   7.7ns ± 1%  -36.16%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      18.8ns ± 0%  14.7ns ± 1%  -21.73%  (p=0.000 n=17+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      25.5ns ± 1%  21.4ns ± 1%  -16.18%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     38.7ns ± 2%  34.5ns ± 1%  -10.78%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     69.7ns ± 1%  62.8ns ± 1%   -9.88%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      130ns ± 1%   117ns ± 1%   -9.45%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      244ns ± 0%   225ns ± 1%   -8.11%  (p=0.000 n=17+20)
```

... and on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          6.48ns ± 1%  5.28ns ± 0%  -18.62%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         7.40ns ± 1%  5.29ns ± 1%  -28.45%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      10.4ns ± 0%   7.0ns ± 0%  -32.34%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         6.56ns ± 1%  5.32ns ± 1%  -18.95%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    10.8ns ± 2%   7.0ns ± 1%  -34.89%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         6.71ns ± 3%  5.38ns ± 2%  -19.84%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             10.3ns ± 3%   5.4ns ± 2%  -47.67%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    10.9ns ± 2%   7.2ns ± 4%  -33.67%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    10.7ns ± 4%   7.3ns ± 4%  -31.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    10.5ns ± 3%   7.3ns ± 4%  -30.71%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        14.1ns ± 3%   7.3ns ± 4%  -48.32%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        14.0ns ± 1%   7.3ns ± 4%  -47.95%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      9.41ns ± 4%  8.68ns ± 4%   -7.71%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    12.2ns ± 2%   8.7ns ± 4%  -28.81%  (p=0.000 n=18+19)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            18.9ns ± 2%   8.7ns ± 4%  -54.17%  (p=0.000 n=17+19)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        15.6ns ± 2%   8.7ns ± 4%  -44.37%  (p=0.000 n=17+19)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        15.5ns ± 2%   8.7ns ± 4%  -44.08%  (p=0.000 n=18+19)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        5.89ns ± 2%  2.64ns ± 3%  -55.26%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        5.73ns ± 3%  2.88ns ± 3%  -49.71%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       10.1ns ± 1%   6.1ns ± 2%  -39.00%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       15.7ns ± 0%   8.3ns ± 1%  -47.27%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       21.2ns ± 0%  13.8ns ± 0%  -34.81%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      37.9ns ± 0%  31.2ns ± 0%  -17.77%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      56.8ns ± 0%  49.8ns ± 0%  -12.21%  (p=0.000 n=20+18)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     93.8ns ± 0%  86.9ns ± 0%   -7.38%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      174ns ± 0%   163ns ± 0%   -6.03%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      330ns ± 0%   312ns ± 0%   -5.25%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      641ns ± 0%   610ns ± 0%   -4.79%  (p=0.000 n=19+19)
```

This is the same as the above delta comparison, but with the "before"
being LLVM's hash function:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          6.85ns ± 1%  3.10ns ± 1%  -54.78%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         6.85ns ± 1%  3.10ns ± 1%  -54.78%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      6.25ns ± 1%  4.09ns ± 1%  -34.58%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         6.87ns ± 1%  3.12ns ± 2%  -54.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    7.35ns ± 1%  4.10ns ± 1%  -44.20%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         7.34ns ± 1%  3.13ns ± 1%  -57.34%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             7.33ns ± 1%  3.13ns ± 2%  -57.27%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    7.27ns ± 1%  3.99ns ± 1%  -45.12%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    14.5ns ± 1%   4.0ns ± 1%  -72.23%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    14.6ns ± 1%   4.3ns ± 2%  -70.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        14.5ns ± 1%   4.0ns ± 1%  -72.21%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        14.6ns ± 1%   4.3ns ± 1%  -70.46%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      7.31ns ± 1%  4.33ns ± 1%  -40.81%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    7.78ns ± 1%  4.32ns ± 1%  -44.45%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            7.78ns ± 2%  4.33ns ± 1%  -44.42%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        7.62ns ± 1%  4.32ns ± 1%  -43.24%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        7.77ns ± 1%  4.33ns ± 1%  -44.34%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        8.15ns ± 3%  1.94ns ± 5%  -76.16%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        7.02ns ± 3%  1.69ns ± 4%  -75.94%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       7.83ns ± 2%  3.50ns ± 3%  -55.34%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       9.17ns ± 1%  4.43ns ± 2%  -51.65%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       11.3ns ± 1%   7.6ns ± 1%  -32.04%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      23.0ns ± 1%  14.7ns ± 1%  -36.14%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      32.9ns ± 0%  21.4ns ± 1%  -34.96%  (p=0.000 n=17+19)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     52.2ns ± 1%  34.4ns ± 1%  -34.01%  (p=0.000 n=19+18)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     92.1ns ± 1%  62.8ns ± 1%  -31.82%  (p=0.000 n=19+19)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      169ns ± 1%   117ns ± 1%  -30.53%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      319ns ± 1%   224ns ± 1%  -29.78%  (p=0.000 n=20+18)
```

... and on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          8.38ns ± 0%  5.27ns ± 0%  -37.04%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         8.39ns ± 1%  5.28ns ± 0%  -37.01%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      8.07ns ± 0%  7.02ns ± 0%  -13.10%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         8.48ns ± 1%  5.32ns ± 1%  -37.25%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    9.34ns ± 2%  7.09ns ± 2%  -24.14%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         9.76ns ± 3%  5.37ns ± 2%  -44.98%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             9.76ns ± 3%  5.37ns ± 2%  -44.98%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    10.1ns ± 2%   7.2ns ± 3%  -29.36%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    11.9ns ± 2%   7.3ns ± 4%  -38.68%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    11.3ns ± 2%   7.3ns ± 4%  -35.16%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        11.9ns ± 2%   7.3ns ± 4%  -38.68%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        11.3ns ± 2%   7.3ns ± 4%  -35.16%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      10.3ns ± 2%   8.7ns ± 3%  -15.81%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        9.39ns ± 2%  2.66ns ± 3%  -71.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        10.7ns ± 3%   2.9ns ± 3%  -72.97%  (p=0.000 n=19+18)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       11.8ns ± 1%   6.1ns ± 2%  -47.75%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       13.9ns ± 1%   8.3ns ± 1%  -40.71%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       16.8ns ± 1%  13.8ns ± 0%  -17.83%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      31.7ns ± 1%  31.2ns ± 0%   -1.76%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      43.5ns ± 0%  49.8ns ± 0%  +14.56%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     66.2ns ± 0%  86.9ns ± 0%  +31.39%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      112ns ± 0%   163ns ± 0%  +46.09%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      201ns ± 0%   312ns ± 0%  +55.49%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      379ns ± 0%   610ns ± 0%  +61.08%  (p=0.000 n=20+20)
```

Note that there is a significant regression on long strings compared to
LLVM's hash function on the ARM server I have access to. This doesn't
show up on the M1 at all, and is likely specific to inadequate
throughput for the 64-bit multiply operations. This seems fine as a) our
priority is for short strings, and b) the M1 and other ARM CPUs are
likely to improve here over time given the prevalent use of this core
technique. For example, Abseil's current hash algorithm has the same
long-string behavior (and performance bottleneck) on this server.

---------

Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Co-authored-by: Geoff Romer <gromer@google.com>
2023-11-20 19:59:06 +00:00
Jonathan B. CoeandChandler Carruth 8c28a0494e Add size="small" to test targets where advised (#3326)
Running `bazel test //...` reported:

```
Test execution time outside of range for MODERATE tests.
Consider setting timeout="short" or size="small".
```

This change adds size="small" to avoid such warnings being reported.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-10-24 06:52:00 +00:00
Richard Smith c7a9e29a89 Add typed nodes to SemIR. (#3280)
Replace `SemIR::Node::GetAsFoo` and `SemIR::Node::Foo::Make` with
`SemIR::Foo` class that represents a particular kind of node, with named
fields.

Rename `SemIR::IntegerLiteral` and `SemIR::RealLiteral` to
`IntegerValue` / `RealValue` to better reflect their purpose and avoid a
name collision with the corresponding `SemIR` node kinds.

Remove `NodeKind::Invalid` and the `SemIR::Node` default constructor
entirely, as they were not used for anything.
2023-10-11 05:39:59 +00:00
Chandler Carruth 4596cd230d Avoid building the non-test file group in :all. (#3191)
This file group exists to allow a `genquery` rule and a Python test to
verify our non-test dependency graph. We don't actually need to build
the binaries in the file group as part of that. The `genquery` rule
seems to do the right thing -- building it directly doesn't cause the
binaries in the group to be built. But without a manual tag, the group
itself is part of `:all` and thus part of `//...` and part of the rules
that will be built even with PR #3106. A consequence is that any change
to the toolchain causes several other binaries to be built as well
because this file group is in the impacted set. Making it manual should
avoid all of this, and without breaking the actual use from `genquery`.

For example, before this change, in a fully cached build after a `bazel
clean`:
```
> bazel test //bazel/check_deps:all
INFO: Invocation ID: 2d83ebee-4c00-425d-be33-23f42b079614
INFO: Analyzed 3 targets (103 packages loaded, 7137 targets configured).
INFO: Found 2 targets and 1 test target...
INFO: Elapsed time: 4.081s, Critical Path: 2.61s
INFO: 3111 processes: 2796 disk cache hit, 315 internal.
INFO: Build completed successfully, 3111 total actions
```

After this change:
```
> bazel test //bazel/check_deps:al
INFO: Invocation ID: c94089e8-a420-4d3c-9902-134e6b55b297
INFO: Analyzed 2 targets (92 packages loaded, 568 targets configured).
INFO: Found 1 target and 1 test target...
INFO: Elapsed time: 0.700s, Critical Path: 0.01s
INFO: 7 processes: 2 disk cache hit, 5 internal.
INFO: Build completed successfully, 7 total actions
```

While here, re-generate the file group, and fix several issues it
uncovers: mark test utilities as `testonly` and update our LLVM package
allowlist to include `clangd`'s package.
2023-10-05 01:30:41 +00:00
Jon Ross-Perkins 53af8f04b2 Provide a Printable CRTP parent to replace HasPrintable templates. (#3166)
With the toolchain splitting namespaces, ostream.h's `operator<<`
templates aren't reliably found with name lookup, likely due to the loss
of associated namespaces (zygoloid commented on this at
https://github.com/carbon-language/carbon-lang/pull/3161#discussion_r1307941999).
This is especially a barrier to moving the lex files into `Carbon::Lex`;
versus other parts of the toolchain, they contain more printable types
which are used cross-namespace, including `Carbon::Testing`. As a
consequence, I'm looking at migrating ostream.h to a more reliable
approach that doesn't rely as much on everything being in the `Carbon`
namespace.
2023-08-30 21:32:19 +00:00
Jon Ross-PerkinsandRichard Smith 7157445f97 Set up a 'Parse' namespace. (#3161)
Continuing on #3070.

I moved ParseTree::Node to just Parse::Node, versus Parse::Tree::Node.
Other name changes are just removing "Parse" or "Parser" prefixes.

In EnumBase, I'm directly defining operator<< because the ostream.h
approach just isn't working, not for either of Parse::State nor
Parse::NodeKind. Errors look like:

```toolchain/parser/parser_context.cpp:449:34: error: invalid operands to binary expression ('llvm::raw_ostream' and 'const Carbon::Parse::State')
    output << "\t" << i << ".\t" << entry.state;
    ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ^  ~~~~~~~~~~~
```

The expected template in `Carbon::` is not in the error list; I only see
the:

```
./common/ostream.h:112:6: note: candidate template ignored: requirement 'std::is_base_of_v<std::ostream, llvm::raw_ostream>' was not satisfied [with S = llvm::raw_ostream, T = Carbon::Parse::State]
auto operator<<(S& standard_out, const T& value) -> S& {
     ^
```

I'm still prodding at this, but not seeing an obvious fix.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-28 21:58:38 +00:00