Commit Graph
174 Commits
Author SHA1 Message Date
Chandler CarruthandJon Ross-Perkins b51dc7f8e2 Introduce version and build info stamping. (#4054)
This adds a defined Carbon version to the Bazel build and codebase that
can be used both to implement features like version checks and to report
a meaningful version on the command line. This replaces a hard-coded
string and a TODO in the driver.

As part of this, it adds support for defining the version in Bazel, and
special build flags for overriding relevant parts such as the
pre-release marker used. The exact structure and meaning of our version
string, including the pre-release parts, is implemented here in line
with the draft proposal:

https://docs.google.com/document/d/11S5VAPe5Pm_BZPlajWrqDDVr9qc7-7tS2VshqO0wWkk/edit?resourcekey=0-2YFC9Uvl4puuDnWlr2MmYw

This also introduces a workspace status command to the repository to
extract the git commit SHA and other information when building, and the
logic to stamp that into binaries as part of the version string when
useful. The technique used leverages weak symbols with whole archive
linking to allow a link-time override of unstamped data with stamped
data in the leaf executable. This makes building with `--stamp` a
reasonable default, especially for development builds. The CI system is
explicitly opted out of this as there it has no benefit.

Last but not least, all of these are wired into the install rules so
that we build installable packages with the version number in a
conventional place in the directory and filename.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-06-18 00:33:05 +00:00
Chandler Carruthandjosh11b 07fadc6474 Add growth API to the new hashtables. (#4044)
This adds two different growth APIs. This is instead of the more
conventional STL `reserve` method. One allows users that aren't trying
to grow in anticipation of an *exact* count of insertions, but generally
trying to size the table to the correct ballpark with a power-of-two
estimate.

The other API allows pre-growing to allow a specific number of
insertions to be performed without further growth. This API takes the
maximum load factor and other implementation details into account.

---------

Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
2024-06-12 19:38:35 +00:00
Chandler Carruth b773ff3695 Reduce the new hashtable test times. (#4047)
While it was convenient once to have an immediate check while inserting,
it is indeed far too quadratic. The test was taking 30-60 seconds for
me. =[ So most of the fix here is just to stop doing the check on every
insertion for all previous elements.

There were a few other somewhat slow steps, and I tried to pull those
back as well. I don't think we lose any utility here. Now everything
runs nice and quickly. =]
2024-06-11 03:43:18 +00:00
Chandler Carruth 26ead9addc Add a build of boost_unordered for benchmarking. (#4045)
This uses a header-only extraction of the Boost unordered hashtable
project to allow a trivial Bazel build and for us to benchmark against
it effectively.
2024-06-10 22:10:10 +00:00
Chandler CarruthandRichard Smith 3be57b71e0 Collect more detailed metrics on hashtables. (#4046)
Previously we just looked at the raw count of probed keys. Now, we
compute the average and max of both the probe _distance_ measured in the
number of _groups_ probed, and the number of probe _compares_ measured
in the compares required _before_ finding the matching entry.

This lets us understand the relative impact of probe-distance vs. tag
collisions on a given set of benchmark keys. Some of this is motivated
by considering additional optimization techniques similar to those used
in Boost's table and the F14 table from Facebook/Meta.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-06-10 21:43:13 +00:00
Chandler Carruthandjosh11b 21a81bc59e Introduce custom hash table data structures. (#3940)
The hash table design is heavily based on Abseil's ["Swiss
Tables"][swiss-tables] design. It uses an array of bytes storing
metadata about each entry and an array of entries where each is a pair
of key and value. The metadata byte consists of 7-bits of hash of the
key (distinct from the bits used to index the table), and one bit
indicating the presence of a special entry -- either empty or deleted.

[swiss-tables]: https://abseil.io/about/design/swisstables

There are a large range of optimizations and other nuanced aspects of
this hash table design and implementation, a good point to understand
that context is `raw_hashtable.h` which has an overview of the design
and references to various other files for relevant details.

---------

Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
2024-06-08 01:50:02 +00:00
Chandler Carruth 8c64f0bfdd Add -Wmissing-prototypes and fix issues it finds. (#4019)
Most of these are places where we failed to include a header file and
simply never got an error about this. The fix is to include the header
file.

Most other cases are functions that should have been marked `static` but
were not. Finding all of these was a main motivation for me enabling the
warning despite how much work it is.

One complicating factor was that we weren't including the `handle.h` for
all the state-based handler functions. While this isn't a tiny amount of
code, it is just declarations and doesn't add any extra dependencies. It
also lets us have the checking for which functions need to be `static`
and which don't. For the `parse` library I had to add the `handle.h`
header as well, I tried to match the design of it in `check`.

I have also had to work around a bug in the warning, but given the value
it seems to be providing, that seems reasonable. I've filed the bug
upstream: https://github.com/llvm/llvm-project/issues/94138

I also had to use some hacks to work around limitations of Bazel rules
that wrap `cc_library` rules and don't expose `copts`. I filed a bug for
`cc_proto_library` specifically:
~https://github.com/bazelbuild/bazel/issues/22610~ 
https://github.com/bazelbuild/bazel/issues/4446
2024-06-04 20:04:45 +00:00
Chandler Carruth dd0890619a Enable a couple of boring warnings. (#4018)
Just spotted these while looking at warnings that seem to fire on our
code are probably are things we'd fix if we saw them. None of these seem
important FWIW.

Also removes a redundant flag that is part of `-Wall`.

I have a follow-up for the high-value warning I spotted that motivated
me to look at all of this. But it's noisy so kept it as a separate PR.
2024-06-03 05:02:26 +00:00
Richard SmithandJon Ross-Perkins 3c01ee69ed Move information on the token associated with a parse node from the .def file into the typed node. (#4001)
Instead of tracking the token associated with a parse node in the `.def`
file macro, track it on the typed node instead. List the token as a
field inside the node structure to show the order of the token relative
to the other components of the grammar production, and to allow the
token index to be accessed when the node is extracted.

Remove the corresponding information from the `.def` file, leaving
behind just a list of parse node kinds in the majority of cases.

This also removes the checking of the token kind associated with a parse
node in the case where the parse node has errors. Previously we had a
flag on the node kind to indicate whether we should check this, but per
[discord
discussion](https://discord.com/channels/655572317891461132/655578254970716160/1246214418979881052),
we have decided to remove this.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-05-31 23:11:51 +00:00
Jon Ross-Perkins 16942e4d07 Move errs tie to LLVM to include test infrastructure. (#4006)
The lack of flushing can sometimes be observed when streaming test
output, particularly with --dump_output. My thought is moving it into
our LLVM init is probably reasonable, given the driver's already doing
this for similar reasons; this means it's used in tests through our
gtest_main.cpp in addition to the driver.
2024-05-30 21:08:44 +00:00
Jon Ross-Perkins cda5f66d22 Refactor NodeCategory to provide a class API (#4004)
Mirroring #4003 for NodeCategory.

Note we template a lot more on NodeCategory's enum, so this is a
slightly more awkward delta.

Also, switch from Enum in KeywordModifierSet to RawEnumType for
consistency with EnumBase. The templating on NodeCategory had me
thinking about that more.
2024-05-29 22:47:50 +00:00
Chandler Carruth 582b2c04dd Allow GoogleTest tests to locate their executable path. (#3992)
This also factors out the code for doing this location from the driver
to a tiny helper library.

The motivation for this is letting tests locate data files like the
prelude or other files needed by the toolchain more easily. A subsequent
PR will use this heavily in the Clang runner and related logic for
example.
2024-05-28 23:50:57 +00:00
Richard Smith 5c8fa6ad5c Replace FoldingSet with DenseMap for instruction canonicalization. (#3979)
Switch from recursing into non-canonical instruction fields to
separately canonicalizing those fields. This means we now form canonical
`InstBlockId`s, `TypeBlockId`s, `IntId`s, `FloatId`s, and `BindNameId`s
at least in the cases when they're referenced by a constant instruction.

This reduces the overall runtime for @chandlerc's 10MLoC example by
27.5% on my machine.
2024-05-23 00:48:49 +00:00
Chandler Carruth b473eac5bc Fix clang-tidy issues in //common. (#3962)
These likely predate the CI integration for `clang-tidy` runs.

Most of these seem good generally, even though I disabled some with
nolint comments. The multilevel pointer one seems almost like a bug in
the check to detect the specific case of `memcpy`, but otherwise seems
like a solid lint.
2024-05-20 23:29:07 +00:00
Jon Ross-Perkins b19a87642b Update LLVM (#3956)
Fix use of newly deprecated API
2024-05-17 19:36:21 +00:00
f5c01ee487 Add a custom main library for benchmarks. (#3872)
This largely reproduces the upstream one, but has a few advantages:

1) It initializes LLVM which can be important if we use the backends.
2) It parses commandline flags with Abseil allowing the use of Abseil
   flags in benchmarks.
3) It flips the default for tabular display of results which we use
   pretty heavily.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com>
2024-04-09 20:36:41 +00:00
Jon Ross-Perkins e66406ec93 Disable bugprone-macro-parentheses and let clang-format insert braces. (#3825)
-
[bugprone-macro-parentheses](https://clang.llvm.org/extra/clang-tidy/checks/bugprone/macro-parentheses.html)
-- this is just a false positive issue, I don't think it's helping us
catch bugs.
-
[InsertBraces](https://clang.llvm.org/docs/ClangFormatStyleOptions.html#insertbraces)
-- although there's a warning about this creating issues due to
incomplete semantic information, it seems to be happy with our code, and
allows clang-format to fix something that clang-tidy would otherwise
warn about.
2024-04-02 11:18:25 +00:00
Richard SmithandJon Ross-Perkins f9ce0b194d Defer parsing of method bodies until the end of a suitable enclosing scope. (#3832)
In parse, form a list of methods that are defined inline, tracking where
they start, where they end, and which other inline methods are nested
within them.

In check, when we reach an inline method body, skip it and add it to a
worklist to be processed later. We also track when we reach the start
and end of a context in which inline method bodies are deferred, so that
we know when to replay the bodies.

When suspending a function definition to be processed later, the
`DeclNameStack` entry is moved to separate storage, including popping
the corresponding scopes from the scope stack and removing the
corresponding lexical names from lexical lookup. Later, when we return
to the function and parse its definition, the `DeclNameStack` entry is
restored. The same is done when we reach the end of a nested context
that can have inline methods, so that we can reenter the nested scope
before processing its members.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-04-01 18:25:27 +00:00
Richard Smith f0e940ddfd Initial support for builtin functions. (#3803)
For now, a builtin function is defined by specifying a string literal
initializer in a function declaration:

```carbon
fn MyBuiltin(a: i32) -> i32 = "builtin.name";
```

End-to-end support is included for a sample `"int.add"` builtin
performing integer addition, covering constant evaluation and code
generation.

The implementation here needs substantial refactoring before we'll be
ready to start adding more builtins. That refactoring work will be
coming next. This change is aiming to checkpoint some incremental
progress.
2024-03-21 20:46:34 +00:00
Jon Ross-Perkins a3b1c433be Remove legacy repo_name settings (#3772)
I'd kept these in to separate the bazel module update from the BUILD
file changes, then forgot about it. I think all of these can be cleanly
removed now. I think it's something we should clean up for consistency
with the bazel central repository names; I think it's best to reduce
that divergence.

llvm_zlib and llvm_zstd remain because of how llvm depends on the
particular names.
2024-03-13 22:58:56 +00:00
cui fliter fea2651e7c chore: fix typos (#3738)
Signed-off-by: cui fliter <imcusg@gmail.com>
2024-03-04 16:27:31 +00:00
Jon Ross-Perkins 1974e44fd9 Rename factory functions from 'Create' to 'Make' (#3706)
Similar to #3705, we actually have a mix of `Make` and `Create` in
factory functions too, so this PR is normalizing on `Make`. It's
intended to be consistent with the naming choice for Carbon factory
functions.

Note, MakeSyntheticBlock is the only one I feel a little weird about
because llvm's own APIs use Create, and this is essentially wrapping
LLVM calls. But the flipside is it also feels like a vague line to draw,
when we also differ from LLVM coding style in other ways.
2024-02-14 18:26:56 +00:00
Calvin 9305d9888b Add help information printing for a specific subcommand (#3699)
The `command_library` provides functionality for end users to query
documentation and help information about commands through the command
line interface itself. At present, this is implemented as:

* A flag, `--help`, on every command
* A meta subcommand, `help`, on every command that has other subcommands

To match with common functionality in other command line interfaces,
this PR amends the `help` meta subcommand to accept an optional
positional string argument. This argument specifies which subcommand to
print help information for. When not specified, the present behavior is
maintained, printing the help information for its parent command.

A potentially unexpected consequence of this implementation is that you
can query for help on meta subcommands as well: `help help` is a valid
input which prints the help information about the `help` meta subcommand
(e.g., same for `version`). I don't see any harm in this behavior so I
chose not to explicitly prevent it, but I'm open to preventing it if
deemed unwanted. (I'm also happy to workshop the strings in this PR.)

Closes #3694.

#### Before

```
$ carbon help compile
```
```
ERROR: Found unexpected positional argument or subcommand: 'compile'
```

#### After

```
$ carbon help compile
```
```
Compile Carbon source code.

This subcommand runs the Carbon compiler over input source code, checking it for errors and producing the requested output.

Error messages are written to the standard error stream.

Different phases of the compiler can be selected to run, and intermediate state can be written to standard output as these phases progress.

Subcommand 'compile' usage:
  carbon [-v] compile [OPTIONS] <FILE>...

...
```
2024-02-11 09:30:11 +00:00
Chandler Carruth 2e236759ca Switch //common to use C++20 concepts. (#3665)
This removes the use of `enable_if` and tries to adopt concepts instead
of type traits when available.

The `ostream.h` change is a bit subtle as it adds a restriction not
previously in place -- that the stream is *contvertible* to
`std::ostream` as well as having it as a base class. This seems to match
the intent of the code.

The `hashing.h` code adds an implementation detail concept, and so I've
also clarified that the dispatch namespace is an internal one that isn't
part of the public API.
2024-01-30 18:14:07 +00:00
Chandler Carruth bf02d1f4b0 Remove headers marked as unused by ClangD. (#3661)
This required adding a few headers that were found transitively before,
but not too many. This is sadly a fairly manual process of opening every
file in my IDE, but I think I got everything in `//common` and
`//toolchain`.

There are a few cases where technically we don't need `foo.h` to be
included into `foo.cpp`, but I've forced those to stay with a pragma.

I've tried to catch the places where we can cut deps in Bazel as well,
but not sure I got all of those.

I had been noticing these in other PRs and it seemed better to isolate
the change.
2024-01-29 16:15:35 +00:00
Chandler Carruth ebbdc11877 Switch to a "better" multiplicative hash constant. (#3629)
Testing this hash function with representative hash table
implementations showed significant differences in quality between
different multiplicative hashing constants. The constants used and
documented were OK but had clear limitations when merely using a single
64-bit multiplication. However, a search uncovered (partly by luck)
a constant that has empirically been shown to be both significantly
better than other constants and generally not have problematic
weaknesses. There are still plenty of collisions for string keys and
heavily loaded hash tables of course, but no examples of severe outlier
collision rates as observed with all other constants we have tried.

There is a slightly longer comment explaining some of this context and
the other constants we have tried in the PR as well.

To this day, we still don't fully understand why the constant used here
behaves so much better than other constants we have tried, including all
of those we've found in other hashing algorithms.
2024-01-20 22:14:19 +00:00
Chandler Carruth 5f62cf752d Fix an oversight that dropped the buffer. (#3628)
This lost any seed or prior hashing done, which isn't good. I've added
some basic testing that would have caught this immediately.
2024-01-20 21:09:17 +00:00
Chandler Carruth 7e9760d9e4 Simplify the index & tag extraction API for hash codes. (#3627)
The fancier API ended up not being helpful and making it harder to
optimize a hash table implemented on top of this.
2024-01-20 20:55:15 +00:00
Jon Ross-Perkins a196b9840f Run clang-tidy on headers (#3572)
This patches bazel_clang_tidy handling of headers. I found an equivalent
change at https://github.com/erenon/bazel_clang_tidy/pull/13, but that
was [already
rejected](https://github.com/erenon/bazel_clang_tidy/pull/13#issuecomment-1047007424).
Per the criticism, this will result in redundant processing of headers.

The project instead uses `HeaderFilterRegex: ".*"`, but that results in
two problems:

1. When running with `-k`, errors are repeated when a header is included
more than once, which is common.
2. clang-tidy including errors from headers that are included from other
modules (e.g., abseil-cpp); filtering correctly is difficult.

Given the trade-offs and options (including forking), I thought patching
was preferable so long as it remains narrow.
2024-01-05 23:01:47 +00:00
Richard Smith 07efa026de Support defining instruction categories with a common representation (#3569)
Add a mechanism to define instruction categories, to support inspecting
the common representation of similar kinds of instruction. Use that
mechanism to make formatting of branch instructions slightly more
type-safe.

The idea here is to use the existing `Inst` mechanism for converting to
and from structs, extended to operate on a struct representing multiple
different kinds of instruction. In this case, the concrete kind of
instruction is stored in the struct in a `kind` field, rather than being
implied by the type.

Factored out of #3555 where this mechanism is used to provide a common
interface for runtime and symbolic name bindings.
2024-01-04 23:10:25 +00:00
Jon Ross-Perkins 379d776084 Add support for '--config=clang-tidy' (#3559)
This sets things up to use `bazel` to run `clang-tidy` using
https://github.com/erenon/bazel_clang_tidy.

I'm fixing issues outside of explorer, and disabling clang-tidy for
targets in explorer that have legacy issues. I was going to disable
clang-tidy for targets in explorer such as interpreter anyways, because
they're slow to parse, and just extended that to the currently failing
targets.
2024-01-04 18:09:52 +00:00
2e97f27b8d Typed wrappers around parse tree nodes (#3534)
These are intended to allow the structure of a parse tree node to be
described more precisely in code, to support these use cases:

- Automated checking that the parse tree conforms to the expected
structure. (Added to `Tree::Verify`.)
- Easier reading and understanding of the structure of the parse tree by
toolchain developers. (See `parse/typed_nodes.h`.)
- Easier navigation of the parse tree, for example for tooling uses and
for use when forming diagnostics.

On this last point, an object representing the file may be inspecting
using `Tree::ExtractFile`, as in:
```
auto file = tree->ExtractFile();
for (AnyDeclId decl_id : file.decls) {
  // `decl_id` is convertible to a `NodeId`.
  if (std::optional<FunctionDecl> fn_decl =
      tree->ExtractAs<FunctionDecl>(decl_id)) {
    // fn_decl->params is a `TuplePatternId` (which extends `NodeId`)
    // that is guaranteed to reference a `TuplePattern`.
    std::optional<TuplePattern> params = tree->Extract(fn_decl->params);
    // `params` has a value unless there was an error in that node.
  } else if (auto class_def = tree->ExtractAs<ClassDefinition>(decl_id)) {
    // ...
  }
}
```

The `Extract...` functions collect the child nodes into the typed parse
node's fields (internally using a `Tree::SiblingIterator`) for easy
access. However, this is not as fast as directly observing the tree
structure using the postorder strategy being used by the check stage.

These functions rely on using struct reflection on the typed parse node
definitions from `parse/typed_nodes.h` to get the expected structure of
child nodes and then populate them.

Note that validating these in `Tree::Verify` adds significant cost to
it, and is currently included in the parsing stage. Without this change,
a 10 mloc test case of lex & parse takes 4.129 s ± 0.041 s. With this
change, it takes 5.768 s ± 0.036 s.

This builds upon and completes #3393.

Co-authored-by: Richard Smith <richard@metafoo.co.uk>

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-12-22 22:14:11 +00:00
Jon Ross-Perkins bd0ef62a8f Fix clang-tidy issues in common and testing (#3470)
Choosing to make the constructor explicit in the test, rather than
NOLINT, because it seems to better reflect how our code is usually
written (and may be more likely to trip an issue).
2023-12-07 19:06:29 +00:00
Richard SmithandJon Ross-Perkins bf8697113a Move llvm::Initialize* calls to main. (#3449)
Per their documentation, the `llvm::Initialize*` functions are only
supposed to be called by the main program, not by a library like
toolchain/codegen. Fixes a hang due to a data race in multithreaded
autoupdate.

Add a utility class `Carbon::InitLLVM` to do the common LLVM
initialization shared by all Carbon tools, optionally including
initializing the LLVM targets. Because the LLVM targets add a lot of
binary size, only initialize them for binaries that opt in by depending
on a new target `//common:all_llvm_targets`.

Also fix `//explorer:file_test` and `//explorer:file_test.trace` to
share a binary rather than linking an identical binary twice.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-12-07 01:45:47 +00:00
Jon Ross-Perkins 35d15a390c Remove nodiscard uses. (#3418)
Per [#toolchain
discussion](https://discord.com/channels/655572317891461132/655578254970716160/1176632520834560211)

We'd at one point been trying to put `[[nodiscard]]` everywhere, but
then we stopped because it had felt verbose without finding many issues
(plus, people plain forgot to add it). Some history in #888.

Since newer code gets added without it, we now have code like:

```
  auto GetLineInfo(Line line) -> LineInfo&;
  [[nodiscard]] auto GetLineInfo(Line line) const -> const LineInfo&;
  auto AddLine(LineInfo info) -> Line;
  auto GetTokenInfo(Token token) -> TokenInfo&;
  [[nodiscard]] auto GetTokenInfo(Token token) const -> const TokenInfo&;
  auto AddToken(TokenInfo info) -> Token;
  [[nodiscard]] auto GetTokenPrintWidths(Token token) const -> PrintWidths;
```

Here, the lack of `[[nodiscard]]` doesn't mean anything: for example,
`GetLineInfo` should not have its result discarded if it's called. But
the mix could be confusing for readers.

As a resolution, remove the attribute. `[[nodiscard]]` should be treated
like other attributes going forward, which essentially means "avoid in
general, add a comment to explain why the attribute is needed" rather
than use-as-default.
2023-11-28 18:46:19 +00:00
f59a6cdbdd Introduce a Carbon hashing framework. (#3327)
# Overview

This is a latency-optimized hashing framework based on Abseil's and
others. At it's core it uses both a normal 64-bit multiply as well as a
64-bit multiply capturing both low and high 64-bit components of the
result and XOR-ing them together. These are the primitives used in
FxHash and Abseil respectively, although they both appear in others.

The implementation has been *substantially* optimized for short inputs
and latency over quality. As a result, this function does not remotely
pass the SMHasher quality tests. However, basic collisions are rare, and
I've included a small subset of the SMHasher collision testing directly
to make sure the quality doesn't slip too far inadvertently.

The customization framework is roughly similar to Abseil's and LLVM's
but has been simplified significantly, inspired in some respects by the
AHash API design and in others by my experience of all performance
sensitive hashing implementations needing to work at a very low level to
hit their performance targets. The abstractions are stripped down to
facilitate this.

# Details of the performance optimization

This function is 2x - 4x faster than LLVM's on small inputs, and up to
2x faster than Abseil. Significant effort has gone into optimizing short
strings in particular compared to Abseil.

Small integer and pointer hashing is also faster than Abseil's by
leveraging a lower quality 64-bit multiply in some cases inspired by
FxHash. One consequence is that this routine is particulary fast for
32-bit integers.

The short string improvements largely come from packing more of the
bytes of string into as few multiplies as possible. While this fails to
mix the bits sufficient to hit SMHasher's strict avalanche criteria and
does leave some collision windows, it provides dramatic latency
improvements. Some of these techniques come from Abseil's own bulk
hashing routine but re-applied here. Others are novel, for example using
small sizes to sample nicely uniform random data to efficiently handle
the very small number of bits of data that need to be hashed.

The other observed improvement is diligent handling of pairs and tuples
and fairly aggressively turning things into integers. Some of the
comparisons with Abseil aren't realistic as the Abseil hash table does
some of these mappings before hashing. I've done this directly in the
hash function as that seems cleaner.

For long strings, the performance is comparable or a bit better than
Abseil, and significantly better than LLVM's hash function.

Overall, for short inputs this is hoped to be the fastest hash function
that still gets "just enough" mixing for modern hash tables to perform
well.

# Details of the quality vs. latency tradeoff

A key insight is that modern hash tables don't need especially high
quality hash functions, but do benefit from something beyond the
identify function. That isn't the target of SMHasher or other quality
assessing tools and has resulted in unnecessarily aggressive hashing for
any functions actually evaluated against it. Many hash functions turn
off the high quality implementations evaluated with SMHasher for integer
or pointer keys to recover latency & performance (AHash for example),
but the same performance-oriented design applies beyond these narrow
types, for example for short strings.

However, a consequence is that there are serious limits to the quality
of the hash function. The avalanche test is failed hilariously, etc.,
but in the exact same ways as Abseil itself fails it for integer keys.
There are also real collisions spaces. For example, for 16-byte strings,
there is one 64-bit value for the first 8 bytes that will have the same
hash regardless of the other 8 bytes of the string. Some minor effort is
taken to make this pattern unlikely to be a practical problem, but it is
a clear theoretical weakness.

It also means that this hash function couldn't be further from providing
any hash-flooding DoS attack protection -- I expect it to be trivially
easy to attack in this way by a motivated adversary. Defending against
these attacks is defined as out-of-scope, in large part because even
attempts that have made a compelling effort to address these issues such
as HighwayHash have found serious limits. Instead, this takes a
principled position that any such defense should be provided entirely at
the data structure level with a strong worst-case bound rather than
through strengthening the hash function.

# Future work

A subsequent PR will introduce a hash table inspired very heavily by the
design of Abseil's "SwissTable" and using this hash function. The goal
is to provide a significant improvement to hot hash tables such as the
identifier table in the lexer of Carbon's toolchain.

# Detailed benchmark data

The benchmarks introduced are heavily inspired by the latency
benchmarking of hash functions in Abseil. I've adapted them to fit
better into Carbon's coding style and to try to have more stable results
with broader coverage of types and string sizes.

Running the benchmarks directly gives horizontal comparisons across
different hash functions. That can be hard to read, so here is *just*
the newly introduced hash function benchmark results on an AMD server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          3.11ns ± 1%
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         3.11ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      4.11ns ± 1%
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         3.12ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    4.13ns ± 1%
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         3.16ns ± 2%
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             3.16ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    4.03ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    4.04ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    4.34ns ± 2%
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        4.04ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        4.34ns ± 2%
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            4.34ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        4.33ns ± 1%
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        1.95ns ± 4%
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        1.70ns ± 3%
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       3.52ns ± 3%
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       4.46ns ± 2%
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       7.69ns ± 1%
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      14.8ns ± 1%
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      21.5ns ± 1%
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     34.6ns ± 0%
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     63.1ns ± 1%
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      118ns ± 1%
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      225ns ± 1%
```

And on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          5.28ns ± 0%
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         5.29ns ± 0%
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      7.02ns ± 0%
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         5.34ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    7.07ns ± 4%
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         5.36ns ± 2%
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             5.36ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    7.19ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    7.29ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    7.31ns ± 4%
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        7.29ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        7.31ns ± 4%
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        8.69ns ± 3%
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        2.64ns ± 2%
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        2.90ns ± 4%
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       6.14ns ± 1%
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       8.27ns ± 1%
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       13.8ns ± 0%
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      31.2ns ± 0%
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      49.9ns ± 0%
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     86.9ns ± 0%
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      163ns ± 0%
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      312ns ± 0%
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      610ns ± 0%
```

I don't have the same nice statistical multi-run error bars, but one run
from my M1 MacBook:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                             3.89 ns
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                            3.87 ns
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>         4.39 ns
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                            3.93 ns
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>       4.98 ns
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                            3.87 ns
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                                3.87 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>       4.86 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>       4.43 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>       4.41 ns
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>           4.44 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>           4.69 ns
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                         4.33 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>       4.38 ns
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>               4.34 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>           4.35 ns
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>           4.38 ns
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                           1.15 ns
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                          0.973 ns
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                          3.03 ns
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                          3.97 ns
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                          6.64 ns
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                         12.5 ns
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                         17.9 ns
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                        27.9 ns
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                        48.1 ns
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                        87.3 ns
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                         166 ns
```

And here I have internally replaced the Carbon hash function with
Abseil's hash function for "before" and then restored it in the "after"
and computed the delta for each benchmark. This basically shows the
speed-up (lower time -> lower latency -> speed-up -> good) over Abseil
on an AMD server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          4.00ns ± 1%  3.10ns ± 0%  -22.45%  (p=0.000 n=20+15)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         4.01ns ± 1%  3.10ns ± 1%  -22.64%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      6.25ns ± 1%  4.10ns ± 1%  -34.30%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         4.02ns ± 1%  3.12ns ± 1%  -22.50%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    6.25ns ± 1%  4.11ns ± 1%  -34.20%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         4.03ns ± 1%  3.14ns ± 1%  -22.17%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             5.95ns ± 1%  3.14ns ± 1%  -47.24%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    6.04ns ± 1%  4.01ns ± 1%  -33.64%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    5.96ns ± 1%  4.02ns ± 1%  -32.51%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    5.93ns ± 1%  4.30ns ± 1%  -27.56%  (p=0.000 n=20+17)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        7.97ns ± 1%  4.02ns ± 1%  -49.50%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        7.98ns ± 1%  4.32ns ± 1%  -45.88%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      4.40ns ± 2%  4.32ns ± 1%   -1.81%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    5.94ns ± 1%  4.32ns ± 1%  -27.25%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            10.0ns ± 1%   4.3ns ± 1%  -56.56%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        8.04ns ± 1%  4.32ns ± 1%  -46.29%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        7.95ns ± 1%  4.33ns ± 1%  -45.59%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        3.28ns ± 3%  1.93ns ± 4%  -41.19%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        3.05ns ± 3%  1.69ns ± 4%  -44.52%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       5.88ns ± 2%  3.50ns ± 3%  -40.42%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       8.92ns ± 1%  4.44ns ± 2%  -50.22%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       12.0ns ± 1%   7.7ns ± 1%  -36.16%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      18.8ns ± 0%  14.7ns ± 1%  -21.73%  (p=0.000 n=17+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      25.5ns ± 1%  21.4ns ± 1%  -16.18%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     38.7ns ± 2%  34.5ns ± 1%  -10.78%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     69.7ns ± 1%  62.8ns ± 1%   -9.88%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      130ns ± 1%   117ns ± 1%   -9.45%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      244ns ± 0%   225ns ± 1%   -8.11%  (p=0.000 n=17+20)
```

... and on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          6.48ns ± 1%  5.28ns ± 0%  -18.62%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         7.40ns ± 1%  5.29ns ± 1%  -28.45%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      10.4ns ± 0%   7.0ns ± 0%  -32.34%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         6.56ns ± 1%  5.32ns ± 1%  -18.95%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    10.8ns ± 2%   7.0ns ± 1%  -34.89%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         6.71ns ± 3%  5.38ns ± 2%  -19.84%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             10.3ns ± 3%   5.4ns ± 2%  -47.67%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    10.9ns ± 2%   7.2ns ± 4%  -33.67%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    10.7ns ± 4%   7.3ns ± 4%  -31.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    10.5ns ± 3%   7.3ns ± 4%  -30.71%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        14.1ns ± 3%   7.3ns ± 4%  -48.32%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        14.0ns ± 1%   7.3ns ± 4%  -47.95%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      9.41ns ± 4%  8.68ns ± 4%   -7.71%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    12.2ns ± 2%   8.7ns ± 4%  -28.81%  (p=0.000 n=18+19)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            18.9ns ± 2%   8.7ns ± 4%  -54.17%  (p=0.000 n=17+19)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        15.6ns ± 2%   8.7ns ± 4%  -44.37%  (p=0.000 n=17+19)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        15.5ns ± 2%   8.7ns ± 4%  -44.08%  (p=0.000 n=18+19)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        5.89ns ± 2%  2.64ns ± 3%  -55.26%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        5.73ns ± 3%  2.88ns ± 3%  -49.71%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       10.1ns ± 1%   6.1ns ± 2%  -39.00%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       15.7ns ± 0%   8.3ns ± 1%  -47.27%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       21.2ns ± 0%  13.8ns ± 0%  -34.81%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      37.9ns ± 0%  31.2ns ± 0%  -17.77%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      56.8ns ± 0%  49.8ns ± 0%  -12.21%  (p=0.000 n=20+18)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     93.8ns ± 0%  86.9ns ± 0%   -7.38%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      174ns ± 0%   163ns ± 0%   -6.03%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      330ns ± 0%   312ns ± 0%   -5.25%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      641ns ± 0%   610ns ± 0%   -4.79%  (p=0.000 n=19+19)
```

This is the same as the above delta comparison, but with the "before"
being LLVM's hash function:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          6.85ns ± 1%  3.10ns ± 1%  -54.78%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         6.85ns ± 1%  3.10ns ± 1%  -54.78%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      6.25ns ± 1%  4.09ns ± 1%  -34.58%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         6.87ns ± 1%  3.12ns ± 2%  -54.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    7.35ns ± 1%  4.10ns ± 1%  -44.20%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         7.34ns ± 1%  3.13ns ± 1%  -57.34%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             7.33ns ± 1%  3.13ns ± 2%  -57.27%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    7.27ns ± 1%  3.99ns ± 1%  -45.12%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    14.5ns ± 1%   4.0ns ± 1%  -72.23%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    14.6ns ± 1%   4.3ns ± 2%  -70.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        14.5ns ± 1%   4.0ns ± 1%  -72.21%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        14.6ns ± 1%   4.3ns ± 1%  -70.46%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      7.31ns ± 1%  4.33ns ± 1%  -40.81%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    7.78ns ± 1%  4.32ns ± 1%  -44.45%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            7.78ns ± 2%  4.33ns ± 1%  -44.42%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        7.62ns ± 1%  4.32ns ± 1%  -43.24%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        7.77ns ± 1%  4.33ns ± 1%  -44.34%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        8.15ns ± 3%  1.94ns ± 5%  -76.16%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        7.02ns ± 3%  1.69ns ± 4%  -75.94%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       7.83ns ± 2%  3.50ns ± 3%  -55.34%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       9.17ns ± 1%  4.43ns ± 2%  -51.65%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       11.3ns ± 1%   7.6ns ± 1%  -32.04%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      23.0ns ± 1%  14.7ns ± 1%  -36.14%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      32.9ns ± 0%  21.4ns ± 1%  -34.96%  (p=0.000 n=17+19)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     52.2ns ± 1%  34.4ns ± 1%  -34.01%  (p=0.000 n=19+18)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     92.1ns ± 1%  62.8ns ± 1%  -31.82%  (p=0.000 n=19+19)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      169ns ± 1%   117ns ± 1%  -30.53%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      319ns ± 1%   224ns ± 1%  -29.78%  (p=0.000 n=20+18)
```

... and on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          8.38ns ± 0%  5.27ns ± 0%  -37.04%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         8.39ns ± 1%  5.28ns ± 0%  -37.01%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      8.07ns ± 0%  7.02ns ± 0%  -13.10%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         8.48ns ± 1%  5.32ns ± 1%  -37.25%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    9.34ns ± 2%  7.09ns ± 2%  -24.14%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         9.76ns ± 3%  5.37ns ± 2%  -44.98%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             9.76ns ± 3%  5.37ns ± 2%  -44.98%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    10.1ns ± 2%   7.2ns ± 3%  -29.36%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    11.9ns ± 2%   7.3ns ± 4%  -38.68%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    11.3ns ± 2%   7.3ns ± 4%  -35.16%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        11.9ns ± 2%   7.3ns ± 4%  -38.68%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        11.3ns ± 2%   7.3ns ± 4%  -35.16%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      10.3ns ± 2%   8.7ns ± 3%  -15.81%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        9.39ns ± 2%  2.66ns ± 3%  -71.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        10.7ns ± 3%   2.9ns ± 3%  -72.97%  (p=0.000 n=19+18)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       11.8ns ± 1%   6.1ns ± 2%  -47.75%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       13.9ns ± 1%   8.3ns ± 1%  -40.71%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       16.8ns ± 1%  13.8ns ± 0%  -17.83%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      31.7ns ± 1%  31.2ns ± 0%   -1.76%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      43.5ns ± 0%  49.8ns ± 0%  +14.56%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     66.2ns ± 0%  86.9ns ± 0%  +31.39%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      112ns ± 0%   163ns ± 0%  +46.09%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      201ns ± 0%   312ns ± 0%  +55.49%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      379ns ± 0%   610ns ± 0%  +61.08%  (p=0.000 n=20+20)
```

Note that there is a significant regression on long strings compared to
LLVM's hash function on the ARM server I have access to. This doesn't
show up on the M1 at all, and is likely specific to inadequate
throughput for the 64-bit multiply operations. This seems fine as a) our
priority is for short strings, and b) the M1 and other ARM CPUs are
likely to improve here over time given the prevalent use of this core
technique. For example, Abseil's current hash algorithm has the same
long-string behavior (and performance bottleneck) on this server.

---------

Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Co-authored-by: Geoff Romer <gromer@google.com>
2023-11-20 19:59:06 +00:00
josh11b 5020fdb3be Use abbreviation "decl" instead of "declaration" (#3382)
Part of switching to the [abbreviations we've decided to
use](https://docs.google.com/document/d/1RRYMm42osyqhI2LyjrjockYCutQ5dOf8Abu50kTrkX0/edit?resourcekey=0-kHyqOESbOHmzZphUbtLrTw#heading=h.pph7i5m5un7q).

I will rename files in a follow-up PR.
2023-11-10 10:43:25 +00:00
josh11bandJon Ross-Perkins 4c09a37448 Expand comments in enum_base.h (#3315)
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-10-27 20:47:47 +00:00
Jon Ross-Perkins 3af7eb2672 Refactor YAML handling to use the llvm::yaml API. (#3337)
Provides an adapter for the llvm::yaml API because it otherwise needs a
bunch of const/non-const definitions, and the traits are difficult to
diagnose issues with. The current approach is pretty simple to use, even
if it's not super efficient (which, yaml output is more of a debugging
thing so I'm not really expecting it to be an issue).

Changes the format of yaml output to provide more index information,
just as reminders when seeing something like `node+0`. Note this would
create more churn in deltas if we were reliant on the output yaml in
tests, but we aren't so it should be okay.
2023-10-26 18:50:30 +00:00
Jonathan B. CoeandChandler Carruth 8c28a0494e Add size="small" to test targets where advised (#3326)
Running `bazel test //...` reported:

```
Test execution time outside of range for MODERATE tests.
Consider setting timeout="short" or size="small".
```

This change adds size="small" to avoid such warnings being reported.

---------

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-10-24 06:52:00 +00:00
Richard Smith c7a9e29a89 Add typed nodes to SemIR. (#3280)
Replace `SemIR::Node::GetAsFoo` and `SemIR::Node::Foo::Make` with
`SemIR::Foo` class that represents a particular kind of node, with named
fields.

Rename `SemIR::IntegerLiteral` and `SemIR::RealLiteral` to
`IntegerValue` / `RealValue` to better reflect their purpose and avoid a
name collision with the corresponding `SemIR` node kinds.

Remove `NodeKind::Invalid` and the `SemIR::Node` default constructor
entirely, as they were not used for anything.
2023-10-11 05:39:59 +00:00
Chandler Carruth 4596cd230d Avoid building the non-test file group in :all. (#3191)
This file group exists to allow a `genquery` rule and a Python test to
verify our non-test dependency graph. We don't actually need to build
the binaries in the file group as part of that. The `genquery` rule
seems to do the right thing -- building it directly doesn't cause the
binaries in the group to be built. But without a manual tag, the group
itself is part of `:all` and thus part of `//...` and part of the rules
that will be built even with PR #3106. A consequence is that any change
to the toolchain causes several other binaries to be built as well
because this file group is in the impacted set. Making it manual should
avoid all of this, and without breaking the actual use from `genquery`.

For example, before this change, in a fully cached build after a `bazel
clean`:
```
> bazel test //bazel/check_deps:all
INFO: Invocation ID: 2d83ebee-4c00-425d-be33-23f42b079614
INFO: Analyzed 3 targets (103 packages loaded, 7137 targets configured).
INFO: Found 2 targets and 1 test target...
INFO: Elapsed time: 4.081s, Critical Path: 2.61s
INFO: 3111 processes: 2796 disk cache hit, 315 internal.
INFO: Build completed successfully, 3111 total actions
```

After this change:
```
> bazel test //bazel/check_deps:al
INFO: Invocation ID: c94089e8-a420-4d3c-9902-134e6b55b297
INFO: Analyzed 2 targets (92 packages loaded, 568 targets configured).
INFO: Found 1 target and 1 test target...
INFO: Elapsed time: 0.700s, Critical Path: 0.01s
INFO: 7 processes: 2 disk cache hit, 5 internal.
INFO: Build completed successfully, 7 total actions
```

While here, re-generate the file group, and fix several issues it
uncovers: mark test utilities as `testonly` and update our LLVM package
allowlist to include `clangd`'s package.
2023-10-05 01:30:41 +00:00
Geoff Romer e3d3122f1d Move tests to the namespace of the code under test (#3244)
Rationale: this convention avoids forcing closely-related code to be far
apart in the namespace hierarchy, and vice versa. By the same token, it
makes the namespace hierarchy more consistent with the directory
hierarchy.
2023-09-18 22:05:14 +00:00
Jon Ross-Perkins 53af8f04b2 Provide a Printable CRTP parent to replace HasPrintable templates. (#3166)
With the toolchain splitting namespaces, ostream.h's `operator<<`
templates aren't reliably found with name lookup, likely due to the loss
of associated namespaces (zygoloid commented on this at
https://github.com/carbon-language/carbon-lang/pull/3161#discussion_r1307941999).
This is especially a barrier to moving the lex files into `Carbon::Lex`;
versus other parts of the toolchain, they contain more printable types
which are used cross-namespace, including `Carbon::Testing`. As a
consequence, I'm looking at migrating ostream.h to a more reliable
approach that doesn't rely as much on everything being in the `Carbon`
namespace.
2023-08-30 21:32:19 +00:00
Jon Ross-Perkins f63834c71d Move PrintAsID into explorer. (#3163)
#2569 added PrintAsID to //common/ostream.h, but given it's
explorer-specific behavior, I don't think it's the right home for it.

Noticed this while pondering better ostream interfaces.
2023-08-28 23:39:56 +00:00
Jon Ross-PerkinsandRichard Smith 7157445f97 Set up a 'Parse' namespace. (#3161)
Continuing on #3070.

I moved ParseTree::Node to just Parse::Node, versus Parse::Tree::Node.
Other name changes are just removing "Parse" or "Parser" prefixes.

In EnumBase, I'm directly defining operator<< because the ostream.h
approach just isn't working, not for either of Parse::State nor
Parse::NodeKind. Errors look like:

```toolchain/parser/parser_context.cpp:449:34: error: invalid operands to binary expression ('llvm::raw_ostream' and 'const Carbon::Parse::State')
    output << "\t" << i << ".\t" << entry.state;
    ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ^  ~~~~~~~~~~~
```

The expected template in `Carbon::` is not in the error list; I only see
the:

```
./common/ostream.h:112:6: note: candidate template ignored: requirement 'std::is_base_of_v<std::ostream, llvm::raw_ostream>' was not satisfied [with S = llvm::raw_ostream, T = Carbon::Parse::State]
auto operator<<(S& standard_out, const T& value) -> S& {
     ^
```

I'm still prodding at this, but not seeing an obvious fix.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-08-28 21:58:38 +00:00
Jon Ross-Perkins 29b6399e4f Modify EnumBase to better support the namespacing of toolchain (#3156)
The different approach to Names avoids the issues with trying to define
a static member (or also member function) of the templated instance of
Carbon::Internal::EnumBase from a non-enclosing namespace such as
Carbon::SemIR.

Note, I'm trying to do this from the cpp file. An alternative might be
to do `inline constexpr llvm::StringLiteral Names[]` in the .h file, but
I think concerns had been raised about that needing deduplication.
2023-08-26 00:17:25 +00:00
josh11b f790a27ace Fix out-of-bounds array access for option characters >= 128 (#3151)
Bug found by fuzzing.
2023-08-24 16:39:49 +00:00
Jon Ross-Perkins 67da700dd5 Split Semantics into Check and SemIR namespaces (#3138)
Splits IR files into SemIR, and logic files into Check. These will be
split into separate directories as part of a later move; the namespaces
are being done first in order to vet the switch, and hopefully make
conflicts a little easier to manage due to the substantial renames.

A lot of this is just automated removal of Semantics prefixes from
names, adding namespace references where needed. A few special-cases
are:

- SemanticsIR -> SemIR::File
- A few things were discussed, like Unit, CompileUnit, or CompiledUnit.
Unit was too vague for chandlerc, and I thought CompileUnit might lead
to incorrect inferences (CompilationUnit would be more precise, but
typically written as SemIR::CompilationUnit which is pretty long). File
seemed to be a short name that we could agree on.
- SemanticsIRFormatter -> SemIR::Formatter
- FormatSemanticsIR -> SemIR::FormatFile
- SemanticsFileTest -> CheckFileTest
- It remains in the Testing namespace, where just "FileTest" might be
too broad a name.
- SemanticsDeclarationNameStack::Context ->
Check::DeclarationNameStack::NameContext
  - This avoids a Check::Context name shadowing.

Changes check_internal.h to include ostream.h to improve finding of
Print/operator<< (otherwise it didn't compile).

This is part of #3070
2023-08-23 22:51:23 +00:00
Jon Ross-Perkins 605763d62d Add lint fixes to the buildifier setup. (#3109)
The main motivation for this is to get python loads in using the
`native-py` lint fix. However, enabling that made me wonder, maybe we
should fix in general?

`native-cc` is delayed, but not wholly cancelled (and `native-py`
picking up might indicate `native-cc` won't be too far behind). There's
also some automated fixes for `.append` and dict sorting -- this felt
okay to me, maybe not something to eagerly add but probably not worth
stopping buildifier from fixing (I've noticed the warnings in the past
and had been ignoring them).

Running everything does mean that load orders are sorted automatically
now, which I think is a positive. Most generally, I think these fixes
aren't _harmful_, and having them done automatically seems beneficial:
my biggest concern about `native-py` and `native-cc` was actually that
regressions wouldn't be caught, but this addresses that issue
automatically.
2023-08-22 21:01:42 +00:00