Commit Graph
13 Commits
Author SHA1 Message Date
Jon Ross-PerkinsandGeoff Romer 4c4c4a4d2c Add RawStringOstream for slightly simpler streaming to strings (#4817)
This adds a RawStringOstream. Versus TestRawOstream, which is
consolidated over to RawStringOstream, it uses a string for storage
instead of a vector, mainly to support move-to-string semantics. Versus
llvm::raw_string_ostream, it owns the string and supports pwrite (which
is needed for driver and its fd_ostream compatibility requirement).

This converts most uses of llvm::raw_string_ostream, leaving behind a
few in InstNamer that explicitly cannot own the string, such as:

```
     llvm::raw_string_ostream(name)
          << "_" << tree.tokens().GetColumnNumber(token);
```

I have this as its own library so that it can use CHECK.

Yes this doesn't save much code, but it's code we repeatedly write.

---------

Co-authored-by: Geoff Romer <gromer@google.com>
2025-01-18 01:11:44 +00:00
Jon Ross-Perkins 61c0a8b676 Make more use of llvm STLExtras (#4668)
This is essentially the result of looking at `.begin()` uses. We also
frequently do `std::shuffle`, but unfortunately STLExtras doesn't
provide a wrapper for that.
2024-12-11 18:16:38 +00:00
4845f40dff Switch CARBON_CHECK to a format string API (#4285)
This switches `DCHECK` and `FATAL` as well.

The goal is to reduce the code size impact of these assertions so that
we can keep more of them enabled. Currently, the largest cost I see from
`CHECK` is not the actual check or the cold code itself, but actually
the failure to inline trivial functions due to the presence of the cold
code. This means that our goal isn't to reduce apparent code size in the
final binary but the LLVM IR cost assessed for these routines in the
inliner, which closely correlates with code size but is a bit different.

As discussed in #4283, experimentation shows that a single function call
with a minimal number of arguments is the lowest cost model for these.
This is easily achieved with a format-string API that internally uses
`llvm::formatv`. This PR is essentially the `CHECK` version of #4283.

However, the check macros are substantially harder to make work with
both format strings and streaming because they also take a condition.
Also, unexpectedly, I was very successful at devising a regular
expression based automated rewrite from the streaming to the format
string form with only low 10s of manual fixes. This includes compacting
strings broken up across lines, etc. Given how well that went, I've
prepared this PR which just directly switches to the format string API
and migrate everything to use it.

One nice side-effect is that the format string approach ends up greatly
simplifying the implementation here as well.

This is ... *shockingly* effective. Parsing speeds up by more than 3%
with just this change. And checking speeds up by **8%** with this change
alone:
```
BM_CompileAPIFileDenseDecls<Phase::Parse>/256      86.3µs ± 1%  82.9µs ± 1%  -3.94%  (p=0.000 n=17+19)
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024      431µs ± 1%   415µs ± 1%  -3.76%  (p=0.000 n=18+19)
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096     1.77ms ± 1%  1.71ms ± 1%  -3.18%  (p=0.000 n=18+19)
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384    7.44ms ± 1%  7.17ms ± 2%  -3.56%  (p=0.000 n=18+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    30.7ms ± 1%  29.7ms ± 1%  -3.15%  (p=0.000 n=18+20)
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144    131ms ± 1%   127ms ± 1%  -2.81%  (p=0.000 n=18+18)
BM_CompileAPIFileDenseDecls<Phase::Check>/256       878µs ± 2%   800µs ± 1%  -8.91%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/1024     1.88ms ± 2%  1.72ms ± 1%  -8.56%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/4096     5.78ms ± 2%  5.28ms ± 1%  -8.70%  (p=0.000 n=20+18)
BM_CompileAPIFileDenseDecls<Phase::Check>/16384    21.9ms ± 1%  20.1ms ± 1%  -8.02%  (p=0.000 n=18+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/65536    90.4ms ± 2%  83.1ms ± 1%  -8.04%  (p=0.000 n=19+20)
BM_CompileAPIFileDenseDecls<Phase::Check>/262144    381ms ± 2%   352ms ± 1%  -7.79%  (p=0.000 n=19+19)
```

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
2024-09-12 16:42:08 +00:00
David Blaikie 5a11048c34 Remove some explicit (Mutable)ArrayRef constructions (#4238)
Rely on implicit conversion in call sites and initialization.

Removing the explicit conversions is only code simplication.
Moving from `auto x = Y(z)` to `Y x = z;` helps ensure that only
implicit constructors/conversions are happening (whereas the prior
syntax allows explicit conversions) which can help with readability
since implicit conversions are generally "less
complex"/risky/attention-requiring.
2024-08-22 19:40:54 +00:00
bf736e6b03 A collection of hashing improvements from using hashtables. (#4094)
LLVM's `APInt` and `APFloat` need specialized handling to be used
effectively in hashtables. We can't inject overrides into LLVM so we
need to handle them in our hashing routine.

There were also problematic limits on hashing pairs and tuples. First,
the unique-object-representation hashing of pairs was more restricted
than tuples which was a problematic asymmetry and isn't needed. But the
larger issue is that we didn't support recursively hashing when
necessary. That requires a careful predicate to avoid infinite recursion
but lets us handle important use cases for hashtables with a tuple as a
key.

Also added support for hashing arrays that recurse in addition to arrays
where we can hash the raw storage, and added overloads to redirect to
common array handling from various array-like types.

Last but not least, re-worked the constraint model for hashing as raw
data to not override custom hashing functions.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com>
2024-07-02 20:53:48 +00:00
Chandler Carruthandjosh11b 21a81bc59e Introduce custom hash table data structures. (#3940)
The hash table design is heavily based on Abseil's ["Swiss
Tables"][swiss-tables] design. It uses an array of bytes storing
metadata about each entry and an array of entries where each is a pair
of key and value. The metadata byte consists of 7-bits of hash of the
key (distinct from the bits used to index the table), and one bit
indicating the presence of a special entry -- either empty or deleted.

[swiss-tables]: https://abseil.io/about/design/swisstables

There are a large range of optimizations and other nuanced aspects of
this hash table design and implementation, a good point to understand
that context is `raw_hashtable.h` which has an overview of the design
and references to various other files for relevant details.

---------

Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com>
2024-06-08 01:50:02 +00:00
Richard Smith 5c8fa6ad5c Replace FoldingSet with DenseMap for instruction canonicalization. (#3979)
Switch from recursing into non-canonical instruction fields to
separately canonicalizing those fields. This means we now form canonical
`InstBlockId`s, `TypeBlockId`s, `IntId`s, `FloatId`s, and `BindNameId`s
at least in the cases when they're referenced by a constant instruction.

This reduces the overall runtime for @chandlerc's 10MLoC example by
27.5% on my machine.
2024-05-23 00:48:49 +00:00
Chandler Carruth b473eac5bc Fix clang-tidy issues in //common. (#3962)
These likely predate the CI integration for `clang-tidy` runs.

Most of these seem good generally, even though I disabled some with
nolint comments. The multilevel pointer one seems almost like a bug in
the check to detect the specific case of `memcpy`, but otherwise seems
like a solid lint.
2024-05-20 23:29:07 +00:00
Chandler Carruth 2e236759ca Switch //common to use C++20 concepts. (#3665)
This removes the use of `enable_if` and tries to adopt concepts instead
of type traits when available.

The `ostream.h` change is a bit subtle as it adds a restriction not
previously in place -- that the stream is *contvertible* to
`std::ostream` as well as having it as a base class. This seems to match
the intent of the code.

The `hashing.h` code adds an implementation detail concept, and so I've
also clarified that the dispatch namespace is an internal one that isn't
part of the public API.
2024-01-30 18:14:07 +00:00
Chandler Carruth ebbdc11877 Switch to a "better" multiplicative hash constant. (#3629)
Testing this hash function with representative hash table
implementations showed significant differences in quality between
different multiplicative hashing constants. The constants used and
documented were OK but had clear limitations when merely using a single
64-bit multiplication. However, a search uncovered (partly by luck)
a constant that has empirically been shown to be both significantly
better than other constants and generally not have problematic
weaknesses. There are still plenty of collisions for string keys and
heavily loaded hash tables of course, but no examples of severe outlier
collision rates as observed with all other constants we have tried.

There is a slightly longer comment explaining some of this context and
the other constants we have tried in the PR as well.

To this day, we still don't fully understand why the constant used here
behaves so much better than other constants we have tried, including all
of those we've found in other hashing algorithms.
2024-01-20 22:14:19 +00:00
Chandler Carruth 5f62cf752d Fix an oversight that dropped the buffer. (#3628)
This lost any seed or prior hashing done, which isn't good. I've added
some basic testing that would have caught this immediately.
2024-01-20 21:09:17 +00:00
Chandler Carruth 7e9760d9e4 Simplify the index & tag extraction API for hash codes. (#3627)
The fancier API ended up not being helpful and making it harder to
optimize a hash table implemented on top of this.
2024-01-20 20:55:15 +00:00
f59a6cdbdd Introduce a Carbon hashing framework. (#3327)
# Overview

This is a latency-optimized hashing framework based on Abseil's and
others. At it's core it uses both a normal 64-bit multiply as well as a
64-bit multiply capturing both low and high 64-bit components of the
result and XOR-ing them together. These are the primitives used in
FxHash and Abseil respectively, although they both appear in others.

The implementation has been *substantially* optimized for short inputs
and latency over quality. As a result, this function does not remotely
pass the SMHasher quality tests. However, basic collisions are rare, and
I've included a small subset of the SMHasher collision testing directly
to make sure the quality doesn't slip too far inadvertently.

The customization framework is roughly similar to Abseil's and LLVM's
but has been simplified significantly, inspired in some respects by the
AHash API design and in others by my experience of all performance
sensitive hashing implementations needing to work at a very low level to
hit their performance targets. The abstractions are stripped down to
facilitate this.

# Details of the performance optimization

This function is 2x - 4x faster than LLVM's on small inputs, and up to
2x faster than Abseil. Significant effort has gone into optimizing short
strings in particular compared to Abseil.

Small integer and pointer hashing is also faster than Abseil's by
leveraging a lower quality 64-bit multiply in some cases inspired by
FxHash. One consequence is that this routine is particulary fast for
32-bit integers.

The short string improvements largely come from packing more of the
bytes of string into as few multiplies as possible. While this fails to
mix the bits sufficient to hit SMHasher's strict avalanche criteria and
does leave some collision windows, it provides dramatic latency
improvements. Some of these techniques come from Abseil's own bulk
hashing routine but re-applied here. Others are novel, for example using
small sizes to sample nicely uniform random data to efficiently handle
the very small number of bits of data that need to be hashed.

The other observed improvement is diligent handling of pairs and tuples
and fairly aggressively turning things into integers. Some of the
comparisons with Abseil aren't realistic as the Abseil hash table does
some of these mappings before hashing. I've done this directly in the
hash function as that seems cleaner.

For long strings, the performance is comparable or a bit better than
Abseil, and significantly better than LLVM's hash function.

Overall, for short inputs this is hoped to be the fastest hash function
that still gets "just enough" mixing for modern hash tables to perform
well.

# Details of the quality vs. latency tradeoff

A key insight is that modern hash tables don't need especially high
quality hash functions, but do benefit from something beyond the
identify function. That isn't the target of SMHasher or other quality
assessing tools and has resulted in unnecessarily aggressive hashing for
any functions actually evaluated against it. Many hash functions turn
off the high quality implementations evaluated with SMHasher for integer
or pointer keys to recover latency & performance (AHash for example),
but the same performance-oriented design applies beyond these narrow
types, for example for short strings.

However, a consequence is that there are serious limits to the quality
of the hash function. The avalanche test is failed hilariously, etc.,
but in the exact same ways as Abseil itself fails it for integer keys.
There are also real collisions spaces. For example, for 16-byte strings,
there is one 64-bit value for the first 8 bytes that will have the same
hash regardless of the other 8 bytes of the string. Some minor effort is
taken to make this pattern unlikely to be a practical problem, but it is
a clear theoretical weakness.

It also means that this hash function couldn't be further from providing
any hash-flooding DoS attack protection -- I expect it to be trivially
easy to attack in this way by a motivated adversary. Defending against
these attacks is defined as out-of-scope, in large part because even
attempts that have made a compelling effort to address these issues such
as HighwayHash have found serious limits. Instead, this takes a
principled position that any such defense should be provided entirely at
the data structure level with a strong worst-case bound rather than
through strengthening the hash function.

# Future work

A subsequent PR will introduce a hash table inspired very heavily by the
design of Abseil's "SwissTable" and using this hash function. The goal
is to provide a significant improvement to hot hash tables such as the
identifier table in the lexer of Carbon's toolchain.

# Detailed benchmark data

The benchmarks introduced are heavily inspired by the latency
benchmarking of hash functions in Abseil. I've adapted them to fit
better into Carbon's coding style and to try to have more stable results
with broader coverage of types and string sizes.

Running the benchmarks directly gives horizontal comparisons across
different hash functions. That can be hard to read, so here is *just*
the newly introduced hash function benchmark results on an AMD server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          3.11ns ± 1%
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         3.11ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      4.11ns ± 1%
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         3.12ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    4.13ns ± 1%
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         3.16ns ± 2%
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             3.16ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    4.03ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    4.04ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    4.34ns ± 2%
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        4.04ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        4.34ns ± 2%
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            4.34ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        4.33ns ± 1%
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        4.33ns ± 1%
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        1.95ns ± 4%
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        1.70ns ± 3%
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       3.52ns ± 3%
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       4.46ns ± 2%
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       7.69ns ± 1%
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      14.8ns ± 1%
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      21.5ns ± 1%
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     34.6ns ± 0%
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     63.1ns ± 1%
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      118ns ± 1%
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      225ns ± 1%
```

And on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          5.28ns ± 0%
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         5.29ns ± 0%
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      7.02ns ± 0%
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         5.34ns ± 1%
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    7.07ns ± 4%
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         5.36ns ± 2%
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             5.36ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    7.19ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    7.29ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    7.31ns ± 4%
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        7.29ns ± 2%
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        7.31ns ± 4%
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        8.69ns ± 3%
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        8.69ns ± 3%
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        2.64ns ± 2%
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        2.90ns ± 4%
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       6.14ns ± 1%
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       8.27ns ± 1%
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       13.8ns ± 0%
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      31.2ns ± 0%
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      49.9ns ± 0%
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     86.9ns ± 0%
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      163ns ± 0%
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      312ns ± 0%
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      610ns ± 0%
```

I don't have the same nice statistical multi-run error bars, but one run
from my M1 MacBook:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                             3.89 ns
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                            3.87 ns
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>         4.39 ns
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                            3.93 ns
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>       4.98 ns
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                            3.87 ns
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                                3.87 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>       4.86 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>       4.43 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>       4.41 ns
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>           4.44 ns
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>           4.69 ns
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                         4.33 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>       4.38 ns
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>               4.34 ns
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>           4.35 ns
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>           4.38 ns
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                           1.15 ns
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                          0.973 ns
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                          3.03 ns
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                          3.97 ns
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                          6.64 ns
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                         12.5 ns
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                         17.9 ns
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                        27.9 ns
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                        48.1 ns
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                        87.3 ns
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                         166 ns
```

And here I have internally replaced the Carbon hash function with
Abseil's hash function for "before" and then restored it in the "after"
and computed the delta for each benchmark. This basically shows the
speed-up (lower time -> lower latency -> speed-up -> good) over Abseil
on an AMD server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          4.00ns ± 1%  3.10ns ± 0%  -22.45%  (p=0.000 n=20+15)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         4.01ns ± 1%  3.10ns ± 1%  -22.64%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      6.25ns ± 1%  4.10ns ± 1%  -34.30%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         4.02ns ± 1%  3.12ns ± 1%  -22.50%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    6.25ns ± 1%  4.11ns ± 1%  -34.20%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         4.03ns ± 1%  3.14ns ± 1%  -22.17%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             5.95ns ± 1%  3.14ns ± 1%  -47.24%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    6.04ns ± 1%  4.01ns ± 1%  -33.64%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    5.96ns ± 1%  4.02ns ± 1%  -32.51%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    5.93ns ± 1%  4.30ns ± 1%  -27.56%  (p=0.000 n=20+17)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        7.97ns ± 1%  4.02ns ± 1%  -49.50%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        7.98ns ± 1%  4.32ns ± 1%  -45.88%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      4.40ns ± 2%  4.32ns ± 1%   -1.81%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    5.94ns ± 1%  4.32ns ± 1%  -27.25%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            10.0ns ± 1%   4.3ns ± 1%  -56.56%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        8.04ns ± 1%  4.32ns ± 1%  -46.29%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        7.95ns ± 1%  4.33ns ± 1%  -45.59%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        3.28ns ± 3%  1.93ns ± 4%  -41.19%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        3.05ns ± 3%  1.69ns ± 4%  -44.52%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       5.88ns ± 2%  3.50ns ± 3%  -40.42%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       8.92ns ± 1%  4.44ns ± 2%  -50.22%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       12.0ns ± 1%   7.7ns ± 1%  -36.16%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      18.8ns ± 0%  14.7ns ± 1%  -21.73%  (p=0.000 n=17+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      25.5ns ± 1%  21.4ns ± 1%  -16.18%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     38.7ns ± 2%  34.5ns ± 1%  -10.78%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     69.7ns ± 1%  62.8ns ± 1%   -9.88%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      130ns ± 1%   117ns ± 1%   -9.45%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      244ns ± 0%   225ns ± 1%   -8.11%  (p=0.000 n=17+20)
```

... and on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          6.48ns ± 1%  5.28ns ± 0%  -18.62%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         7.40ns ± 1%  5.29ns ± 1%  -28.45%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      10.4ns ± 0%   7.0ns ± 0%  -32.34%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         6.56ns ± 1%  5.32ns ± 1%  -18.95%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    10.8ns ± 2%   7.0ns ± 1%  -34.89%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         6.71ns ± 3%  5.38ns ± 2%  -19.84%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             10.3ns ± 3%   5.4ns ± 2%  -47.67%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    10.9ns ± 2%   7.2ns ± 4%  -33.67%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    10.7ns ± 4%   7.3ns ± 4%  -31.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    10.5ns ± 3%   7.3ns ± 4%  -30.71%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        14.1ns ± 3%   7.3ns ± 4%  -48.32%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        14.0ns ± 1%   7.3ns ± 4%  -47.95%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      9.41ns ± 4%  8.68ns ± 4%   -7.71%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    12.2ns ± 2%   8.7ns ± 4%  -28.81%  (p=0.000 n=18+19)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            18.9ns ± 2%   8.7ns ± 4%  -54.17%  (p=0.000 n=17+19)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        15.6ns ± 2%   8.7ns ± 4%  -44.37%  (p=0.000 n=17+19)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        15.5ns ± 2%   8.7ns ± 4%  -44.08%  (p=0.000 n=18+19)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        5.89ns ± 2%  2.64ns ± 3%  -55.26%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        5.73ns ± 3%  2.88ns ± 3%  -49.71%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       10.1ns ± 1%   6.1ns ± 2%  -39.00%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       15.7ns ± 0%   8.3ns ± 1%  -47.27%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       21.2ns ± 0%  13.8ns ± 0%  -34.81%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      37.9ns ± 0%  31.2ns ± 0%  -17.77%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      56.8ns ± 0%  49.8ns ± 0%  -12.21%  (p=0.000 n=20+18)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     93.8ns ± 0%  86.9ns ± 0%   -7.38%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      174ns ± 0%   163ns ± 0%   -6.03%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      330ns ± 0%   312ns ± 0%   -5.25%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      641ns ± 0%   610ns ± 0%   -4.79%  (p=0.000 n=19+19)
```

This is the same as the above delta comparison, but with the "before"
being LLVM's hash function:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          6.85ns ± 1%  3.10ns ± 1%  -54.78%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         6.85ns ± 1%  3.10ns ± 1%  -54.78%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      6.25ns ± 1%  4.09ns ± 1%  -34.58%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         6.87ns ± 1%  3.12ns ± 2%  -54.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    7.35ns ± 1%  4.10ns ± 1%  -44.20%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         7.34ns ± 1%  3.13ns ± 1%  -57.34%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             7.33ns ± 1%  3.13ns ± 2%  -57.27%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    7.27ns ± 1%  3.99ns ± 1%  -45.12%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    14.5ns ± 1%   4.0ns ± 1%  -72.23%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    14.6ns ± 1%   4.3ns ± 2%  -70.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        14.5ns ± 1%   4.0ns ± 1%  -72.21%  (p=0.000 n=20+19)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        14.6ns ± 1%   4.3ns ± 1%  -70.46%  (p=0.000 n=20+18)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      7.31ns ± 1%  4.33ns ± 1%  -40.81%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    7.78ns ± 1%  4.32ns ± 1%  -44.45%  (p=0.000 n=18+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            7.78ns ± 2%  4.33ns ± 1%  -44.42%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        7.62ns ± 1%  4.32ns ± 1%  -43.24%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        7.77ns ± 1%  4.33ns ± 1%  -44.34%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        8.15ns ± 3%  1.94ns ± 5%  -76.16%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        7.02ns ± 3%  1.69ns ± 4%  -75.94%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       7.83ns ± 2%  3.50ns ± 3%  -55.34%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       9.17ns ± 1%  4.43ns ± 2%  -51.65%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       11.3ns ± 1%   7.6ns ± 1%  -32.04%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      23.0ns ± 1%  14.7ns ± 1%  -36.14%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      32.9ns ± 0%  21.4ns ± 1%  -34.96%  (p=0.000 n=17+19)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     52.2ns ± 1%  34.4ns ± 1%  -34.01%  (p=0.000 n=19+18)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                     92.1ns ± 1%  62.8ns ± 1%  -31.82%  (p=0.000 n=19+19)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      169ns ± 1%   117ns ± 1%  -30.53%  (p=0.000 n=20+19)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      319ns ± 1%   224ns ± 1%  -29.78%  (p=0.000 n=20+18)
```

... and on an ARM server:

```
BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench>                          8.38ns ± 0%  5.27ns ± 0%  -37.04%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench>                         8.39ns ± 1%  5.28ns ± 0%  -37.01%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench>      8.07ns ± 0%  7.02ns ± 0%  -13.10%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench>                         8.48ns ± 1%  5.32ns ± 1%  -37.25%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench>    9.34ns ± 2%  7.09ns ± 2%  -24.14%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench>                         9.76ns ± 3%  5.37ns ± 2%  -44.98%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<int*>, CarbonHashBench>                             9.76ns ± 3%  5.37ns ± 2%  -44.98%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench>    10.1ns ± 2%   7.2ns ± 3%  -29.36%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench>    11.9ns ± 2%   7.3ns ± 4%  -38.68%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench>    11.3ns ± 2%   7.3ns ± 4%  -35.16%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench>        11.9ns ± 2%   7.3ns ± 4%  -38.68%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench>        11.3ns ± 2%   7.3ns ± 4%  -35.16%  (p=0.000 n=19+19)
BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench>                      10.3ns ± 2%   8.7ns ± 3%  -15.81%  (p=0.000 n=19+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench>    11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench>            11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench>        11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench>        11.6ns ± 3%   8.7ns ± 3%  -25.44%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench>                        9.39ns ± 2%  2.66ns ± 3%  -71.66%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench>                        10.7ns ± 3%   2.9ns ± 3%  -72.97%  (p=0.000 n=19+18)
BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench>                       11.8ns ± 1%   6.1ns ± 2%  -47.75%  (p=0.000 n=19+20)
BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench>                       13.9ns ± 1%   8.3ns ± 1%  -40.71%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench>                       16.8ns ± 1%  13.8ns ± 0%  -17.83%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench>                      31.7ns ± 1%  31.2ns ± 0%   -1.76%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench>                      43.5ns ± 0%  49.8ns ± 0%  +14.56%  (p=0.000 n=18+20)
BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench>                     66.2ns ± 0%  86.9ns ± 0%  +31.39%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench>                      112ns ± 0%   163ns ± 0%  +46.09%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench>                      201ns ± 0%   312ns ± 0%  +55.49%  (p=0.000 n=20+20)
BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench>                      379ns ± 0%   610ns ± 0%  +61.08%  (p=0.000 n=20+20)
```

Note that there is a significant regression on long strings compared to
LLVM's hash function on the ARM server I have access to. This doesn't
show up on the M1 at all, and is likely specific to inadequate
throughput for the 64-bit multiply operations. This seems fine as a) our
priority is for short strings, and b) the M1 and other ARM CPUs are
likely to improve here over time given the prevalent use of this core
technique. For example, Abseil's current hash algorithm has the same
long-string behavior (and performance bottleneck) on this server.

---------

Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Co-authored-by: Geoff Romer <gromer@google.com>
2023-11-20 19:59:06 +00:00