mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-09-24 21:40:12 +01:00
# Overview This is a latency-optimized hashing framework based on Abseil's and others. At it's core it uses both a normal 64-bit multiply as well as a 64-bit multiply capturing both low and high 64-bit components of the result and XOR-ing them together. These are the primitives used in FxHash and Abseil respectively, although they both appear in others. The implementation has been *substantially* optimized for short inputs and latency over quality. As a result, this function does not remotely pass the SMHasher quality tests. However, basic collisions are rare, and I've included a small subset of the SMHasher collision testing directly to make sure the quality doesn't slip too far inadvertently. The customization framework is roughly similar to Abseil's and LLVM's but has been simplified significantly, inspired in some respects by the AHash API design and in others by my experience of all performance sensitive hashing implementations needing to work at a very low level to hit their performance targets. The abstractions are stripped down to facilitate this. # Details of the performance optimization This function is 2x - 4x faster than LLVM's on small inputs, and up to 2x faster than Abseil. Significant effort has gone into optimizing short strings in particular compared to Abseil. Small integer and pointer hashing is also faster than Abseil's by leveraging a lower quality 64-bit multiply in some cases inspired by FxHash. One consequence is that this routine is particulary fast for 32-bit integers. The short string improvements largely come from packing more of the bytes of string into as few multiplies as possible. While this fails to mix the bits sufficient to hit SMHasher's strict avalanche criteria and does leave some collision windows, it provides dramatic latency improvements. Some of these techniques come from Abseil's own bulk hashing routine but re-applied here. Others are novel, for example using small sizes to sample nicely uniform random data to efficiently handle the very small number of bits of data that need to be hashed. The other observed improvement is diligent handling of pairs and tuples and fairly aggressively turning things into integers. Some of the comparisons with Abseil aren't realistic as the Abseil hash table does some of these mappings before hashing. I've done this directly in the hash function as that seems cleaner. For long strings, the performance is comparable or a bit better than Abseil, and significantly better than LLVM's hash function. Overall, for short inputs this is hoped to be the fastest hash function that still gets "just enough" mixing for modern hash tables to perform well. # Details of the quality vs. latency tradeoff A key insight is that modern hash tables don't need especially high quality hash functions, but do benefit from something beyond the identify function. That isn't the target of SMHasher or other quality assessing tools and has resulted in unnecessarily aggressive hashing for any functions actually evaluated against it. Many hash functions turn off the high quality implementations evaluated with SMHasher for integer or pointer keys to recover latency & performance (AHash for example), but the same performance-oriented design applies beyond these narrow types, for example for short strings. However, a consequence is that there are serious limits to the quality of the hash function. The avalanche test is failed hilariously, etc., but in the exact same ways as Abseil itself fails it for integer keys. There are also real collisions spaces. For example, for 16-byte strings, there is one 64-bit value for the first 8 bytes that will have the same hash regardless of the other 8 bytes of the string. Some minor effort is taken to make this pattern unlikely to be a practical problem, but it is a clear theoretical weakness. It also means that this hash function couldn't be further from providing any hash-flooding DoS attack protection -- I expect it to be trivially easy to attack in this way by a motivated adversary. Defending against these attacks is defined as out-of-scope, in large part because even attempts that have made a compelling effort to address these issues such as HighwayHash have found serious limits. Instead, this takes a principled position that any such defense should be provided entirely at the data structure level with a strong worst-case bound rather than through strengthening the hash function. # Future work A subsequent PR will introduce a hash table inspired very heavily by the design of Abseil's "SwissTable" and using this hash function. The goal is to provide a significant improvement to hot hash tables such as the identifier table in the lexer of Carbon's toolchain. # Detailed benchmark data The benchmarks introduced are heavily inspired by the latency benchmarking of hash functions in Abseil. I've adapted them to fit better into Carbon's coding style and to try to have more stable results with broader coverage of types and string sizes. Running the benchmarks directly gives horizontal comparisons across different hash functions. That can be hard to read, so here is *just* the newly introduced hash function benchmark results on an AMD server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 3.11ns ± 1% BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 3.11ns ± 1% BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 4.11ns ± 1% BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 3.12ns ± 1% BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 4.13ns ± 1% BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 3.16ns ± 2% BM_LatencyHash<RandValues<int*>, CarbonHashBench> 3.16ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 4.03ns ± 2% BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 4.04ns ± 1% BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 4.34ns ± 2% BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 4.04ns ± 1% BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 4.34ns ± 2% BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 4.34ns ± 1% BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 1.95ns ± 4% BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 1.70ns ± 3% BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 3.52ns ± 3% BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 4.46ns ± 2% BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 7.69ns ± 1% BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 14.8ns ± 1% BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 21.5ns ± 1% BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 34.6ns ± 0% BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 63.1ns ± 1% BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 118ns ± 1% BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 225ns ± 1% ``` And on an ARM server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 5.28ns ± 0% BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 5.29ns ± 0% BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 7.02ns ± 0% BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 5.34ns ± 1% BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 7.07ns ± 4% BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 5.36ns ± 2% BM_LatencyHash<RandValues<int*>, CarbonHashBench> 5.36ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 7.19ns ± 3% BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 7.29ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 7.31ns ± 4% BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 7.29ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 7.31ns ± 4% BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 2.64ns ± 2% BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 2.90ns ± 4% BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 6.14ns ± 1% BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 8.27ns ± 1% BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 13.8ns ± 0% BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 31.2ns ± 0% BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 49.9ns ± 0% BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 86.9ns ± 0% BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 163ns ± 0% BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 312ns ± 0% BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 610ns ± 0% ``` I don't have the same nice statistical multi-run error bars, but one run from my M1 MacBook: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 3.89 ns BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 3.87 ns BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 4.39 ns BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 3.93 ns BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 4.98 ns BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 3.87 ns BM_LatencyHash<RandValues<int*>, CarbonHashBench> 3.87 ns BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 4.86 ns BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 4.43 ns BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 4.41 ns BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 4.44 ns BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 4.69 ns BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 4.33 ns BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 4.38 ns BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 4.34 ns BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 4.35 ns BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 4.38 ns BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 1.15 ns BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 0.973 ns BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 3.03 ns BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 3.97 ns BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 6.64 ns BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 12.5 ns BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 17.9 ns BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 27.9 ns BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 48.1 ns BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 87.3 ns BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 166 ns ``` And here I have internally replaced the Carbon hash function with Abseil's hash function for "before" and then restored it in the "after" and computed the delta for each benchmark. This basically shows the speed-up (lower time -> lower latency -> speed-up -> good) over Abseil on an AMD server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 4.00ns ± 1% 3.10ns ± 0% -22.45% (p=0.000 n=20+15) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 4.01ns ± 1% 3.10ns ± 1% -22.64% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 6.25ns ± 1% 4.10ns ± 1% -34.30% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 4.02ns ± 1% 3.12ns ± 1% -22.50% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 6.25ns ± 1% 4.11ns ± 1% -34.20% (p=0.000 n=20+19) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 4.03ns ± 1% 3.14ns ± 1% -22.17% (p=0.000 n=19+19) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 5.95ns ± 1% 3.14ns ± 1% -47.24% (p=0.000 n=20+18) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 6.04ns ± 1% 4.01ns ± 1% -33.64% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 5.96ns ± 1% 4.02ns ± 1% -32.51% (p=0.000 n=18+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 5.93ns ± 1% 4.30ns ± 1% -27.56% (p=0.000 n=20+17) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 7.97ns ± 1% 4.02ns ± 1% -49.50% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 7.98ns ± 1% 4.32ns ± 1% -45.88% (p=0.000 n=19+20) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 4.40ns ± 2% 4.32ns ± 1% -1.81% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 5.94ns ± 1% 4.32ns ± 1% -27.25% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 10.0ns ± 1% 4.3ns ± 1% -56.56% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 8.04ns ± 1% 4.32ns ± 1% -46.29% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 7.95ns ± 1% 4.33ns ± 1% -45.59% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 3.28ns ± 3% 1.93ns ± 4% -41.19% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 3.05ns ± 3% 1.69ns ± 4% -44.52% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 5.88ns ± 2% 3.50ns ± 3% -40.42% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 8.92ns ± 1% 4.44ns ± 2% -50.22% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 12.0ns ± 1% 7.7ns ± 1% -36.16% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 18.8ns ± 0% 14.7ns ± 1% -21.73% (p=0.000 n=17+20) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 25.5ns ± 1% 21.4ns ± 1% -16.18% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 38.7ns ± 2% 34.5ns ± 1% -10.78% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 69.7ns ± 1% 62.8ns ± 1% -9.88% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 130ns ± 1% 117ns ± 1% -9.45% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 244ns ± 0% 225ns ± 1% -8.11% (p=0.000 n=17+20) ``` ... and on an ARM server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 6.48ns ± 1% 5.28ns ± 0% -18.62% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 7.40ns ± 1% 5.29ns ± 1% -28.45% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 10.4ns ± 0% 7.0ns ± 0% -32.34% (p=0.000 n=19+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 6.56ns ± 1% 5.32ns ± 1% -18.95% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 10.8ns ± 2% 7.0ns ± 1% -34.89% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 6.71ns ± 3% 5.38ns ± 2% -19.84% (p=0.000 n=20+20) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 10.3ns ± 3% 5.4ns ± 2% -47.67% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 10.9ns ± 2% 7.2ns ± 4% -33.67% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 10.7ns ± 4% 7.3ns ± 4% -31.66% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 10.5ns ± 3% 7.3ns ± 4% -30.71% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 14.1ns ± 3% 7.3ns ± 4% -48.32% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 14.0ns ± 1% 7.3ns ± 4% -47.95% (p=0.000 n=19+19) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 9.41ns ± 4% 8.68ns ± 4% -7.71% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 12.2ns ± 2% 8.7ns ± 4% -28.81% (p=0.000 n=18+19) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 18.9ns ± 2% 8.7ns ± 4% -54.17% (p=0.000 n=17+19) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 15.6ns ± 2% 8.7ns ± 4% -44.37% (p=0.000 n=17+19) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 15.5ns ± 2% 8.7ns ± 4% -44.08% (p=0.000 n=18+19) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 5.89ns ± 2% 2.64ns ± 3% -55.26% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 5.73ns ± 3% 2.88ns ± 3% -49.71% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 10.1ns ± 1% 6.1ns ± 2% -39.00% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 15.7ns ± 0% 8.3ns ± 1% -47.27% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 21.2ns ± 0% 13.8ns ± 0% -34.81% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 37.9ns ± 0% 31.2ns ± 0% -17.77% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 56.8ns ± 0% 49.8ns ± 0% -12.21% (p=0.000 n=20+18) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 93.8ns ± 0% 86.9ns ± 0% -7.38% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 174ns ± 0% 163ns ± 0% -6.03% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 330ns ± 0% 312ns ± 0% -5.25% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 641ns ± 0% 610ns ± 0% -4.79% (p=0.000 n=19+19) ``` This is the same as the above delta comparison, but with the "before" being LLVM's hash function: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 6.85ns ± 1% 3.10ns ± 1% -54.78% (p=0.000 n=20+19) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 6.85ns ± 1% 3.10ns ± 1% -54.78% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 6.25ns ± 1% 4.09ns ± 1% -34.58% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 6.87ns ± 1% 3.12ns ± 2% -54.66% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 7.35ns ± 1% 4.10ns ± 1% -44.20% (p=0.000 n=20+19) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 7.34ns ± 1% 3.13ns ± 1% -57.34% (p=0.000 n=20+18) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 7.33ns ± 1% 3.13ns ± 2% -57.27% (p=0.000 n=20+18) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 7.27ns ± 1% 3.99ns ± 1% -45.12% (p=0.000 n=20+18) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 14.5ns ± 1% 4.0ns ± 1% -72.23% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 14.6ns ± 1% 4.3ns ± 2% -70.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 14.5ns ± 1% 4.0ns ± 1% -72.21% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 14.6ns ± 1% 4.3ns ± 1% -70.46% (p=0.000 n=20+18) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 7.31ns ± 1% 4.33ns ± 1% -40.81% (p=0.000 n=18+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 7.78ns ± 1% 4.32ns ± 1% -44.45% (p=0.000 n=18+20) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 7.78ns ± 2% 4.33ns ± 1% -44.42% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 7.62ns ± 1% 4.32ns ± 1% -43.24% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 7.77ns ± 1% 4.33ns ± 1% -44.34% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 8.15ns ± 3% 1.94ns ± 5% -76.16% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 7.02ns ± 3% 1.69ns ± 4% -75.94% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 7.83ns ± 2% 3.50ns ± 3% -55.34% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 9.17ns ± 1% 4.43ns ± 2% -51.65% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 11.3ns ± 1% 7.6ns ± 1% -32.04% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 23.0ns ± 1% 14.7ns ± 1% -36.14% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 32.9ns ± 0% 21.4ns ± 1% -34.96% (p=0.000 n=17+19) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 52.2ns ± 1% 34.4ns ± 1% -34.01% (p=0.000 n=19+18) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 92.1ns ± 1% 62.8ns ± 1% -31.82% (p=0.000 n=19+19) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 169ns ± 1% 117ns ± 1% -30.53% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 319ns ± 1% 224ns ± 1% -29.78% (p=0.000 n=20+18) ``` ... and on an ARM server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 8.38ns ± 0% 5.27ns ± 0% -37.04% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 8.39ns ± 1% 5.28ns ± 0% -37.01% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 8.07ns ± 0% 7.02ns ± 0% -13.10% (p=0.000 n=19+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 8.48ns ± 1% 5.32ns ± 1% -37.25% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 9.34ns ± 2% 7.09ns ± 2% -24.14% (p=0.000 n=19+20) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 9.76ns ± 3% 5.37ns ± 2% -44.98% (p=0.000 n=20+20) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 9.76ns ± 3% 5.37ns ± 2% -44.98% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 10.1ns ± 2% 7.2ns ± 3% -29.36% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 11.9ns ± 2% 7.3ns ± 4% -38.68% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 11.3ns ± 2% 7.3ns ± 4% -35.16% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 11.9ns ± 2% 7.3ns ± 4% -38.68% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 11.3ns ± 2% 7.3ns ± 4% -35.16% (p=0.000 n=19+19) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 10.3ns ± 2% 8.7ns ± 3% -15.81% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 9.39ns ± 2% 2.66ns ± 3% -71.66% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 10.7ns ± 3% 2.9ns ± 3% -72.97% (p=0.000 n=19+18) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 11.8ns ± 1% 6.1ns ± 2% -47.75% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 13.9ns ± 1% 8.3ns ± 1% -40.71% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 16.8ns ± 1% 13.8ns ± 0% -17.83% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 31.7ns ± 1% 31.2ns ± 0% -1.76% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 43.5ns ± 0% 49.8ns ± 0% +14.56% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 66.2ns ± 0% 86.9ns ± 0% +31.39% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 112ns ± 0% 163ns ± 0% +46.09% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 201ns ± 0% 312ns ± 0% +55.49% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 379ns ± 0% 610ns ± 0% +61.08% (p=0.000 n=20+20) ``` Note that there is a significant regression on long strings compared to LLVM's hash function on the ARM server I have access to. This doesn't show up on the M1 at all, and is likely specific to inadequate throughput for the 64-bit multiply operations. This seems fine as a) our priority is for short strings, and b) the M1 and other ARM CPUs are likely to improve here over time given the prevalent use of this core technique. For example, Abseil's current hash algorithm has the same long-string behavior (and performance bottleneck) on this server. --------- Co-authored-by: josh11b <josh11b@users.noreply.github.com> Co-authored-by: Geoff Romer <gromer@google.com>
720 lines
28 KiB
C++
720 lines
28 KiB
C++
// Part of the Carbon Language project, under the Apache License v2.0 with LLVM
|
|
// Exceptions. See /LICENSE for license information.
|
|
// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
|
|
|
|
#include "common/hashing.h"
|
|
|
|
#include <gmock/gmock.h>
|
|
#include <gtest/gtest.h>
|
|
|
|
#include <type_traits>
|
|
|
|
#include "llvm/ADT/Sequence.h"
|
|
#include "llvm/ADT/StringExtras.h"
|
|
#include "llvm/Support/FormatVariadic.h"
|
|
#include "llvm/Support/TypeName.h"
|
|
|
|
namespace Carbon {
|
|
namespace {
|
|
|
|
using ::testing::Eq;
|
|
using ::testing::Le;
|
|
using ::testing::Ne;
|
|
|
|
TEST(HashingTest, HashCodeAPI) {
|
|
// Manually compute a few hash codes where we can exercise the underlying API.
|
|
HashCode empty = HashValue("");
|
|
HashCode a = HashValue("a");
|
|
HashCode b = HashValue("b");
|
|
ASSERT_THAT(HashValue(""), Eq(empty));
|
|
ASSERT_THAT(HashValue("a"), Eq(a));
|
|
ASSERT_THAT(HashValue("b"), Eq(b));
|
|
ASSERT_THAT(empty, Ne(a));
|
|
ASSERT_THAT(empty, Ne(b));
|
|
ASSERT_THAT(a, Ne(b));
|
|
|
|
// Exercise the methods in basic ways across a few sizes. This doesn't check
|
|
// much beyond stability across re-computed values, crashing, or hitting UB.
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(2), Eq(a.ExtractIndex(2)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(4), Eq(a.ExtractIndex(4)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(8), Eq(a.ExtractIndex(8)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(1 << 10),
|
|
Eq(a.ExtractIndex(1 << 10)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(1 << 20),
|
|
Eq(a.ExtractIndex(1 << 20)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(1 << 30),
|
|
Eq(a.ExtractIndex(1 << 30)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(1LL << 40),
|
|
Eq(a.ExtractIndex(1LL << 40)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndex(1LL << 50),
|
|
Eq(a.ExtractIndex(1LL << 50)));
|
|
|
|
EXPECT_THAT(a.ExtractIndex(8), Ne(b.ExtractIndex(8)));
|
|
EXPECT_THAT(a.ExtractIndex(8), Ne(empty.ExtractIndex(8)));
|
|
|
|
// Note that the index produced with a tag may be different from the index
|
|
// alone!
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<2>(2),
|
|
Eq(a.ExtractIndexAndTag<2>(2)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<16>(4),
|
|
Eq(a.ExtractIndexAndTag<16>(4)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<7>(8),
|
|
Eq(a.ExtractIndexAndTag<7>(8)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<7>(1 << 10),
|
|
Eq(a.ExtractIndexAndTag<7>(1 << 10)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<7>(1 << 20),
|
|
Eq(a.ExtractIndexAndTag<7>(1 << 20)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<7>(1 << 30),
|
|
Eq(a.ExtractIndexAndTag<7>(1 << 30)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<7>(1LL << 40),
|
|
Eq(a.ExtractIndexAndTag<7>(1LL << 40)));
|
|
EXPECT_THAT(HashValue("a").ExtractIndexAndTag<7>(1LL << 50),
|
|
Eq(a.ExtractIndexAndTag<7>(1LL << 50)));
|
|
|
|
const auto [a_index, a_tag] = a.ExtractIndexAndTag<4>(8);
|
|
const auto [b_index, b_tag] = b.ExtractIndexAndTag<4>(8);
|
|
EXPECT_THAT(a_index, Ne(b_index));
|
|
EXPECT_THAT(a_tag, Ne(b_tag));
|
|
}
|
|
|
|
TEST(HashingTest, Integers) {
|
|
for (int64_t i : {0, 1, 2, 3, 42, -1, -2, -3, -13}) {
|
|
SCOPED_TRACE(llvm::formatv("Hashing: {0}", i).str());
|
|
auto test_int_hash = [](auto i) {
|
|
using T = decltype(i);
|
|
SCOPED_TRACE(
|
|
llvm::formatv("Hashing type: {0}", llvm::getTypeName<T>()).str());
|
|
HashCode hash = HashValue(i);
|
|
// Hashes should be stable within the execution.
|
|
EXPECT_THAT(HashValue(i), Eq(hash));
|
|
|
|
// Zero should match, and other integers shouldn't collide trivially.
|
|
HashCode hash_zero = HashValue(static_cast<T>(0));
|
|
if (i == 0) {
|
|
EXPECT_THAT(hash, Eq(hash_zero));
|
|
} else {
|
|
EXPECT_THAT(hash, Ne(hash_zero));
|
|
}
|
|
};
|
|
test_int_hash(i);
|
|
test_int_hash(static_cast<int8_t>(i));
|
|
test_int_hash(static_cast<uint8_t>(i));
|
|
test_int_hash(static_cast<int16_t>(i));
|
|
test_int_hash(static_cast<uint16_t>(i));
|
|
test_int_hash(static_cast<int32_t>(i));
|
|
test_int_hash(static_cast<uint32_t>(i));
|
|
test_int_hash(static_cast<int64_t>(i));
|
|
test_int_hash(static_cast<uint64_t>(i));
|
|
}
|
|
}
|
|
|
|
TEST(HashingTest, Pointers) {
|
|
int object1 = 42;
|
|
std::string object2 =
|
|
"Hello World! This is a long-ish string so it ends up on the heap!";
|
|
|
|
HashCode hash_null = HashValue(nullptr);
|
|
// Hashes should be stable.
|
|
EXPECT_THAT(HashValue(nullptr), Eq(hash_null));
|
|
|
|
// Hash other kinds of pointers without trivial collisions.
|
|
HashCode hash1 = HashValue(&object1);
|
|
HashCode hash2 = HashValue(&object2);
|
|
HashCode hash3 = HashValue(object2.data());
|
|
EXPECT_THAT(hash1, Ne(hash_null));
|
|
EXPECT_THAT(hash2, Ne(hash_null));
|
|
EXPECT_THAT(hash3, Ne(hash_null));
|
|
EXPECT_THAT(hash1, Ne(hash2));
|
|
EXPECT_THAT(hash1, Ne(hash3));
|
|
EXPECT_THAT(hash2, Ne(hash3));
|
|
|
|
// Hash values reflect the address and not the type.
|
|
EXPECT_THAT(HashValue(static_cast<void*>(nullptr)), Eq(hash_null));
|
|
EXPECT_THAT(HashValue(static_cast<int*>(nullptr)), Eq(hash_null));
|
|
EXPECT_THAT(HashValue(static_cast<std::string*>(nullptr)), Eq(hash_null));
|
|
EXPECT_THAT(HashValue(reinterpret_cast<void*>(&object1)), Eq(hash1));
|
|
EXPECT_THAT(HashValue(reinterpret_cast<int*>(&object2)), Eq(hash2));
|
|
EXPECT_THAT(HashValue(reinterpret_cast<std::string*>(object2.data())),
|
|
Eq(hash3));
|
|
}
|
|
|
|
TEST(HashingTest, PairsAndTuples) {
|
|
// Note that we can't compare hash codes across arity, or in general, compare
|
|
// hash codes for different types as the type isn't part of the hash. These
|
|
// hashes are targeted at use in hash tables which pick a single type that's
|
|
// the basis of any comparison.
|
|
HashCode hash_00 = HashValue(std::pair(0, 0));
|
|
HashCode hash_01 = HashValue(std::pair(0, 1));
|
|
HashCode hash_10 = HashValue(std::pair(1, 0));
|
|
HashCode hash_11 = HashValue(std::pair(1, 1));
|
|
EXPECT_THAT(hash_00, Ne(hash_01));
|
|
EXPECT_THAT(hash_00, Ne(hash_10));
|
|
EXPECT_THAT(hash_00, Ne(hash_11));
|
|
EXPECT_THAT(hash_01, Ne(hash_10));
|
|
EXPECT_THAT(hash_01, Ne(hash_11));
|
|
EXPECT_THAT(hash_10, Ne(hash_11));
|
|
|
|
HashCode hash_000 = HashValue(std::tuple(0, 0, 0));
|
|
HashCode hash_001 = HashValue(std::tuple(0, 0, 1));
|
|
HashCode hash_010 = HashValue(std::tuple(0, 1, 0));
|
|
HashCode hash_011 = HashValue(std::tuple(0, 1, 1));
|
|
HashCode hash_100 = HashValue(std::tuple(1, 0, 0));
|
|
HashCode hash_101 = HashValue(std::tuple(1, 0, 1));
|
|
HashCode hash_110 = HashValue(std::tuple(1, 1, 0));
|
|
HashCode hash_111 = HashValue(std::tuple(1, 1, 1));
|
|
EXPECT_THAT(hash_000, Ne(hash_001));
|
|
EXPECT_THAT(hash_000, Ne(hash_010));
|
|
EXPECT_THAT(hash_000, Ne(hash_011));
|
|
EXPECT_THAT(hash_000, Ne(hash_100));
|
|
EXPECT_THAT(hash_000, Ne(hash_101));
|
|
EXPECT_THAT(hash_000, Ne(hash_110));
|
|
EXPECT_THAT(hash_000, Ne(hash_111));
|
|
EXPECT_THAT(hash_001, Ne(hash_010));
|
|
EXPECT_THAT(hash_001, Ne(hash_011));
|
|
EXPECT_THAT(hash_001, Ne(hash_100));
|
|
EXPECT_THAT(hash_001, Ne(hash_101));
|
|
EXPECT_THAT(hash_001, Ne(hash_110));
|
|
EXPECT_THAT(hash_001, Ne(hash_111));
|
|
EXPECT_THAT(hash_010, Ne(hash_011));
|
|
EXPECT_THAT(hash_010, Ne(hash_100));
|
|
EXPECT_THAT(hash_010, Ne(hash_101));
|
|
EXPECT_THAT(hash_010, Ne(hash_110));
|
|
EXPECT_THAT(hash_010, Ne(hash_111));
|
|
EXPECT_THAT(hash_011, Ne(hash_100));
|
|
EXPECT_THAT(hash_011, Ne(hash_101));
|
|
EXPECT_THAT(hash_011, Ne(hash_110));
|
|
EXPECT_THAT(hash_011, Ne(hash_111));
|
|
EXPECT_THAT(hash_100, Ne(hash_101));
|
|
EXPECT_THAT(hash_100, Ne(hash_110));
|
|
EXPECT_THAT(hash_100, Ne(hash_111));
|
|
EXPECT_THAT(hash_101, Ne(hash_110));
|
|
EXPECT_THAT(hash_101, Ne(hash_111));
|
|
EXPECT_THAT(hash_110, Ne(hash_111));
|
|
|
|
// Hashing a 2-tuple and a pair should produce identical results, so pairs
|
|
// are compatible with code using things like variadic tuple construction.
|
|
EXPECT_THAT(HashValue(std::tuple(0, 0)), Eq(hash_00));
|
|
EXPECT_THAT(HashValue(std::tuple(0, 1)), Eq(hash_01));
|
|
EXPECT_THAT(HashValue(std::tuple(1, 0)), Eq(hash_10));
|
|
EXPECT_THAT(HashValue(std::tuple(1, 1)), Eq(hash_11));
|
|
|
|
// Integers in tuples should also work.
|
|
for (int i : {0, 1, 2, 3, 42, -1, -2, -3, -13}) {
|
|
SCOPED_TRACE(llvm::formatv("Hashing: ({0}, {0}, {0})", i).str());
|
|
auto test_int_tuple_hash = [](auto i) {
|
|
using T = decltype(i);
|
|
SCOPED_TRACE(
|
|
llvm::formatv("Hashing integer type: {0}", llvm::getTypeName<T>())
|
|
.str());
|
|
std::tuple v = {i, i, i};
|
|
HashCode hash = HashValue(v);
|
|
|
|
// Hashes should be stable within the execution.
|
|
EXPECT_THAT(HashValue(v), Eq(hash));
|
|
|
|
// Zero should match, and other integers shouldn't collide trivially.
|
|
T zero = 0;
|
|
std::tuple zero_tuple = {zero, zero, zero};
|
|
HashCode hash_zero = HashValue(zero_tuple);
|
|
if (i == 0) {
|
|
EXPECT_THAT(hash, Eq(hash_zero));
|
|
} else {
|
|
EXPECT_THAT(hash, Ne(hash_zero));
|
|
}
|
|
};
|
|
test_int_tuple_hash(i);
|
|
test_int_tuple_hash(static_cast<int8_t>(i));
|
|
test_int_tuple_hash(static_cast<uint8_t>(i));
|
|
test_int_tuple_hash(static_cast<int16_t>(i));
|
|
test_int_tuple_hash(static_cast<uint16_t>(i));
|
|
test_int_tuple_hash(static_cast<int32_t>(i));
|
|
test_int_tuple_hash(static_cast<uint32_t>(i));
|
|
test_int_tuple_hash(static_cast<int64_t>(i));
|
|
test_int_tuple_hash(static_cast<uint64_t>(i));
|
|
|
|
// Heterogeneous integer types should also work, but we only support
|
|
// comparing against hashes of tuples with the exact same type.
|
|
using T1 = std::tuple<int8_t, uint32_t, int16_t>;
|
|
using T2 = std::tuple<uint32_t, int16_t, uint64_t>;
|
|
if (i == 0) {
|
|
EXPECT_THAT(HashValue(T1{i, i, i}), Eq(HashValue(T1{0, 0, 0})));
|
|
EXPECT_THAT(HashValue(T2{i, i, i}), Eq(HashValue(T2{0, 0, 0})));
|
|
} else {
|
|
EXPECT_THAT(HashValue(T1{i, i, i}), Ne(HashValue(T1{0, 0, 0})));
|
|
EXPECT_THAT(HashValue(T2{i, i, i}), Ne(HashValue(T2{0, 0, 0})));
|
|
}
|
|
}
|
|
|
|
// Hash values of pointers in pairs and tuples reflect the address and not the
|
|
// type. Pairs and 2-tuples give the same hash values.
|
|
HashCode hash_2null = HashValue(std::pair(nullptr, nullptr));
|
|
EXPECT_THAT(HashValue(std::tuple(static_cast<int*>(nullptr),
|
|
static_cast<double*>(nullptr))),
|
|
Eq(hash_2null));
|
|
|
|
// Hash other kinds of pointers without trivial collisions.
|
|
int object1 = 42;
|
|
std::string object2 = "Hello world!";
|
|
HashCode hash_3ptr =
|
|
HashValue(std::tuple(&object1, &object2, object2.data()));
|
|
EXPECT_THAT(hash_3ptr, Ne(HashValue(std::tuple(nullptr, nullptr, nullptr))));
|
|
|
|
// Hash values reflect the address and not the type.
|
|
EXPECT_THAT(
|
|
HashValue(std::tuple(reinterpret_cast<void*>(&object1),
|
|
reinterpret_cast<int*>(&object2),
|
|
reinterpret_cast<std::string*>(object2.data()))),
|
|
Eq(hash_3ptr));
|
|
}
|
|
|
|
TEST(HashingTest, BasicStrings) {
|
|
llvm::SmallVector<std::pair<std::string, HashCode>> hashes;
|
|
for (int size : {0, 1, 2, 4, 16, 64, 256, 1024}) {
|
|
std::string s(size, 'a');
|
|
hashes.push_back({s, HashValue(s)});
|
|
}
|
|
for (const auto& [s1, hash1] : hashes) {
|
|
EXPECT_THAT(HashValue(s1), Eq(hash1));
|
|
// Also check that we get the same hashes even when using string-wrapping
|
|
// types.
|
|
EXPECT_THAT(HashValue(std::string_view(s1)), Eq(hash1));
|
|
EXPECT_THAT(HashValue(llvm::StringRef(s1)), Eq(hash1));
|
|
|
|
// And some basic tests that simple things don't collide.
|
|
for (const auto& [s2, hash2] : hashes) {
|
|
if (s1 != s2) {
|
|
EXPECT_THAT(hash1, Ne(hash2))
|
|
<< "Matching hashes for '" << s1 << "' and '" << s2 << "'";
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
struct HashableType {
|
|
int x;
|
|
int y;
|
|
|
|
int ignored = 0;
|
|
|
|
friend auto CarbonHashValue(const HashableType& value, uint64_t seed)
|
|
-> HashCode {
|
|
Hasher hasher(seed);
|
|
hasher.Hash(value.x, value.y);
|
|
return static_cast<HashCode>(hasher);
|
|
}
|
|
};
|
|
|
|
TEST(HashingTest, CustomType) {
|
|
HashableType a = {.x = 1, .y = 2};
|
|
HashableType b = {.x = 3, .y = 4};
|
|
|
|
EXPECT_THAT(HashValue(a), Eq(HashValue(a)));
|
|
EXPECT_THAT(HashValue(a), Ne(HashValue(b)));
|
|
|
|
// Differences in an ignored field have no impact.
|
|
HashableType c = {.x = 3, .y = 4, .ignored = 42};
|
|
EXPECT_THAT(HashValue(c), Eq(HashValue(b)));
|
|
}
|
|
|
|
// The only significantly bad seed is zero, so pick a non-zero seed with a tiny
|
|
// amount of entropy to make sure that none of the testing relies on the entropy
|
|
// from this.
|
|
constexpr uint64_t TestSeed = 42ULL * 1024;
|
|
|
|
auto ToHexBytes(llvm::StringRef s) -> std::string {
|
|
std::string rendered;
|
|
llvm::raw_string_ostream os(rendered);
|
|
os << "{";
|
|
llvm::ListSeparator sep(", ");
|
|
for (const char c : s) {
|
|
os << sep << llvm::formatv("{0:x2}", static_cast<uint8_t>(c));
|
|
}
|
|
os << "}";
|
|
return rendered;
|
|
}
|
|
|
|
template <typename T>
|
|
struct HashedValue {
|
|
HashCode hash;
|
|
T v;
|
|
};
|
|
|
|
using HashedString = HashedValue<std::string>;
|
|
|
|
template <typename T>
|
|
auto PrintFullWidthHex(llvm::raw_ostream& os, T value) {
|
|
static_assert(sizeof(T) == 1 || sizeof(T) == 2 || sizeof(T) == 4 ||
|
|
sizeof(T) == 8);
|
|
os << llvm::formatv(sizeof(T) == 1 ? "{0:x2}"
|
|
: sizeof(T) == 2 ? "{0:x4}"
|
|
: sizeof(T) == 4 ? "{0:x8}"
|
|
: "{0:x16}",
|
|
static_cast<uint64_t>(value));
|
|
}
|
|
|
|
template <typename T, typename = std::enable_if_t<std::is_integral_v<T>>>
|
|
auto operator<<(llvm::raw_ostream& os, HashedValue<T> hv)
|
|
-> llvm::raw_ostream& {
|
|
os << "hash " << hv.hash << " for value ";
|
|
PrintFullWidthHex(os, hv.v);
|
|
return os;
|
|
}
|
|
|
|
template <typename T, typename U,
|
|
typename = std::enable_if_t<std::is_integral_v<T>>,
|
|
typename = std::enable_if_t<std::is_integral_v<U>>>
|
|
auto operator<<(llvm::raw_ostream& os, HashedValue<std::pair<T, U>> hv)
|
|
-> llvm::raw_ostream& {
|
|
os << "hash " << hv.hash << " for pair of ";
|
|
PrintFullWidthHex(os, hv.v.first);
|
|
os << " and ";
|
|
PrintFullWidthHex(os, hv.v.second);
|
|
return os;
|
|
}
|
|
|
|
struct Collisions {
|
|
int total;
|
|
int median;
|
|
int max;
|
|
};
|
|
|
|
// Analyzes a list of hashed values to find all of the hash codes which collide
|
|
// within a specific bit-range.
|
|
//
|
|
// With `BitBegin=0` and `BitEnd=64`, this is equivalent to finding full
|
|
// collisions. But when the begin and end of the bit range are narrower than the
|
|
// 64-bits of the hash code, it allows this function to analyze a specific
|
|
// window of bits within the 64-bit hash code to understand how many collisions
|
|
// emerge purely within that bit range.
|
|
//
|
|
// With narrow ranges (we often look at the first N and last N bits for small
|
|
// N), collisions are common and so this function summarizes this with the total
|
|
// number of collisions and the median number of collisions for an input value.
|
|
template <int BitBegin, int BitEnd, typename T>
|
|
auto FindBitRangeCollisions(llvm::ArrayRef<HashedValue<T>> hashes)
|
|
-> Collisions {
|
|
static_assert(BitBegin < BitEnd);
|
|
constexpr int BitCount = BitEnd - BitBegin;
|
|
static_assert(BitCount <= 32);
|
|
constexpr int BitShift = BitBegin;
|
|
constexpr uint64_t BitMask = ((1ULL << BitCount) - 1) << BitShift;
|
|
|
|
// We collect counts of collisions in a vector. Initially, we just have a zero
|
|
// and all inputs map to that collision count. As we discover collisions,
|
|
// we'll create a dedicated counter for it and count how many inputs collide.
|
|
llvm::SmallVector<int> collision_counts;
|
|
collision_counts.push_back(0);
|
|
// The "map" for collision counts. Each input hashed value has a corresponding
|
|
// index stored here. That index is the index of the collision count in the
|
|
// container above. We resize this to fill it with zeros to start as the zero
|
|
// index above has a collision count of zero.
|
|
//
|
|
// The result of this is that the number of collisions for `hashes[i]` is
|
|
// `collision_counts[collision_map[i]]`.
|
|
llvm::SmallVector<int> collision_map;
|
|
collision_map.resize(hashes.size());
|
|
|
|
// First, we extract the bit subsequence we want to examine from each hash and
|
|
// store it with an index back into the hashed values (or the collision map).
|
|
//
|
|
// The result is that, `bits_and_indices[i].bits` has the hash bits of
|
|
// interest from `hashes[bits_and_indices[i].index]`.
|
|
//
|
|
// And because `collision_map` above uses the same indices as `hashes`,
|
|
// `collision_counts[collision_map[bits_and_indices[i].index]]` is the number
|
|
// of collisions for `bits_and_indices[i].bits`.
|
|
struct BitSequenceAndHashIndex {
|
|
// The bit subsequence of a hash input, adjusted into the low bits.
|
|
uint32_t bits;
|
|
// The index of the hash input corresponding to this bit sequence.
|
|
int index;
|
|
};
|
|
llvm::SmallVector<BitSequenceAndHashIndex> bits_and_indices;
|
|
bits_and_indices.reserve(hashes.size());
|
|
for (const auto& [hash, v] : hashes) {
|
|
CARBON_DCHECK(v == hashes[bits_and_indices.size()].v);
|
|
auto hash_bits = (static_cast<uint64_t>(hash) & BitMask) >> BitShift;
|
|
bits_and_indices.push_back(
|
|
{.bits = static_cast<uint32_t>(hash_bits),
|
|
.index = static_cast<int>(bits_and_indices.size())});
|
|
}
|
|
|
|
// Now we sort by the extracted bit sequence so we can efficiently scan for
|
|
// colliding bit patterns.
|
|
std::sort(
|
|
bits_and_indices.begin(), bits_and_indices.end(),
|
|
[](const auto& lhs, const auto& rhs) { return lhs.bits < rhs.bits; });
|
|
|
|
// Scan the sorted bit sequences we've extracted looking for collisions. We
|
|
// count the total collisions, but we also track the number of individual
|
|
// inputs that collide with each specific bit pattern.
|
|
uint32_t prev_hash_bits = bits_and_indices[0].bits;
|
|
int prev_index = bits_and_indices[0].index;
|
|
bool in_collision = false;
|
|
int total = 0;
|
|
for (const auto& [hash_bits, hash_index] :
|
|
llvm::ArrayRef(bits_and_indices).slice(1)) {
|
|
// Check if we've found a new hash (and thus a new value), reset everything.
|
|
CARBON_CHECK(hashes[prev_index].v != hashes[hash_index].v);
|
|
if (hash_bits != prev_hash_bits) {
|
|
CARBON_CHECK(hashes[prev_index].hash != hashes[hash_index].hash);
|
|
prev_hash_bits = hash_bits;
|
|
prev_index = hash_index;
|
|
in_collision = false;
|
|
continue;
|
|
}
|
|
|
|
// Otherwise, we have a colliding bit sequence.
|
|
++total;
|
|
|
|
// If we've already created a collision count to track this, just increment
|
|
// it and map this hash to it.
|
|
if (in_collision) {
|
|
++collision_counts.back();
|
|
collision_map[hash_index] = collision_counts.size() - 1;
|
|
continue;
|
|
}
|
|
|
|
// If this is a new collision, create a dedicated count to track it and
|
|
// begin counting.
|
|
in_collision = true;
|
|
collision_map[prev_index] = collision_counts.size();
|
|
collision_map[hash_index] = collision_counts.size();
|
|
collision_counts.push_back(1);
|
|
}
|
|
|
|
// Sort by collision count for each hash.
|
|
std::sort(bits_and_indices.begin(), bits_and_indices.end(),
|
|
[&](const auto& lhs, const auto& rhs) {
|
|
return collision_counts[collision_map[lhs.index]] <
|
|
collision_counts[collision_map[rhs.index]];
|
|
});
|
|
|
|
// And compute the median and max.
|
|
int median = collision_counts
|
|
[collision_map[bits_and_indices[bits_and_indices.size() / 2].index]];
|
|
int max = *std::max_element(collision_counts.begin(), collision_counts.end());
|
|
CARBON_CHECK(max ==
|
|
collision_counts[collision_map[bits_and_indices.back().index]]);
|
|
return {.total = total, .median = median, .max = max};
|
|
}
|
|
|
|
auto CheckNoDuplicateValues(llvm::ArrayRef<HashedString> hashes) -> void {
|
|
for (int i = 0, size = hashes.size(); i < size - 1; ++i) {
|
|
const auto& [_, value] = hashes[i];
|
|
CARBON_CHECK(value != hashes[i + 1].v) << "Duplicate value: " << value;
|
|
}
|
|
}
|
|
|
|
template <int N>
|
|
auto AllByteStringsHashedAndSorted() {
|
|
static_assert(N < 5, "Can only generate all 4-byte strings or shorter.");
|
|
|
|
llvm::SmallVector<HashedString> hashes;
|
|
int64_t count = 1LL << (N * 8);
|
|
for (int64_t i : llvm::seq(count)) {
|
|
uint8_t bytes[N];
|
|
for (int j : llvm::seq(N)) {
|
|
bytes[j] = (static_cast<uint64_t>(i) >> (8 * j)) & 0xff;
|
|
}
|
|
std::string s(std::begin(bytes), std::end(bytes));
|
|
hashes.push_back({HashValue(s, TestSeed), s});
|
|
}
|
|
|
|
std::sort(hashes.begin(), hashes.end(),
|
|
[](const HashedString& lhs, const HashedString& rhs) {
|
|
return static_cast<uint64_t>(lhs.hash) <
|
|
static_cast<uint64_t>(rhs.hash);
|
|
});
|
|
CheckNoDuplicateValues(hashes);
|
|
|
|
return hashes;
|
|
}
|
|
|
|
auto ExpectNoHashCollisions(llvm::ArrayRef<HashedString> hashes) -> void {
|
|
HashCode prev_hash = hashes[0].hash;
|
|
llvm::StringRef prev_s = hashes[0].v;
|
|
for (const auto& [hash, s] : hashes.slice(1)) {
|
|
if (hash != prev_hash) {
|
|
prev_hash = hash;
|
|
prev_s = s;
|
|
continue;
|
|
}
|
|
|
|
FAIL() << "Colliding hash '" << hash << "' of strings "
|
|
<< ToHexBytes(prev_s) << " and " << ToHexBytes(s);
|
|
}
|
|
}
|
|
|
|
TEST(HashingTest, Collisions1ByteSized) {
|
|
auto hashes_storage = AllByteStringsHashedAndSorted<1>();
|
|
auto hashes = llvm::ArrayRef(hashes_storage);
|
|
ExpectNoHashCollisions(hashes);
|
|
|
|
auto low_32bit_collisions = FindBitRangeCollisions<0, 32>(hashes);
|
|
EXPECT_THAT(low_32bit_collisions.total, Eq(0));
|
|
auto high_32bit_collisions = FindBitRangeCollisions<32, 64>(hashes);
|
|
EXPECT_THAT(high_32bit_collisions.total, Eq(0));
|
|
|
|
// We expect collisions when only looking at 7-bits of the hash. However,
|
|
// modern hash table designs need to use either the low or high 7 bits as tags
|
|
// for faster searching. So we add some direct testing that the median and max
|
|
// collisions for any given key stay within bounds. We express the bounds in
|
|
// terms of the minimum expected "perfect" rate of collisions if uniformly
|
|
// distributed.
|
|
int min_7bit_collisions = llvm::NextPowerOf2(hashes.size() - 1) / (1 << 7);
|
|
auto low_7bit_collisions = FindBitRangeCollisions<0, 7>(hashes);
|
|
EXPECT_THAT(low_7bit_collisions.median, Le(2 * min_7bit_collisions));
|
|
EXPECT_THAT(low_7bit_collisions.max, Le(4 * min_7bit_collisions));
|
|
auto high_7bit_collisions = FindBitRangeCollisions<64 - 7, 64>(hashes);
|
|
EXPECT_THAT(high_7bit_collisions.median, Le(2 * min_7bit_collisions));
|
|
EXPECT_THAT(high_7bit_collisions.max, Le(4 * min_7bit_collisions));
|
|
}
|
|
|
|
TEST(HashingTest, Collisions2ByteSized) {
|
|
auto hashes_storage = AllByteStringsHashedAndSorted<2>();
|
|
auto hashes = llvm::ArrayRef(hashes_storage);
|
|
ExpectNoHashCollisions(hashes);
|
|
|
|
auto low_32bit_collisions = FindBitRangeCollisions<0, 32>(hashes);
|
|
EXPECT_THAT(low_32bit_collisions.total, Eq(0));
|
|
auto high_32bit_collisions = FindBitRangeCollisions<32, 64>(hashes);
|
|
EXPECT_THAT(high_32bit_collisions.total, Eq(0));
|
|
|
|
// Similar to 1-byte keys, we do expect a certain rate of collisions here but
|
|
// bound the median and max.
|
|
int min_7bit_collisions = llvm::NextPowerOf2(hashes.size() - 1) / (1 << 7);
|
|
auto low_7bit_collisions = FindBitRangeCollisions<0, 7>(hashes);
|
|
EXPECT_THAT(low_7bit_collisions.median, Le(2 * min_7bit_collisions));
|
|
EXPECT_THAT(low_7bit_collisions.max, Le(2 * min_7bit_collisions));
|
|
auto high_7bit_collisions = FindBitRangeCollisions<64 - 7, 64>(hashes);
|
|
EXPECT_THAT(high_7bit_collisions.median, Le(2 * min_7bit_collisions));
|
|
EXPECT_THAT(high_7bit_collisions.max, Le(2 * min_7bit_collisions));
|
|
}
|
|
|
|
// Generate and hash all strings of of [BeginByteCount, EndByteCount) bytes,
|
|
// with [BeginSetBitCount, EndSetBitCount) contiguous bits at each possible bit
|
|
// offset set to one and all other bits set to zero.
|
|
template <int BeginByteCount, int EndByteCount, int BeginSetBitCount,
|
|
int EndSetBitCount>
|
|
struct SparseHashTestParamRanges {
|
|
static_assert(BeginByteCount >= 0);
|
|
static_assert(BeginByteCount < EndByteCount);
|
|
static_assert(BeginSetBitCount >= 0);
|
|
static_assert(BeginSetBitCount < EndSetBitCount);
|
|
// Note that we intentionally allow the end-set-bit-count to result in more
|
|
// set bits than are available -- we truncate the number of set bits to fit
|
|
// within the byte string.
|
|
static_assert(BeginSetBitCount <= BeginByteCount * 8);
|
|
|
|
struct ByteCount {
|
|
static constexpr int Begin = BeginByteCount;
|
|
static constexpr int End = EndByteCount;
|
|
};
|
|
struct SetBitCount {
|
|
static constexpr int Begin = BeginSetBitCount;
|
|
static constexpr int End = EndSetBitCount;
|
|
};
|
|
};
|
|
|
|
template <typename ParamRanges>
|
|
struct SparseHashTest : ::testing::Test {
|
|
using ByteCount = typename ParamRanges::ByteCount;
|
|
using SetBitCount = typename ParamRanges::SetBitCount;
|
|
|
|
static auto GetHashedByteStrings() {
|
|
llvm::SmallVector<HashedString> hashes;
|
|
for (int byte_count :
|
|
llvm::seq_inclusive(ByteCount::Begin, ByteCount::End)) {
|
|
int bits = byte_count * 8;
|
|
for (int set_bit_count : llvm::seq_inclusive(
|
|
SetBitCount::Begin, std::min(bits, SetBitCount::End))) {
|
|
if (set_bit_count == 0) {
|
|
std::string s(byte_count, '\0');
|
|
hashes.push_back({HashValue(s, TestSeed), std::move(s)});
|
|
continue;
|
|
}
|
|
for (int begin_set_bit : llvm::seq_inclusive(0, bits - set_bit_count)) {
|
|
std::string s(byte_count, '\0');
|
|
|
|
int begin_set_bit_byte_index = begin_set_bit / 8;
|
|
int begin_set_bit_bit_index = begin_set_bit % 8;
|
|
int end_set_bit_byte_index = (begin_set_bit + set_bit_count) / 8;
|
|
int end_set_bit_bit_index = (begin_set_bit + set_bit_count) % 8;
|
|
|
|
// We build a begin byte and end byte. We set the begin byte, set
|
|
// subsequent bytes up to *and including* the end byte to all ones,
|
|
// and then mask the end byte. For multi-byte runs, the mask just sets
|
|
// the end byte and for single-byte runs the mask computes the
|
|
// intersecting bits.
|
|
//
|
|
// Consider a 4-set-bit count, starting at bit 2. The begin bit index
|
|
// is 2, and the end bit index is 6.
|
|
//
|
|
// Begin byte: 0b11111111 -(shl 2)-----> 0b11111100
|
|
// End byte: 0b11111111 -(shr (8-6))-> 0b00111111
|
|
// Masked byte: 0b00111100
|
|
//
|
|
// Or a 10-set-bit-count starting at bit 2. The begin bit index is 2,
|
|
// the end byte index is (12 / 8) or 1, and the end bit index is (12 %
|
|
// 8) or 4.
|
|
//
|
|
// Begin byte: 0b11111111 -(shl 2)-----> 0b11111100 -> 6 bits
|
|
// End byte: 0b11111111 -(shr (8-4))-> 0b00001111 -> 4 bits
|
|
// 10 total bits
|
|
//
|
|
uint8_t begin_set_bit_byte = 0xFFU << begin_set_bit_bit_index;
|
|
uint8_t end_set_bit_byte = 0xFFU >> (8 - end_set_bit_bit_index);
|
|
bool has_end_byte_bits = end_set_bit_byte != 0;
|
|
s[begin_set_bit_byte_index] = begin_set_bit_byte;
|
|
for (int i : llvm::seq(begin_set_bit_byte_index + 1,
|
|
end_set_bit_byte_index + has_end_byte_bits)) {
|
|
s[i] = '\xFF';
|
|
}
|
|
// If there are no bits set in the end byte, it may be past-the-end
|
|
// and we can't even mask a zero byte safely.
|
|
if (has_end_byte_bits) {
|
|
s[end_set_bit_byte_index] &= end_set_bit_byte;
|
|
}
|
|
hashes.push_back({HashValue(s, TestSeed), std::move(s)});
|
|
}
|
|
}
|
|
}
|
|
|
|
std::sort(hashes.begin(), hashes.end(),
|
|
[](const HashedString& lhs, const HashedString& rhs) {
|
|
return static_cast<uint64_t>(lhs.hash) <
|
|
static_cast<uint64_t>(rhs.hash);
|
|
});
|
|
CheckNoDuplicateValues(hashes);
|
|
|
|
return hashes;
|
|
}
|
|
};
|
|
|
|
using SparseHashTestParams = ::testing::Types<
|
|
SparseHashTestParamRanges</*BeginByteCount=*/0, /*EndByteCount=*/256,
|
|
/*BeginSetBitCount=*/0, /*EndSetBitCount=*/1>,
|
|
SparseHashTestParamRanges</*BeginByteCount=*/1, /*EndByteCount=*/128,
|
|
/*BeginSetBitCount=*/2, /*EndSetBitCount=*/4>,
|
|
SparseHashTestParamRanges</*BeginByteCount=*/1, /*EndByteCount=*/64,
|
|
/*BeginSetBitCount=*/4, /*EndSetBitCount=*/16>>;
|
|
TYPED_TEST_SUITE(SparseHashTest, SparseHashTestParams);
|
|
|
|
TYPED_TEST(SparseHashTest, Collisions) {
|
|
auto hashes_storage = this->GetHashedByteStrings();
|
|
auto hashes = llvm::ArrayRef(hashes_storage);
|
|
ExpectNoHashCollisions(hashes);
|
|
|
|
int min_7bit_collisions = llvm::NextPowerOf2(hashes.size() - 1) / (1 << 7);
|
|
auto low_7bit_collisions = FindBitRangeCollisions<0, 7>(hashes);
|
|
EXPECT_THAT(low_7bit_collisions.median, Le(2 * min_7bit_collisions));
|
|
EXPECT_THAT(low_7bit_collisions.max, Le(2 * min_7bit_collisions));
|
|
auto high_7bit_collisions = FindBitRangeCollisions<64 - 7, 64>(hashes);
|
|
EXPECT_THAT(high_7bit_collisions.median, Le(2 * min_7bit_collisions));
|
|
EXPECT_THAT(high_7bit_collisions.max, Le(2 * min_7bit_collisions));
|
|
}
|
|
|
|
} // namespace
|
|
} // namespace Carbon
|