mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-10-03 09:15:49 +01:00
bbca8668aef03ca995bc2dbddff4e36be3ff082d
16
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ee97511496 |
Fix IsCarbonMap invocations to avoid build failures for non-Carbon map types (not sure when this broke). (#6662)
Also, update the multiplication constant for carbon hashing for improved probing. Co-authored-by: Evan Brown <ezb@google.com> |
||
|
|
384e1cbb92 |
Update Read*To* to improve operation dependency graph. (#4746)
Benchmarks for StringRef key seem slightly positive. ``` name old CYCLES/op new CYCLES/op delta BM_MapContainsHit<Map<llvm::StringRef, int>>/1/256 24.2 ± 0% 23.9 ± 0% -1.14% (p=0.000 n=54+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/2/256 24.2 ± 0% 23.9 ± 0% -1.15% (p=0.000 n=53+54) BM_MapContainsHit<Map<llvm::StringRef, int>>/3/256 24.2 ± 0% 23.9 ± 0% -1.15% (p=0.000 n=53+54) BM_MapContainsHit<Map<llvm::StringRef, int>>/4/256 24.2 ± 0% 23.9 ± 0% -1.14% (p=0.000 n=56+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/8/256 25.4 ± 3% 26.3 ± 4% +3.61% (p=0.000 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/16/256 29.1 ± 2% 29.0 ± 2% -0.28% (p=0.030 n=56+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/32/256 29.2 ± 2% 29.0 ± 1% -0.59% (p=0.000 n=57+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/64/256 30.1 ± 2% 30.0 ± 2% -0.43% (p=0.045 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/256/256 30.5 ± 1% 30.3 ± 1% -0.56% (p=0.000 n=56+56) BM_MapContainsHit<Map<llvm::StringRef, int>>/256/64 29.2 ± 1% 29.2 ± 2% ~ (p=0.513 n=55+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/256/128 29.6 ± 1% 29.5 ± 1% -0.34% (p=0.002 n=55+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/4096/256 32.0 ± 2% 31.9 ± 2% ~ (p=0.082 n=55+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/4096/1024 37.8 ± 2% 37.8 ± 1% ~ (p=0.751 n=57+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/4096/2048 45.3 ± 2% 45.5 ± 2% +0.46% (p=0.001 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/65536/256 34.3 ± 2% 34.2 ± 2% -0.46% (p=0.000 n=57+56) BM_MapContainsHit<Map<llvm::StringRef, int>>/65536/16384 72.4 ± 3% 72.3 ± 2% ~ (p=0.458 n=54+50) BM_MapContainsHit<Map<llvm::StringRef, int>>/65536/32768 77.7 ± 3% 77.6 ± 3% ~ (p=0.774 n=56+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/1048576/256 34.9 ± 1% 34.8 ± 2% ~ (p=0.051 n=56+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/1048576/262144 115 ± 5% 115 ± 5% ~ (p=0.660 n=57+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/1048576/524288 145 ± 4% 145 ± 5% ~ (p=0.917 n=57+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/16777216/256 36.5 ± 2% 36.5 ± 2% ~ (p=0.250 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/16777216/4194304 288 ± 3% 287 ± 4% ~ (p=0.058 n=56+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/16777216/8388608 303 ± 2% 302 ± 3% -0.47% (p=0.044 n=53+54) BM_MapContainsHit<Map<llvm::StringRef, int>>/56/256 29.1 ± 3% 29.0 ± 3% ~ (p=0.147 n=56+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/224/256 30.7 ± 2% 30.6 ± 2% ~ (p=0.140 n=56+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/3584/256 31.4 ± 1% 31.3 ± 1% -0.42% (p=0.003 n=53+54) BM_MapContainsHit<Map<llvm::StringRef, int>>/3584/896 35.8 ± 2% 36.0 ± 2% +0.58% (p=0.000 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/3584/1792 43.5 ± 1% 43.6 ± 2% +0.21% (p=0.032 n=51+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/57344/256 34.3 ± 2% 34.1 ± 1% -0.43% (p=0.003 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/57344/14336 67.1 ± 2% 66.8 ± 2% ~ (p=0.057 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/57344/28672 72.8 ± 3% 72.5 ± 3% -0.45% (p=0.032 n=57+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/917504/256 34.7 ± 2% 34.6 ± 2% ~ (p=0.065 n=56+57) BM_MapContainsHit<Map<llvm::StringRef, int>>/917504/229376 104 ± 4% 104 ± 5% ~ (p=0.853 n=55+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/917504/458752 114 ± 6% 114 ± 5% ~ (p=0.643 n=56+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/14680064/256 36.4 ± 2% 36.2 ± 2% -0.58% (p=0.001 n=56+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/14680064/3670016 271 ± 2% 271 ± 4% ~ (p=0.632 n=55+55) BM_MapContainsHit<Map<llvm::StringRef, int>>/14680064/7340032 285 ± 3% 285 ± 3% ~ (p=0.658 n=57+55) BM_MapContainsMiss<Map<llvm::StringRef, int>>/1 19.3 ± 1% 19.3 ± 2% ~ (p=0.201 n=55+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/2 19.4 ± 1% 19.3 ± 1% ~ (p=0.191 n=56+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/3 19.4 ± 1% 19.4 ± 2% ~ (p=0.422 n=55+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/4 19.4 ± 1% 19.4 ± 1% ~ (p=0.179 n=56+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/8 19.5 ± 2% 19.5 ± 1% ~ (p=0.148 n=54+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/16 19.7 ± 2% 19.6 ± 2% ~ (p=0.204 n=54+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/32 20.0 ± 3% 20.0 ± 3% ~ (p=0.917 n=56+54) BM_MapContainsMiss<Map<llvm::StringRef, int>>/64 19.8 ± 3% 19.8 ± 3% ~ (p=0.245 n=57+54) BM_MapContainsMiss<Map<llvm::StringRef, int>>/256 20.1 ± 3% 20.1 ± 3% ~ (p=0.307 n=57+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/4096 20.1 ± 3% 20.2 ± 2% ~ (p=0.070 n=57+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/65536 20.5 ± 3% 20.5 ± 3% ~ (p=0.174 n=56+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/1048576 20.9 ± 2% 20.8 ± 3% ~ (p=0.476 n=53+55) BM_MapContainsMiss<Map<llvm::StringRef, int>>/16777216 22.2 ± 4% 22.2 ± 3% ~ (p=0.807 n=57+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/56 24.9 ±28% 23.9 ±16% ~ (p=0.058 n=57+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/224 27.1 ±19% 26.6 ±19% ~ (p=0.122 n=57+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/3584 28.9 ±10% 28.7 ±10% ~ (p=0.405 n=56+57) BM_MapContainsMiss<Map<llvm::StringRef, int>>/57344 30.5 ± 7% 31.2 ± 7% +2.32% (p=0.000 n=57+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/917504 31.8 ± 7% 31.7 ± 7% ~ (p=0.713 n=57+56) BM_MapContainsMiss<Map<llvm::StringRef, int>>/14680064 33.4 ± 9% 33.5 ± 7% ~ (p=0.921 n=56+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/1/256 49.3 ± 0% 48.2 ± 0% -2.17% (p=0.000 n=55+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/2/256 49.3 ± 0% 48.2 ± 0% -2.17% (p=0.000 n=56+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/3/256 49.3 ± 0% 48.2 ± 0% -2.17% (p=0.000 n=54+55) BM_MapLookupHit<Map<llvm::StringRef, int>>/4/256 49.3 ± 0% 48.2 ± 0% -2.17% (p=0.000 n=54+53) BM_MapLookupHit<Map<llvm::StringRef, int>>/8/256 49.0 ± 0% 48.0 ± 0% -2.02% (p=0.000 n=51+51) BM_MapLookupHit<Map<llvm::StringRef, int>>/16/256 51.8 ± 1% 51.3 ± 1% -0.89% (p=0.000 n=50+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/32/256 51.8 ± 1% 51.3 ± 1% -1.07% (p=0.000 n=56+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/64/256 52.4 ± 1% 51.8 ± 1% -1.12% (p=0.000 n=57+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/256/256 54.6 ± 1% 54.1 ± 1% -0.94% (p=0.000 n=53+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/256/64 51.9 ± 1% 51.4 ± 1% -0.95% (p=0.000 n=55+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/256/128 52.5 ± 1% 52.0 ± 1% -1.07% (p=0.000 n=57+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/4096/256 62.0 ± 3% 61.6 ± 3% -0.62% (p=0.002 n=55+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/4096/1024 74.6 ± 1% 73.5 ± 1% -1.38% (p=0.000 n=56+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/4096/2048 80.9 ± 1% 79.8 ± 1% -1.34% (p=0.000 n=57+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/65536/256 72.0 ± 2% 71.4 ± 2% -0.77% (p=0.000 n=56+55) BM_MapLookupHit<Map<llvm::StringRef, int>>/65536/16384 145 ± 4% 145 ± 3% ~ (p=0.662 n=57+55) BM_MapLookupHit<Map<llvm::StringRef, int>>/65536/32768 155 ± 4% 156 ± 4% ~ (p=0.541 n=57+54) BM_MapLookupHit<Map<llvm::StringRef, int>>/1048576/256 73.1 ± 2% 72.5 ± 2% -0.73% (p=0.000 n=56+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/1048576/262144 281 ± 7% 283 ± 5% ~ (p=0.284 n=57+49) BM_MapLookupHit<Map<llvm::StringRef, int>>/1048576/524288 342 ± 5% 342 ± 5% ~ (p=0.684 n=57+53) BM_MapLookupHit<Map<llvm::StringRef, int>>/16777216/256 77.5 ± 2% 76.9 ± 2% -0.74% (p=0.000 n=55+54) BM_MapLookupHit<Map<llvm::StringRef, int>>/16777216/4194304 750 ± 3% 749 ± 3% ~ (p=0.458 n=57+53) BM_MapLookupHit<Map<llvm::StringRef, int>>/16777216/8388608 802 ± 2% 801 ± 3% ~ (p=0.518 n=57+55) BM_MapLookupHit<Map<llvm::StringRef, int>>/56/256 51.9 ± 1% 51.3 ± 1% -1.10% (p=0.000 n=57+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/224/256 54.0 ± 1% 53.5 ± 1% -1.01% (p=0.000 n=56+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/3584/256 58.8 ± 2% 58.1 ± 2% -1.28% (p=0.000 n=56+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/3584/896 69.7 ± 2% 68.7 ± 1% -1.35% (p=0.000 n=57+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/3584/1792 77.1 ± 1% 76.0 ± 1% -1.45% (p=0.000 n=55+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/57344/256 71.3 ± 2% 70.7 ± 3% -0.85% (p=0.000 n=55+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/57344/14336 128 ± 3% 128 ± 3% ~ (p=0.556 n=57+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/57344/28672 140 ± 4% 140 ± 3% ~ (p=0.735 n=57+51) BM_MapLookupHit<Map<llvm::StringRef, int>>/917504/256 72.8 ± 2% 72.3 ± 2% -0.76% (p=0.000 n=57+57) BM_MapLookupHit<Map<llvm::StringRef, int>>/917504/229376 242 ± 7% 243 ± 6% ~ (p=0.303 n=57+55) BM_MapLookupHit<Map<llvm::StringRef, int>>/917504/458752 264 ± 7% 264 ± 6% ~ (p=0.823 n=57+55) BM_MapLookupHit<Map<llvm::StringRef, int>>/14680064/256 76.4 ± 2% 75.8 ± 3% -0.78% (p=0.000 n=57+56) BM_MapLookupHit<Map<llvm::StringRef, int>>/14680064/3670016 696 ± 3% 698 ± 3% ~ (p=0.189 n=56+53) BM_MapLookupHit<Map<llvm::StringRef, int>>/14680064/7340032 749 ± 3% 750 ± 3% ~ (p=0.266 n=55+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/1/256 34.9 ± 0% 35.0 ± 0% +0.36% (p=0.000 n=56+50) BM_MapUpdateHit<Map<llvm::StringRef, int>>/2/256 34.9 ± 0% 35.0 ± 0% +0.35% (p=0.000 n=55+53) BM_MapUpdateHit<Map<llvm::StringRef, int>>/3/256 34.9 ± 0% 35.0 ± 0% +0.35% (p=0.000 n=55+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/4/256 34.9 ± 0% 35.0 ± 0% +0.36% (p=0.000 n=56+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/8/256 37.5 ± 3% 37.6 ± 2% ~ (p=0.081 n=57+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/16/256 39.4 ± 1% 39.5 ± 2% ~ (p=0.054 n=55+57) BM_MapUpdateHit<Map<llvm::StringRef, int>>/32/256 40.0 ± 3% 39.9 ± 4% ~ (p=0.449 n=56+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/64/256 40.0 ± 1% 40.1 ± 2% ~ (p=0.796 n=54+54) BM_MapUpdateHit<Map<llvm::StringRef, int>>/256/256 41.1 ± 2% 41.2 ± 2% ~ (p=0.061 n=53+50) BM_MapUpdateHit<Map<llvm::StringRef, int>>/256/64 39.6 ± 2% 39.6 ± 2% ~ (p=0.695 n=55+52) BM_MapUpdateHit<Map<llvm::StringRef, int>>/256/128 40.2 ± 2% 40.1 ± 2% ~ (p=0.507 n=53+49) BM_MapUpdateHit<Map<llvm::StringRef, int>>/4096/256 43.4 ± 2% 43.5 ± 2% ~ (p=0.300 n=53+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/4096/1024 50.9 ± 2% 51.8 ± 2% +1.79% (p=0.000 n=56+57) BM_MapUpdateHit<Map<llvm::StringRef, int>>/4096/2048 58.2 ± 1% 58.3 ± 1% ~ (p=0.072 n=57+57) BM_MapUpdateHit<Map<llvm::StringRef, int>>/65536/256 46.1 ± 1% 46.1 ± 2% ~ (p=0.197 n=54+53) BM_MapUpdateHit<Map<llvm::StringRef, int>>/65536/16384 88.1 ± 5% 88.9 ± 4% +0.90% (p=0.011 n=57+57) BM_MapUpdateHit<Map<llvm::StringRef, int>>/65536/32768 92.4 ± 3% 93.6 ± 3% +1.35% (p=0.000 n=57+57) BM_MapUpdateHit<Map<llvm::StringRef, int>>/1048576/256 46.6 ± 2% 46.7 ± 2% ~ (p=0.687 n=51+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/1048576/262144 144 ± 7% 145 ± 6% ~ (p=0.130 n=57+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/1048576/524288 181 ± 4% 182 ± 4% ~ (p=0.057 n=56+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/16777216/256 48.9 ± 2% 48.7 ± 2% -0.30% (p=0.042 n=56+53) BM_MapUpdateHit<Map<llvm::StringRef, int>>/16777216/4194304 351 ± 2% 350 ± 3% ~ (p=0.287 n=57+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/16777216/8388608 368 ± 3% 367 ± 3% ~ (p=0.710 n=57+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/56/256 39.7 ± 3% 39.6 ± 3% ~ (p=0.572 n=57+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/224/256 41.7 ± 2% 41.6 ± 3% ~ (p=0.233 n=55+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/3584/256 42.6 ± 2% 42.5 ± 2% ~ (p=0.309 n=54+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/3584/896 49.1 ± 1% 49.8 ± 1% +1.51% (p=0.000 n=57+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/3584/1792 57.0 ± 1% 57.1 ± 2% +0.30% (p=0.022 n=56+57) BM_MapUpdateHit<Map<llvm::StringRef, int>>/57344/256 46.1 ± 2% 46.0 ± 1% -0.28% (p=0.013 n=55+53) BM_MapUpdateHit<Map<llvm::StringRef, int>>/57344/14336 82.0 ± 2% 82.6 ± 2% +0.71% (p=0.000 n=57+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/57344/28672 88.7 ± 2% 89.8 ± 2% +1.22% (p=0.000 n=57+53) BM_MapUpdateHit<Map<llvm::StringRef, int>>/917504/256 46.5 ± 1% 46.5 ± 2% ~ (p=0.961 n=53+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/917504/229376 126 ± 5% 128 ± 5% +1.64% (p=0.000 n=57+54) BM_MapUpdateHit<Map<llvm::StringRef, int>>/917504/458752 140 ± 5% 141 ± 6% ~ (p=0.162 n=57+55) BM_MapUpdateHit<Map<llvm::StringRef, int>>/14680064/256 48.5 ± 2% 48.3 ± 2% ~ (p=0.094 n=55+54) BM_MapUpdateHit<Map<llvm::StringRef, int>>/14680064/3670016 328 ± 4% 328 ± 3% ~ (p=0.925 n=57+56) BM_MapUpdateHit<Map<llvm::StringRef, int>>/14680064/7340032 345 ± 3% 345 ± 3% ~ (p=0.489 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/1/256 76.0 ± 0% 75.9 ± 0% -0.03% (p=0.006 n=54+47) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/2/256 71.3 ± 1% 72.1 ± 5% ~ (p=0.750 n=52+54) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/3/256 70.9 ± 2% 70.7 ± 2% ~ (p=0.095 n=54+52) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/4/256 70.4 ± 2% 70.6 ± 3% ~ (p=0.458 n=47+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/8/256 75.0 ± 1% 74.2 ± 1% -1.16% (p=0.000 n=52+54) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/16/256 80.5 ± 3% 79.0 ± 3% -1.88% (p=0.000 n=51+53) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/32/256 83.2 ± 4% 82.3 ± 5% -1.01% (p=0.009 n=56+57) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/64/256 80.6 ± 3% 79.4 ± 4% -1.48% (p=0.000 n=52+54) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/256/256 83.6 ± 3% 82.6 ± 5% -1.23% (p=0.000 n=54+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/256/64 79.1 ± 6% 78.8 ± 8% ~ (p=0.359 n=55+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/256/128 81.1 ± 6% 80.3 ± 9% -1.04% (p=0.010 n=55+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/4096/256 85.5 ± 5% 84.1 ± 4% -1.61% (p=0.000 n=54+57) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/4096/1024 95.7 ± 3% 95.2 ± 2% -0.47% (p=0.033 n=56+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/4096/2048 101 ± 2% 101 ± 1% -0.68% (p=0.000 n=56+54) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/65536/256 90.5 ± 4% 88.1 ± 4% -2.57% (p=0.000 n=56+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/65536/16384 134 ± 3% 133 ± 2% -0.71% (p=0.002 n=57+57) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/65536/32768 142 ± 3% 141 ± 2% -0.90% (p=0.000 n=57+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/1048576/256 91.2 ± 3% 89.3 ± 4% -2.08% (p=0.000 n=56+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/1048576/262144 209 ± 5% 208 ± 5% ~ (p=0.170 n=57+54) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/1048576/524288 243 ± 5% 240 ± 5% -1.05% (p=0.020 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/16777216/256 94.3 ± 3% 92.5 ± 5% -1.91% (p=0.000 n=55+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/16777216/4194304 542 ± 3% 537 ± 4% -1.02% (p=0.000 n=57+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/16777216/8388608 566 ± 3% 561 ± 4% -1.01% (p=0.000 n=57+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/56/256 83.7 ±10% 81.3 ± 8% -2.84% (p=0.000 n=55+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/224/256 88.7 ± 8% 86.6 ± 9% -2.40% (p=0.001 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/3584/256 94.0 ± 5% 91.3 ± 4% -2.83% (p=0.000 n=56+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/3584/896 118 ± 4% 118 ± 5% ~ (p=0.930 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/3584/1792 143 ± 4% 141 ± 4% -1.10% (p=0.002 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/57344/256 102 ± 4% 100 ± 4% -2.31% (p=0.000 n=56+57) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/57344/14336 191 ± 2% 190 ± 1% -0.32% (p=0.024 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/57344/28672 197 ± 2% 197 ± 1% ~ (p=0.059 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/917504/256 103 ± 4% 101 ± 4% -1.99% (p=0.000 n=57+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/917504/229376 280 ± 3% 279 ± 3% ~ (p=0.145 n=57+52) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/917504/458752 298 ± 4% 296 ± 3% ~ (p=0.116 n=57+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/14680064/256 107 ± 4% 104 ± 4% -2.11% (p=0.000 n=55+56) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/14680064/3670016 613 ± 3% 612 ± 2% ~ (p=0.224 n=57+55) BM_MapEraseUpdateHit<Map<llvm::StringRef, int>>/14680064/7340032 637 ± 2% 635 ± 1% ~ (p=0.075 n=56+55) BM_MapInsertSeq<Map<llvm::StringRef, int>>/1 132 ± 0% 132 ± 0% -0.26% (p=0.000 n=47+41) BM_MapInsertSeq<Map<llvm::StringRef, int>>/2 160 ± 0% 161 ± 4% +0.57% (p=0.001 n=45+57) BM_MapInsertSeq<Map<llvm::StringRef, int>>/3 188 ± 2% 189 ± 3% ~ (p=0.327 n=54+54) BM_MapInsertSeq<Map<llvm::StringRef, int>>/4 217 ± 3% 218 ± 5% ~ (p=0.240 n=54+56) BM_MapInsertSeq<Map<llvm::StringRef, int>>/8 342 ± 5% 341 ± 4% ~ (p=0.282 n=53+54) BM_MapInsertSeq<Map<llvm::StringRef, int>>/16 640 ± 3% 648 ± 8% +1.26% (p=0.023 n=49+54) BM_MapInsertSeq<Map<llvm::StringRef, int>>/32 1.20k ± 8% 1.20k ± 8% ~ (p=0.423 n=53+54) BM_MapInsertSeq<Map<llvm::StringRef, int>>/64 3.57k ± 8% 3.55k ± 6% ~ (p=0.557 n=57+54) BM_MapInsertSeq<Map<llvm::StringRef, int>>/256 18.6k ± 5% 18.6k ± 6% ~ (p=0.799 n=57+57) BM_MapInsertSeq<Map<llvm::StringRef, int>>/4096 492k ± 4% 491k ± 3% ~ (p=0.378 n=56+57) BM_MapInsertSeq<Map<llvm::StringRef, int>>/65536 10.5M ± 2% 10.4M ± 1% ~ (p=0.143 n=57+48) BM_MapInsertSeq<Map<llvm::StringRef, int>>/1048576 323M ± 2% 322M ± 3% ~ (p=0.098 n=56+56) BM_MapInsertSeq<Map<ll::StringRef, int>>/16777216 7.07G ± 3% 7.05G ± 4% ~ (p=0.195 n=56+57) BM_MapInsertSeq<Map<llvm::StringRef, int>>/56 2.04k ± 8% 2.03k ± 7% ~ (p=0.124 n=52+55) BM_MapInsertSeq<Map<llvm::StringRef, int>>/224 12.0k ± 5% 12.0k ± 4% ~ (p=0.467 n=57+55) BM_MapInsertSeq<Map<llvm::StringRef, int>>/3584 294k ± 5% 292k ± 4% ~ (p=0.188 n=56+57) BM_MapInsertSeq<Map<llvm::StringRef, int>>/57344 6.40M ± 2% 6.39M ± 1% ~ (p=0.381 n=57+56) BM_MapInsertSeq<Map<llvm::StringRef, int>>/917504 199M ± 3% 199M ± 3% ~ (p=0.977 n=57+57) BM_MapInsertSeq<Map<llvm::StringRef, int>>/14680064 4.56G ± 3% 4.55G ± 3% ~ (p=0.129 n=55+56) ``` |
||
|
|
9d09061301 |
Avoid misaligned loads from StaticRandomData in the size [4, 8] hashing case. (#4743)
Avoid misaligned loads from StaticRandomData in the size [4, 8] hashing case. We can use aligned loads in this case for lower latency. We introduce the SampleAlignedRandomData function for this purpose. |
||
|
|
afdf846636 |
Align StaticRandomData to cacheline size. (#4741)
Align StaticRandomData to cacheline size to ensure the whole array is on the same cacheline. |
||
|
|
c832d523be |
Update files and clang-tidy config to pass with clang-tidy-20 (#4691)
Disables three new warnings because they lean more towards style conflicts than fixes. I've brought these up on #style. Other than that, mostly fixing basic issues, and things that clang-tidy-20 seems to fire where clang-tiday-16 didn't. One particular curious case is `llvm::StringLiteral::data()` uses, which are flagged as not strictly null-terminated; I'm switching to `const char*` in those spots which matches `llvm::formatv`'s format argument, but feels worse. I'm removing `run_clang_tidy.py` here because I'm observing it give fewer warnings than `bazel build --config=clang-tidy -k //toolchain/...`. The latter matches how we enforce in GitHub actions (and also caches results, and suppresses output for files that have no issues), so I'm dropping the bespoke script. |
||
|
|
bf736e6b03 |
A collection of hashing improvements from using hashtables. (#4094)
LLVM's `APInt` and `APFloat` need specialized handling to be used effectively in hashtables. We can't inject overrides into LLVM so we need to handle them in our hashing routine. There were also problematic limits on hashing pairs and tuples. First, the unique-object-representation hashing of pairs was more restricted than tuples which was a problematic asymmetry and isn't needed. But the larger issue is that we didn't support recursively hashing when necessary. That requires a careful predicate to avoid infinite recursion but lets us handle important use cases for hashtables with a tuple as a key. Also added support for hashing arrays that recurse in addition to arrays where we can hash the raw storage, and added overloads to redirect to common array handling from various array-like types. Last but not least, re-worked the constraint model for hashing as raw data to not override custom hashing functions. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com> |
||
|
|
21a81bc59e |
Introduce custom hash table data structures. (#3940)
The hash table design is heavily based on Abseil's ["Swiss Tables"][swiss-tables] design. It uses an array of bytes storing metadata about each entry and an array of entries where each is a pair of key and value. The metadata byte consists of 7-bits of hash of the key (distinct from the bits used to index the table), and one bit indicating the presence of a special entry -- either empty or deleted. [swiss-tables]: https://abseil.io/about/design/swisstables There are a large range of optimizations and other nuanced aspects of this hash table design and implementation, a good point to understand that context is `raw_hashtable.h` which has an overview of the design and references to various other files for relevant details. --------- Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com> |
||
|
|
5c8fa6ad5c |
Replace FoldingSet with DenseMap for instruction canonicalization. (#3979)
Switch from recursing into non-canonical instruction fields to separately canonicalizing those fields. This means we now form canonical `InstBlockId`s, `TypeBlockId`s, `IntId`s, `FloatId`s, and `BindNameId`s at least in the cases when they're referenced by a constant instruction. This reduces the overall runtime for @chandlerc's 10MLoC example by 27.5% on my machine. |
||
|
|
b473eac5bc |
Fix clang-tidy issues in //common. (#3962)
These likely predate the CI integration for `clang-tidy` runs. Most of these seem good generally, even though I disabled some with nolint comments. The multilevel pointer one seems almost like a bug in the check to detect the specific case of `memcpy`, but otherwise seems like a solid lint. |
||
|
|
fea2651e7c |
chore: fix typos (#3738)
Signed-off-by: cui fliter <imcusg@gmail.com> |
||
|
|
2e236759ca |
Switch //common to use C++20 concepts. (#3665)
This removes the use of `enable_if` and tries to adopt concepts instead of type traits when available. The `ostream.h` change is a bit subtle as it adds a restriction not previously in place -- that the stream is *contvertible* to `std::ostream` as well as having it as a base class. This seems to match the intent of the code. The `hashing.h` code adds an implementation detail concept, and so I've also clarified that the dispatch namespace is an internal one that isn't part of the public API. |
||
|
|
bf02d1f4b0 |
Remove headers marked as unused by ClangD. (#3661)
This required adding a few headers that were found transitively before, but not too many. This is sadly a fairly manual process of opening every file in my IDE, but I think I got everything in `//common` and `//toolchain`. There are a few cases where technically we don't need `foo.h` to be included into `foo.cpp`, but I've forced those to stay with a pragma. I've tried to catch the places where we can cut deps in Bazel as well, but not sure I got all of those. I had been noticing these in other PRs and it seemed better to isolate the change. |
||
|
|
ebbdc11877 |
Switch to a "better" multiplicative hash constant. (#3629)
Testing this hash function with representative hash table implementations showed significant differences in quality between different multiplicative hashing constants. The constants used and documented were OK but had clear limitations when merely using a single 64-bit multiplication. However, a search uncovered (partly by luck) a constant that has empirically been shown to be both significantly better than other constants and generally not have problematic weaknesses. There are still plenty of collisions for string keys and heavily loaded hash tables of course, but no examples of severe outlier collision rates as observed with all other constants we have tried. There is a slightly longer comment explaining some of this context and the other constants we have tried in the PR as well. To this day, we still don't fully understand why the constant used here behaves so much better than other constants we have tried, including all of those we've found in other hashing algorithms. |
||
|
|
5f62cf752d |
Fix an oversight that dropped the buffer. (#3628)
This lost any seed or prior hashing done, which isn't good. I've added some basic testing that would have caught this immediately. |
||
|
|
7e9760d9e4 |
Simplify the index & tag extraction API for hash codes. (#3627)
The fancier API ended up not being helpful and making it harder to optimize a hash table implemented on top of this. |
||
|
|
f59a6cdbdd |
Introduce a Carbon hashing framework. (#3327)
# Overview This is a latency-optimized hashing framework based on Abseil's and others. At it's core it uses both a normal 64-bit multiply as well as a 64-bit multiply capturing both low and high 64-bit components of the result and XOR-ing them together. These are the primitives used in FxHash and Abseil respectively, although they both appear in others. The implementation has been *substantially* optimized for short inputs and latency over quality. As a result, this function does not remotely pass the SMHasher quality tests. However, basic collisions are rare, and I've included a small subset of the SMHasher collision testing directly to make sure the quality doesn't slip too far inadvertently. The customization framework is roughly similar to Abseil's and LLVM's but has been simplified significantly, inspired in some respects by the AHash API design and in others by my experience of all performance sensitive hashing implementations needing to work at a very low level to hit their performance targets. The abstractions are stripped down to facilitate this. # Details of the performance optimization This function is 2x - 4x faster than LLVM's on small inputs, and up to 2x faster than Abseil. Significant effort has gone into optimizing short strings in particular compared to Abseil. Small integer and pointer hashing is also faster than Abseil's by leveraging a lower quality 64-bit multiply in some cases inspired by FxHash. One consequence is that this routine is particulary fast for 32-bit integers. The short string improvements largely come from packing more of the bytes of string into as few multiplies as possible. While this fails to mix the bits sufficient to hit SMHasher's strict avalanche criteria and does leave some collision windows, it provides dramatic latency improvements. Some of these techniques come from Abseil's own bulk hashing routine but re-applied here. Others are novel, for example using small sizes to sample nicely uniform random data to efficiently handle the very small number of bits of data that need to be hashed. The other observed improvement is diligent handling of pairs and tuples and fairly aggressively turning things into integers. Some of the comparisons with Abseil aren't realistic as the Abseil hash table does some of these mappings before hashing. I've done this directly in the hash function as that seems cleaner. For long strings, the performance is comparable or a bit better than Abseil, and significantly better than LLVM's hash function. Overall, for short inputs this is hoped to be the fastest hash function that still gets "just enough" mixing for modern hash tables to perform well. # Details of the quality vs. latency tradeoff A key insight is that modern hash tables don't need especially high quality hash functions, but do benefit from something beyond the identify function. That isn't the target of SMHasher or other quality assessing tools and has resulted in unnecessarily aggressive hashing for any functions actually evaluated against it. Many hash functions turn off the high quality implementations evaluated with SMHasher for integer or pointer keys to recover latency & performance (AHash for example), but the same performance-oriented design applies beyond these narrow types, for example for short strings. However, a consequence is that there are serious limits to the quality of the hash function. The avalanche test is failed hilariously, etc., but in the exact same ways as Abseil itself fails it for integer keys. There are also real collisions spaces. For example, for 16-byte strings, there is one 64-bit value for the first 8 bytes that will have the same hash regardless of the other 8 bytes of the string. Some minor effort is taken to make this pattern unlikely to be a practical problem, but it is a clear theoretical weakness. It also means that this hash function couldn't be further from providing any hash-flooding DoS attack protection -- I expect it to be trivially easy to attack in this way by a motivated adversary. Defending against these attacks is defined as out-of-scope, in large part because even attempts that have made a compelling effort to address these issues such as HighwayHash have found serious limits. Instead, this takes a principled position that any such defense should be provided entirely at the data structure level with a strong worst-case bound rather than through strengthening the hash function. # Future work A subsequent PR will introduce a hash table inspired very heavily by the design of Abseil's "SwissTable" and using this hash function. The goal is to provide a significant improvement to hot hash tables such as the identifier table in the lexer of Carbon's toolchain. # Detailed benchmark data The benchmarks introduced are heavily inspired by the latency benchmarking of hash functions in Abseil. I've adapted them to fit better into Carbon's coding style and to try to have more stable results with broader coverage of types and string sizes. Running the benchmarks directly gives horizontal comparisons across different hash functions. That can be hard to read, so here is *just* the newly introduced hash function benchmark results on an AMD server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 3.11ns ± 1% BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 3.11ns ± 1% BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 4.11ns ± 1% BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 3.12ns ± 1% BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 4.13ns ± 1% BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 3.16ns ± 2% BM_LatencyHash<RandValues<int*>, CarbonHashBench> 3.16ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 4.03ns ± 2% BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 4.04ns ± 1% BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 4.34ns ± 2% BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 4.04ns ± 1% BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 4.34ns ± 2% BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 4.34ns ± 1% BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 4.33ns ± 1% BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 1.95ns ± 4% BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 1.70ns ± 3% BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 3.52ns ± 3% BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 4.46ns ± 2% BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 7.69ns ± 1% BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 14.8ns ± 1% BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 21.5ns ± 1% BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 34.6ns ± 0% BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 63.1ns ± 1% BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 118ns ± 1% BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 225ns ± 1% ``` And on an ARM server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 5.28ns ± 0% BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 5.29ns ± 0% BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 7.02ns ± 0% BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 5.34ns ± 1% BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 7.07ns ± 4% BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 5.36ns ± 2% BM_LatencyHash<RandValues<int*>, CarbonHashBench> 5.36ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 7.19ns ± 3% BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 7.29ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 7.31ns ± 4% BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 7.29ns ± 2% BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 7.31ns ± 4% BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 8.69ns ± 3% BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 2.64ns ± 2% BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 2.90ns ± 4% BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 6.14ns ± 1% BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 8.27ns ± 1% BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 13.8ns ± 0% BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 31.2ns ± 0% BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 49.9ns ± 0% BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 86.9ns ± 0% BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 163ns ± 0% BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 312ns ± 0% BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 610ns ± 0% ``` I don't have the same nice statistical multi-run error bars, but one run from my M1 MacBook: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 3.89 ns BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 3.87 ns BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 4.39 ns BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 3.93 ns BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 4.98 ns BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 3.87 ns BM_LatencyHash<RandValues<int*>, CarbonHashBench> 3.87 ns BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 4.86 ns BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 4.43 ns BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 4.41 ns BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 4.44 ns BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 4.69 ns BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 4.33 ns BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 4.38 ns BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 4.34 ns BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 4.35 ns BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 4.38 ns BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 1.15 ns BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 0.973 ns BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 3.03 ns BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 3.97 ns BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 6.64 ns BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 12.5 ns BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 17.9 ns BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 27.9 ns BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 48.1 ns BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 87.3 ns BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 166 ns ``` And here I have internally replaced the Carbon hash function with Abseil's hash function for "before" and then restored it in the "after" and computed the delta for each benchmark. This basically shows the speed-up (lower time -> lower latency -> speed-up -> good) over Abseil on an AMD server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 4.00ns ± 1% 3.10ns ± 0% -22.45% (p=0.000 n=20+15) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 4.01ns ± 1% 3.10ns ± 1% -22.64% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 6.25ns ± 1% 4.10ns ± 1% -34.30% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 4.02ns ± 1% 3.12ns ± 1% -22.50% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 6.25ns ± 1% 4.11ns ± 1% -34.20% (p=0.000 n=20+19) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 4.03ns ± 1% 3.14ns ± 1% -22.17% (p=0.000 n=19+19) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 5.95ns ± 1% 3.14ns ± 1% -47.24% (p=0.000 n=20+18) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 6.04ns ± 1% 4.01ns ± 1% -33.64% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 5.96ns ± 1% 4.02ns ± 1% -32.51% (p=0.000 n=18+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 5.93ns ± 1% 4.30ns ± 1% -27.56% (p=0.000 n=20+17) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 7.97ns ± 1% 4.02ns ± 1% -49.50% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 7.98ns ± 1% 4.32ns ± 1% -45.88% (p=0.000 n=19+20) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 4.40ns ± 2% 4.32ns ± 1% -1.81% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 5.94ns ± 1% 4.32ns ± 1% -27.25% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 10.0ns ± 1% 4.3ns ± 1% -56.56% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 8.04ns ± 1% 4.32ns ± 1% -46.29% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 7.95ns ± 1% 4.33ns ± 1% -45.59% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 3.28ns ± 3% 1.93ns ± 4% -41.19% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 3.05ns ± 3% 1.69ns ± 4% -44.52% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 5.88ns ± 2% 3.50ns ± 3% -40.42% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 8.92ns ± 1% 4.44ns ± 2% -50.22% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 12.0ns ± 1% 7.7ns ± 1% -36.16% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 18.8ns ± 0% 14.7ns ± 1% -21.73% (p=0.000 n=17+20) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 25.5ns ± 1% 21.4ns ± 1% -16.18% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 38.7ns ± 2% 34.5ns ± 1% -10.78% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 69.7ns ± 1% 62.8ns ± 1% -9.88% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 130ns ± 1% 117ns ± 1% -9.45% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 244ns ± 0% 225ns ± 1% -8.11% (p=0.000 n=17+20) ``` ... and on an ARM server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 6.48ns ± 1% 5.28ns ± 0% -18.62% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 7.40ns ± 1% 5.29ns ± 1% -28.45% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 10.4ns ± 0% 7.0ns ± 0% -32.34% (p=0.000 n=19+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 6.56ns ± 1% 5.32ns ± 1% -18.95% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 10.8ns ± 2% 7.0ns ± 1% -34.89% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 6.71ns ± 3% 5.38ns ± 2% -19.84% (p=0.000 n=20+20) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 10.3ns ± 3% 5.4ns ± 2% -47.67% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 10.9ns ± 2% 7.2ns ± 4% -33.67% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 10.7ns ± 4% 7.3ns ± 4% -31.66% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 10.5ns ± 3% 7.3ns ± 4% -30.71% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 14.1ns ± 3% 7.3ns ± 4% -48.32% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 14.0ns ± 1% 7.3ns ± 4% -47.95% (p=0.000 n=19+19) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 9.41ns ± 4% 8.68ns ± 4% -7.71% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 12.2ns ± 2% 8.7ns ± 4% -28.81% (p=0.000 n=18+19) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 18.9ns ± 2% 8.7ns ± 4% -54.17% (p=0.000 n=17+19) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 15.6ns ± 2% 8.7ns ± 4% -44.37% (p=0.000 n=17+19) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 15.5ns ± 2% 8.7ns ± 4% -44.08% (p=0.000 n=18+19) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 5.89ns ± 2% 2.64ns ± 3% -55.26% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 5.73ns ± 3% 2.88ns ± 3% -49.71% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 10.1ns ± 1% 6.1ns ± 2% -39.00% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 15.7ns ± 0% 8.3ns ± 1% -47.27% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 21.2ns ± 0% 13.8ns ± 0% -34.81% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 37.9ns ± 0% 31.2ns ± 0% -17.77% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 56.8ns ± 0% 49.8ns ± 0% -12.21% (p=0.000 n=20+18) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 93.8ns ± 0% 86.9ns ± 0% -7.38% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 174ns ± 0% 163ns ± 0% -6.03% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 330ns ± 0% 312ns ± 0% -5.25% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 641ns ± 0% 610ns ± 0% -4.79% (p=0.000 n=19+19) ``` This is the same as the above delta comparison, but with the "before" being LLVM's hash function: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 6.85ns ± 1% 3.10ns ± 1% -54.78% (p=0.000 n=20+19) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 6.85ns ± 1% 3.10ns ± 1% -54.78% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 6.25ns ± 1% 4.09ns ± 1% -34.58% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 6.87ns ± 1% 3.12ns ± 2% -54.66% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 7.35ns ± 1% 4.10ns ± 1% -44.20% (p=0.000 n=20+19) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 7.34ns ± 1% 3.13ns ± 1% -57.34% (p=0.000 n=20+18) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 7.33ns ± 1% 3.13ns ± 2% -57.27% (p=0.000 n=20+18) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 7.27ns ± 1% 3.99ns ± 1% -45.12% (p=0.000 n=20+18) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 14.5ns ± 1% 4.0ns ± 1% -72.23% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 14.6ns ± 1% 4.3ns ± 2% -70.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 14.5ns ± 1% 4.0ns ± 1% -72.21% (p=0.000 n=20+19) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 14.6ns ± 1% 4.3ns ± 1% -70.46% (p=0.000 n=20+18) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 7.31ns ± 1% 4.33ns ± 1% -40.81% (p=0.000 n=18+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 7.78ns ± 1% 4.32ns ± 1% -44.45% (p=0.000 n=18+20) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 7.78ns ± 2% 4.33ns ± 1% -44.42% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 7.62ns ± 1% 4.32ns ± 1% -43.24% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 7.77ns ± 1% 4.33ns ± 1% -44.34% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 8.15ns ± 3% 1.94ns ± 5% -76.16% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 7.02ns ± 3% 1.69ns ± 4% -75.94% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 7.83ns ± 2% 3.50ns ± 3% -55.34% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 9.17ns ± 1% 4.43ns ± 2% -51.65% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 11.3ns ± 1% 7.6ns ± 1% -32.04% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 23.0ns ± 1% 14.7ns ± 1% -36.14% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 32.9ns ± 0% 21.4ns ± 1% -34.96% (p=0.000 n=17+19) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 52.2ns ± 1% 34.4ns ± 1% -34.01% (p=0.000 n=19+18) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 92.1ns ± 1% 62.8ns ± 1% -31.82% (p=0.000 n=19+19) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 169ns ± 1% 117ns ± 1% -30.53% (p=0.000 n=20+19) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 319ns ± 1% 224ns ± 1% -29.78% (p=0.000 n=20+18) ``` ... and on an ARM server: ``` BM_LatencyHash<RandValues<uint8_t>, CarbonHashBench> 8.38ns ± 0% 5.27ns ± 0% -37.04% (p=0.000 n=20+20) BM_LatencyHash<RandValues<uint16_t>, CarbonHashBench> 8.39ns ± 1% 5.28ns ± 0% -37.01% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<uint8_t, uint8_t>>, CarbonHashBench> 8.07ns ± 0% 7.02ns ± 0% -13.10% (p=0.000 n=19+20) BM_LatencyHash<RandValues<uint32_t>, CarbonHashBench> 8.48ns ± 1% 5.32ns ± 1% -37.25% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint16_t, uint16_t>>, CarbonHashBench> 9.34ns ± 2% 7.09ns ± 2% -24.14% (p=0.000 n=19+20) BM_LatencyHash<RandValues<uint64_t>, CarbonHashBench> 9.76ns ± 3% 5.37ns ± 2% -44.98% (p=0.000 n=20+20) BM_LatencyHash<RandValues<int*>, CarbonHashBench> 9.76ns ± 3% 5.37ns ± 2% -44.98% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint32_t>>, CarbonHashBench> 10.1ns ± 2% 7.2ns ± 3% -29.36% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint32_t>>, CarbonHashBench> 11.9ns ± 2% 7.3ns ± 4% -38.68% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, uint64_t>>, CarbonHashBench> 11.3ns ± 2% 7.3ns ± 4% -35.16% (p=0.000 n=19+19) BM_LatencyHash<RandValues<std::pair<int*, uint32_t>>, CarbonHashBench> 11.9ns ± 2% 7.3ns ± 4% -38.68% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint32_t, int*>>, CarbonHashBench> 11.3ns ± 2% 7.3ns ± 4% -35.16% (p=0.000 n=19+19) BM_LatencyHash<RandValues<__uint128_t>, CarbonHashBench> 10.3ns ± 2% 8.7ns ± 3% -15.81% (p=0.000 n=19+20) BM_LatencyHash<RandValues<std::pair<uint64_t, uint64_t>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, int*>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<uint64_t, int*>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandValues<std::pair<int*, uint64_t>>, CarbonHashBench> 11.6ns ± 3% 8.7ns ± 3% -25.44% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4>, CarbonHashBench> 9.39ns ± 2% 2.66ns ± 3% -71.66% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 8>, CarbonHashBench> 10.7ns ± 3% 2.9ns ± 3% -72.97% (p=0.000 n=19+18) BM_LatencyHash<RandStrings< true, 16>, CarbonHashBench> 11.8ns ± 1% 6.1ns ± 2% -47.75% (p=0.000 n=19+20) BM_LatencyHash<RandStrings< true, 32>, CarbonHashBench> 13.9ns ± 1% 8.3ns ± 1% -40.71% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 64>, CarbonHashBench> 16.8ns ± 1% 13.8ns ± 0% -17.83% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 256>, CarbonHashBench> 31.7ns ± 1% 31.2ns ± 0% -1.76% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 512>, CarbonHashBench> 43.5ns ± 0% 49.8ns ± 0% +14.56% (p=0.000 n=18+20) BM_LatencyHash<RandStrings< true, 1024>, CarbonHashBench> 66.2ns ± 0% 86.9ns ± 0% +31.39% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 2048>, CarbonHashBench> 112ns ± 0% 163ns ± 0% +46.09% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 4096>, CarbonHashBench> 201ns ± 0% 312ns ± 0% +55.49% (p=0.000 n=20+20) BM_LatencyHash<RandStrings< true, 8192>, CarbonHashBench> 379ns ± 0% 610ns ± 0% +61.08% (p=0.000 n=20+20) ``` Note that there is a significant regression on long strings compared to LLVM's hash function on the ARM server I have access to. This doesn't show up on the M1 at all, and is likely specific to inadequate throughput for the 64-bit multiply operations. This seems fine as a) our priority is for short strings, and b) the M1 and other ARM CPUs are likely to improve here over time given the prevalent use of this core technique. For example, Abseil's current hash algorithm has the same long-string behavior (and performance bottleneck) on this server. --------- Co-authored-by: josh11b <josh11b@users.noreply.github.com> Co-authored-by: Geoff Romer <gromer@google.com> |