Port the toolchain to use the new Carbon hashtable (#4097)

This works to leverage the capabilities of the hashtable as much as
possible, for example using the key context in the value stores.
However, there may still be opportunities to refactor more deeply and
use the functionality even better. Hopefully this is at least
a reasonable start and gets us a clean baseline.

On an Arm M1, this is a 15% improvement on my large lexing stress test,
but ends up a wash on my x86-64 server. This is a smaller benefit than
I expected, and it's because we're using a set-of-IDs and looking up
values with a key context for things like identifiers. This pattern has
a surprising tradeoff. The new hashtable uses significantly less memory,
a 10% peak RSS reduction just from the hashtable change. But indirecting
through the vector of values makes growing the hashtable dramatically
less cache-friendly: it causes growth to randomly access every key when
rehashing. On x86, everything gained by the faster hashtable is lost in
even slower growth. And even on Arm, this eats into the benefits.

But I have a plan to tweak how identifiers specifically work to avoid
most of the growth, and so I suspect this is the right tradeoff on the
whole. It gives us significant working set size reduction and we can
likely avoid the regressed operation (growth with rehash) in most cases
by clever reserving and if necessary by adding a hash caching layer to
the table infrastructure.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
This commit is contained in:
Chandler Carruth
2024-07-03 01:10:44 +00:00
committed by GitHub
co-authored by Jon Ross-Perkins
parent f5f8342542
commit 8992d22ab3
27 changed files with 292 additions and 353 deletions
+11 -7
View File
@@ -147,13 +147,17 @@ auto DeclNameStack::AddName(NameContext name_context, SemIR::InstId target_id,
context_->AddExport(target_id);
}
int scope_index = name_scope.names.size();
name_scope.names.push_back({.name_id = name_context.unresolved_name_id,
.inst_id = target_id,
.access_kind = access_kind});
auto [_, success] = name_scope.name_map.insert(
{name_context.unresolved_name_id, scope_index});
CARBON_CHECK(success)
auto add_scope = [&] {
int index = name_scope.names.size();
name_scope.names.push_back(
{.name_id = name_context.unresolved_name_id,
.inst_id = target_id,
.access_kind = access_kind});
return index;
};
auto result = name_scope.name_map.Insert(
name_context.unresolved_name_id, add_scope);
CARBON_CHECK(result.is_inserted())
<< "Duplicate names should have been resolved previously: "
<< name_context.unresolved_name_id << " in "
<< name_context.parent_scope_id;