This is reducing ValueStore inference of types from `using`, and removes `using ValueType = ...` from affected id types. I'm adding a number of `using FooStore = ValueStore<FooId, Foo>` because I think it's a little repetitive otherwise; often 4 cases where I'm doing this: getter, const getter, member, and getter on `Context`. Note we also have a number of `-> decltype(auto)` that were added I think mainly to avoid repeating the type, but I'm not sure whether there'll be agreement on replacing those and so am not changing them here. I'm placing these aliases with the value type in general, because I think it's probably easier to view that way. An alternative would be to put all the types on `File`, but: - That would be inconsistent with things like `InstStore`, which are very `ValueStore`-adjacent and put with their value type. - `File` would have a _lot_ of using's, and the accessors are already noisy -- I think it would just make the file harder to skim. Note this is the heart of what I'd brought up [on Discord](https://discord.com/channels/655572317891461132/655578254970716160/1388199282250613019). This PR still leaves CanonicalValueStore and BlockValueStore as things to also add parameters to, but I thought it best to try breaking the set of changes apart by type. Both of those rely on ValueStore, so ValueStore needs to change first.
Toolchain architecture
Table of contents
Goals
The toolchain represents the production portion of Carbon. At a high level, the toolchain's top priorities are:
- Correctness.
- Quality of generated code, including performance.
- Compilation performance.
- Quality of diagnostics for incorrect or questionable code.
TODO: Add an expanded document that details the goals and priorities and link to it here.
High-level architecture
The main components are:
-
Driver: Provides commands and ties together compilation flow.
-
Diagnostics: Produces diagnostic output.
-
Compilation flow:
- Source: Load the file into a SourceBuffer.
- Lex: Transform a SourceBuffer into a Lex::TokenizedBuffer.
- Parse: Transform a TokenizedBuffer into a Parse::Tree.
- Check: Transform a Tree to produce SemIR::File.
- Lower: Transform the SemIR to an LLVM Module.
- CodeGen: Transform the LLVM Module into an Object File.
Design patterns
A few common design patterns are:
-
Distinct steps: Each step of processing produces an output structure, avoiding callbacks passing data between structures.
-
For example, the parser takes a
Lex::TokenizedBufferas input and produces aParse::Treeas output. -
Performance: It should yield better locality versus a callback approach.
-
Understandability: Each step has a clear input and output, versus callbacks which obscure the flow of data.
-
-
Vectorized storage: Data is stored in vectors and flyweights are passed around, avoiding more typical heap allocation with pointers.
-
For example, the parse tree is stored as a
llvm::SmallVector<Parse::Tree::NodeImpl>indexed byParse::Nodewhich wraps anint32_t. -
Performance: Vectorization both minimizes memory allocation overhead and enables better read caching because adjacent entries will be cached together.
-
-
Iterative processing: We rely on state stacks and iterative loops for parsing, avoiding recursive function calls.
-
For example, the parser has a
Parse::Stateenum tracked instate_stack_, and loops inParse::Tree::Parse. -
Scalability: Complex code must not cause recursion issues. We have experience in Clang seeing stack frame recursion limits being hit in unexpected ways, and non-recursive approaches largely avoid that risk.
-
See also Idioms for abbreviations and more implementation techniques.
Adding features
We have a walkthrough for adding features.