Commit Graph
11 Commits
Author SHA1 Message Date
3ba4997855 Canonicalize away bit width and embed small integers into IntIds (#4487)
The first change here is to canonicalize away bit width when tracking
integers in our shared value store. This lets us have a more definitive
model of "what is the mathematical value". It also frees us to use more
efficient bit widths when available, such as bits inside the ID itself.

For canonicalizing, we try to minimize the width adjustments and
maximize the use of the SSO in APInt, and so we never shrink belowe
64-bits and grow in multiples of the word bit width in the
implementation. We also canonicalize to the signed 2s compliment
representation so we can represent negative numbers in an intuitive way.

The canonicalizing requires getting the bit width out of the type and
adjusting to it within the toolchain when doing any kind of math, and
this PR updates various places to do that, as well as adding some
convenience APIs to assist.

Then we take advantage of the canonical form and embed small integers
into the ID itself rather than allocating storage for them and
referencing them with an index. This is especially helpful for the
pervasive small integers such as the sizes of types, arrays, etc. Those
no longer require indirection at all. Various short-cut APIs to take
advantage of this have also been added.

This PR improves lexing by about 5% when there are lots of `i32` types.

---------

Co-authored-by: Dana Jansens <danakj@orodu.net>
Co-authored-by: Carbon Infra Bot <carbon-external-infra@google.com>
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-11-13 09:36:20 +00:00
Chandler Carruth e71e6ca07f Use separate value stores for identifiers and string literals (#4106)
This undoes a previous change to unify them, and I think at my advice.
=[ Sorry about that, I think I was just wrong.

Specifically, I think I had suggested that it would be more efficient to
have a single shared hashtable of strings. The more I look at profiles
of the toolchain, the less likely that seems. Specifically for
identifiers and string literals it seems especially problematic.

Using a single, joint hashtable is likely a good idea when all of the
different querying code paths are equally likely, the strings follow the
same distribution of sizes, and either there is no clustering of access
to different sets of strings or none of the sets are meaningfully small
enough to fit into a lower level of resident cache.

I think essentially none of these predicates actually hold for
identifiers vs. string literals:
- Identifiers are *much* more hot
- They have wildly different size distributions.
- The access patterns are very clustered

Sorry for the misleading advice on that one.

While splitting them, I've worked to simplify the code a bit by building
a way to have the `StringRef` holding canonical value stores not require
specializations, and so we get a pretty large code cleanup in the
process here.
2024-07-03 15:54:04 +00:00
Jon Ross-Perkins 83413479d7 Move some of the test information to TIP lines (#4007)
This has is a nice-to-have for me. Frequently I want to run a specific
test, and end up digging through output to be able to copy-paste the run
line. This uses TIP lines to inject the command into the file when using
AUTOUPDATE.

Note, one of the reasons I want this is because "bazel test
//toolchain/testing:file_test --test_output=all" has been regularly
exceeding bazel's output limit for me (workaround is either opening the
output file or specifying an obscure output limit flag), making it a
little harder for me to get the commands. However, frequently I'm adding
a file and want to iterate on it, so that's really the use case I have
in mind here.
2024-06-04 17:57:23 +00:00
Richard Smith 5c8fa6ad5c Replace FoldingSet with DenseMap for instruction canonicalization. (#3979)
Switch from recursing into non-canonical instruction fields to
separately canonicalizing those fields. This means we now form canonical
`InstBlockId`s, `TypeBlockId`s, `IntId`s, `FloatId`s, and `BindNameId`s
at least in the cases when they're referenced by a constant instruction.

This reduces the overall runtime for @chandlerc's 10MLoC example by
27.5% on my machine.
2024-05-23 00:48:49 +00:00
Richard SmithandJon Ross-Perkins f5e386a12b Make driver find all prelude files. Add build rule for //examples:sieve. (#3895)
The driver now looks for all files under core/prelude/ and considers
them all to be part of the prelude. The driver also now only processes
the prelude in `--phase=check` and later, when it would actually be
imported.

With that done, add a simple `carbon_binary` build rule and use it to
build the example in `//examples`. This should cause the example to be
built as part of our continuous integration.

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2024-04-19 23:32:00 +00:00
Chandler Carruth e8430dda81 Move the injection of the prelude compile to the driver. (#3877)
The `file_test` layer still adds it to the VFS and filters it from the
dump output, but the driver now injects it into the compilation unites
depending on the `--import-prelude` file (enabled by default).

This doesn't address improvements to how we find the prelude code, just
sinking the needed logic into the driver itself.
2024-04-10 09:40:53 +00:00
Jon Ross-Perkins 0db63ff17a Abbreviate Integer and FloatingPoint (#3435)
I was suggesting this because `FloatingPoint` is pretty long. `int` and
`float` should be familiar abbreviations. `unsigned` should be familiar
to developers too, but `UnsignedInt` still feels usefully clearer for
the additional chars.
2023-11-29 23:29:48 +00:00
Jon Ross-Perkins d096655cc6 Split out the SharedValueStores to be per-compilation unit. (#3353)
Advantages:

- Allows lexing/parsing in parallel, since they are modifying fully
separate ValueStores.
- Allows SemIR to reliably be stored hermetically.

Disadvantages:

- Creates overlapping storage of duplicate strings when multiple files
are compiled together.
- Prevents Ids from being uniquely compared cross-file.

Per discussion, the decision is that the advantages are more important.

The looser ownership remains because both SemIR checking and
metaprogramming may still generate things we would want to deduplicate.
It would be somewhat odd if TokenizedBuffer owned something that
checking modified.
2023-11-01 19:44:17 +00:00
Jon Ross-Perkins 3af7eb2672 Refactor YAML handling to use the llvm::yaml API. (#3337)
Provides an adapter for the llvm::yaml API because it otherwise needs a
bunch of const/non-const definitions, and the traits are difficult to
diagnose issues with. The current approach is pretty simple to use, even
if it's not super efficient (which, yaml output is more of a debugging
thing so I'm not really expecting it to be an issue).

Changes the format of yaml output to provide more index information,
just as reminders when seeing something like `node+0`. Note this would
create more churn in deltas if we were reliant on the output yaml in
tests, but we aren't so it should be okay.
2023-10-26 18:50:30 +00:00
Jon Ross-Perkins 74c3c665fa Refactor SemIR YAML printing to use dashed lists. (#3330)
This standardizes on having ValueStore and related structures provide
printing, removing the handlers in file.cpp.

The `[]` is provided for empty sequences versus if there was simply
nothing, in which case it would be a sequence when non-empty, and a null
value when empty. Consistently (and explicitly) providing sequences
feels easier to understand.

The changes to the output yaml are overall more terse. My hope is that
this is an improvement for most readers.

Also fixes printing of APInt, defaulting to unsigned for consistency
with Carbon's use.
2023-10-25 16:32:24 +00:00
Jon Ross-Perkins 843dd40f22 Adding ValueStore printing and --dump-shared-values (#3320)
This replaces the printing that was removed from SemIR's raw dump. It's
separate because (for example) lexing generates shared values, and so
reviewing them is not specific to any particular phase.
2023-10-20 21:31:50 +00:00