mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-10-05 18:21:16 +01:00
dfd5fe368d1c13ac6ef302ffe67f6f391ea05021
35
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8c3fa80691 |
Add cc rule wrappers for cc_env (#5277)
Rules executed by bazel don't necessarily have the right environment to find the symbolizer, which was the intent of `cc_env` setting `LLVM_SYMBOLIZER_PATH`. So far, this has kind of been a case-by-case fix, but every so often I'm trying to debug a crash in a test that doesn't provide it. Rather continuing down this route, instead add drop-in wrappers for cc rules so that it's hard to forget. Note `bazel/cc_rules` is intended to mirror `bazel/carbon_rules` and `bazel/cc_toolchains`, rather than `@rules_cc`. AFAICT there isn't a great way to add this as a default for the `bazel run` environment. It's not typically going to be set on its own, forwarding `$PATH` would be too broad, and the [action `env_sets`](https://bazel.build/docs/cc-toolchain-config-reference#using-action-config) I think are not quite what we need (I think those don't include output execution, only compilation). |
||
|
|
422df75a92 |
Switch tree-sitter from explorer to toolchain testdata (#5292)
Noticed as part of #5290; `srcs` is needed to make `$(locations)` work. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> |
||
|
|
3ae62f8130 |
Rewrite Dump calls to use std::string returns (#5195)
Thought this might be interesting for you to allow more continuous
stream use. Also eliminates the need for `DumpNoNewline`.
```
expr Dump(context, complete_type_id)
(std::string) $0 = "type(inst1553): <builtin i32>; {kind: IntType, arg0: signed, arg1: inst1508, type: type(TypeType)}"
complete_type_id.Dump()
(std::string) $1 = "type(inst1553)"
expr Dump(context, specific_id)
(std::string) $2 = "specific166: {generic: generic0, args: inst_block772}"
expr Dump(context, query_self_const_id)
(std::string) $3 = "concrete_constant(inst1510): {kind: ClassType, arg0: class0, arg1: specific166, type: type(TypeType)}"
expr Dump(context, MakeFacetTypeId(arg))
(std::string) $4 = "facet_type22: {impls interface: interface10}
- interface10: {name: name26, parent_scope: name_scope0} `BitAnd`
complete: complete_facet_type22
- interface10: {name: name26, parent_scope: name_scope0} `BitAnd` (to impl)"
```
---------
Co-authored-by: Dana Jansens <danakj@orodu.net>
|
||
|
|
c44e688e5d |
Add parsing for 'fn destroy' (#5045)
Syntax is proposed in #5017, but has already been discussed with leads. Semantics is left as a TODO. --------- Co-authored-by: josh11b <15258583+josh11b@users.noreply.github.com> |
||
|
|
90b6f5a22c |
Refactor NodeCategory for X-macros (#5029)
Taking an approach similar to NameId in #5018 |
||
|
|
7eee9a3489 |
Refactor resolving a location into a SemIR library (#4876)
At present, lower depends on `Check::SemIRDiagnosticConverter` for debug info. That was to support a quick implementation of debug info, but isn't great because it's both an unusual dependency on check's implementation, and relying on diagnostic structures for debug info. This cleans that up by splitting relevant logic out to a library in sem_ir, and having lowering use sem_ir's library instead of check's. Additionally, a small refactoring of `Parse::TreeAndSubtrees` to allow getting locations in lowering without going through a `DiagnosticLoc`. I'm adding `Parse::GetTreeAndSubtreesFn` in because it's a complex signature to have in so many spots. I chose to have `ResolveNodeId` return a `SmallVector` because it seemed likely to be fairly compact, but that could also be using an optional callback to handle resolved node IDs, possibly just returning the last entry. This could be switched if preferred. Note this change shouldn't affect behavior, it's just moving code around. --------- Co-authored-by: Chandler Carruth <chandlerc@gmail.com> |
||
|
|
133717cd7e |
Eliminate NodeLocConverter (#4870)
I'm looking at eliminating `DiagnosticConverter`. This change removes `NodeLocConverter` (albeit adding `UnitAndImportsDiagnosticConverter`), and in doing so, refactors lex conversion functions to extract them out from the `DiagnosticConverter` functions. I'll be following up with changes that collapse `DiagnosticConverter` logic into `DiagnosticEmitter` locations. The intent is that we shouldn't need separate ownership of both types. |
||
|
|
4c4c4a4d2c |
Add RawStringOstream for slightly simpler streaming to strings (#4817)
This adds a RawStringOstream. Versus TestRawOstream, which is
consolidated over to RawStringOstream, it uses a string for storage
instead of a vector, mainly to support move-to-string semantics. Versus
llvm::raw_string_ostream, it owns the string and supports pwrite (which
is needed for driver and its fd_ostream compatibility requirement).
This converts most uses of llvm::raw_string_ostream, leaving behind a
few in InstNamer that explicitly cannot own the string, such as:
```
llvm::raw_string_ostream(name)
<< "_" << tree.tokens().GetColumnNumber(token);
```
I have this as its own library so that it can use CHECK.
Yes this doesn't save much code, but it's code we repeatedly write.
---------
Co-authored-by: Geoff Romer <gromer@google.com>
|
||
|
|
8f685b6953 |
Change how diagnostics are ordered (#4778)
This change deliberately breaks away from the line/column ordering, and instead focuses on a last byte offset corresponding to the final token processed as part of producing the message. Where that's equal, this maintains stable ordering in order to reflect the order that diagnostics were produced. The intent of this approach is that lex, parse, and check diagnostics are interleaved based on where they are produced, but that subexpressions still have diagnostics emitted prior to containing expressions. In particular, the prior line/column sort essentially sorted on the _start_ of where a diagnostic was associated, and this is closer to sorting based on the _end_. As a consequence, something like `F(1 2)` will have the error for `1 2` emitted _before_ a diagnostic for `F(1 2)` not matching parameters, instead of _after_. In check, we track the last handled node. This provides a last_byte_offset _separate_ from where a diagnostic is associated. The intent is that this creates an ordering of diagnostics which may be associated with earlier code, to cause the diagnostics to be emitted later. An example consequence of this is the change in ordering of modifier diagnostics: we are diagnosing those from the same place, but they have the same last_byte_offset, so we print them out in the order produced. I've added similar tracking to parse, but cannot identify any test which is affected by it (note the separate commit, I thought about this late). I'm not sure whether we have good out-of-order errors we could produce for this. A significant number of tests have reordered diagnostics as a consequence of this change, so this change does not add further testing. |
||
|
|
3ce0df67bb |
Add Dump functions to Check, Parse, and Lex (#4669)
- Provide `Check::Dump(context, arg)` and similar. - gdb and lldb should do contextual lookup, and `call Dump(*this, Lex::TokenIndex::Invalid)` has been tested with gdb. - Since this is only for debug, keeps the functions fully separated from code. - Uses alwayslink to ensure objects are correctly linked, even though there are no calls. - `-Wno-missing-prototypes` is needed when we don't have forward declarations. - Code is not linked in opt builds, using `#ifndef NDEBUG`. - This probably could be doing something in BUILD files with a `select()`, but the `#ifndef` seemed easier. This is based on #4620, but uses free functions instead of member functions. Co-authored-by: Dana Jansens <danakj@orodu.net> --------- Co-authored-by: danakj <danakj@orodu.net> |
||
|
|
4148161e24 |
Refactor value store code to use separate files. (#4477)
This is in anticipation of making the integer value store be customized heavily. I'd like to extract it from the common code when doing that, so first disentangling them here without any intended change in functionality or behavior to enable that. I've tried to update `#include`s to be as minimal as I can and added a few missing includes spotted in the process. I've split the test for value store to include what was easy focused on just the value store templates rather than the unified shared value stores. This might surface some opportunities for adding more tests, but for this PR, just doing the minimal restructuring. |
||
|
|
e58ce3e1bb |
Add coverage testing for parse node kinds. (#4436)
This refactors the diagnostic kind coverage check into something that also works for node kinds. Then, since this points out a few node kinds that aren't having their parse verified, I'm adding minor tests for those. |
||
|
|
96964ee534 |
Implement basic bool and int formatting for diagnostics (#4411)
Note, this supports plurals, but doesn't apply it anywhere. I'm mainly doing that to demonstrate the approach regarding syntax. See format_providers.h for details. |
||
|
|
0f350255ce |
Refactor compile-related tests to share construction. (#4396)
Note in particular that this fixes an issue where SharedValueStore had been shared across files, when they should be per-file. This is only visible when doing multiple compilations in a single test, which was rare before. This also moves these tests into the Testing namespace. My memory of the various namespacing changes is that we'd generally agreed to have tests in Testing so that we'd see SemIR:: and similar, same as we would in a lot of the implementation. |
||
|
|
f67791cfee |
Separate subtree size information from parse nodes. (#4174)
Move subtree sizes over to TreeAndSubtrees, using the different structure to represent the additional parse work that occurs, as well as making it clear which functions require the extra information. My intent is to make it hard to use this by accident. The subtree size is still tracked during Parse::Tree construction. I think a lot of that can be cleaned up, although we use it during placeholder assignment so it may take some work. I wanted to see what people thought about this before taking action on such a change. I'm using a 1m line source file generated by #4124 for testing. Command is `time bazel-bin/toolchain/install/prefix_root/bin/carbon compile --phase=check --dump-mem-usage ~/tmp/data.carbon` At head, what I'm seeing is: ``` ... parse_tree_.node_impls_: used_bytes: 61516116 reserved_bytes: 61516116 ... Total: used_bytes: 447814230 reserved_bytes: 551663894 ... 1.43s user 0.14s system 99% cpu 1.565 total ``` With `Tree::Verify` disabled completely, it looks like: ``` parse_tree_.node_impls_: used_bytes: 41010744 reserved_bytes: 41010744 ... Total: used_bytes: 427308858 reserved_bytes: 531158522 ... 1.20s user 0.13s system 99% cpu 1.332 total ``` Re-enabling just the basic verification (what is now `Tree::Verify`), I'm seeing maybe 0.05s slower, but that's within noise for my system. I do see variability in my timing results, and overall I think this is a 0.2s +/- 0.1s improvement versus the earlier (always testing `Extract` code) implementation. That's opt; debug builds will be unaffected, because the same checking occurs as before. Note, the subtree size is a third of the node representation, which is why I'm showing the decrease in memory usage here. |
||
|
|
60db3df24b |
Remove glob_sh_run (#4041)
file_test's --dump_output flag has essentially supplanted this functionality, and is necessary for execution on multi-file tests. |
||
|
|
8c64f0bfdd |
Add -Wmissing-prototypes and fix issues it finds. (#4019)
Most of these are places where we failed to include a header file and simply never got an error about this. The fix is to include the header file. Most other cases are functions that should have been marked `static` but were not. Finding all of these was a main motivation for me enabling the warning despite how much work it is. One complicating factor was that we weren't including the `handle.h` for all the state-based handler functions. While this isn't a tiny amount of code, it is just declarations and doesn't add any extra dependencies. It also lets us have the checking for which functions need to be `static` and which don't. For the `parse` library I had to add the `handle.h` header as well, I tried to match the design of it in `check`. I have also had to work around a bug in the warning, but given the value it seems to be providing, that seems reasonable. I've filed the bug upstream: https://github.com/llvm/llvm-project/issues/94138 I also had to use some hacks to work around limitations of Bazel rules that wrap `cc_library` rules and don't expose `copts`. I filed a bug for `cc_proto_library` specifically: ~https://github.com/bazelbuild/bazel/issues/22610~ https://github.com/bazelbuild/bazel/issues/4446 |
||
|
|
3c01ee69ed |
Move information on the token associated with a parse node from the .def file into the typed node. (#4001)
Instead of tracking the token associated with a parse node in the `.def` file macro, track it on the typed node instead. List the token as a field inside the node structure to show the order of the token relative to the other components of the grammar production, and to allow the token index to be accessed when the node is extracted. Remove the corresponding information from the `.def` file, leaving behind just a list of parse node kinds in the majority of cases. This also removes the checking of the token kind associated with a parse node in the case where the parse node has errors. Previously we had a flag on the node kind to indicate whether we should check this, but per [discord discussion](https://discord.com/channels/655572317891461132/655578254970716160/1246214418979881052), we have decided to remove this. --------- Co-authored-by: Jon Ross-Perkins <jperkins@google.com> |
||
|
|
cda5f66d22 |
Refactor NodeCategory to provide a class API (#4004)
Mirroring #4003 for NodeCategory. Note we template a lot more on NodeCategory's enum, so this is a slightly more awkward delta. Also, switch from Enum in KeywordModifierSet to RawEnumType for consistency with EnumBase. The templating on NodeCategory had me thinking about that more. |
||
|
|
0bd45f0d6b |
Rename DiagnosticLocationTranslator -> DiagnosticConverter (#3804)
Since the addition of TranslateArg, I don't think this type is going to go away (cutting a TODO). Refactoring names slightly to fit the current role, and adding const to ConvertLocation. |
||
|
|
a3b1c433be |
Remove legacy repo_name settings (#3772)
I'd kept these in to separate the bazel module update from the BUILD file changes, then forgot about it. I think all of these can be cleanly removed now. I think it's something we should clean up for consistency with the bazel central repository names; I think it's best to reduce that divergence. llvm_zlib and llvm_zstd remain because of how llvm depends on the particular names. |
||
|
|
bf02d1f4b0 |
Remove headers marked as unused by ClangD. (#3661)
This required adding a few headers that were found transitively before, but not too many. This is sadly a fairly manual process of opening every file in my IDE, but I think I got everything in `//common` and `//toolchain`. There are a few cases where technically we don't need `foo.h` to be included into `foo.cpp`, but I've forced those to stay with a pragma. I've tried to catch the places where we can cut deps in Bazel as well, but not sure I got all of those. I had been noticing these in other PRs and it seemed better to isolate the change. |
||
|
|
95746df788 |
Split out context targets for check and parse. (#3558)
Per request on #3556. Note, I'm assuming handlers should remain in the same target as the caller for longer-term performance reasons. |
||
|
|
c8b30d3eec |
Split Parse out to its own target. (#3556)
This is mirroring the structure of codegen/codegen.h, lower/lower.h, and check/check.h. I recently did lex/lex.h, so parse/parse.h is the last. Now, the directory's main API file is eponymous with the directory. I could've used a friend function to avoid making the Tree constructor public, but in other places we make less use of `friend`, just leaving things public. This felt more consistent, and simple because it only affects the constructor. |
||
|
|
2e97f27b8d |
Typed wrappers around parse tree nodes (#3534)
These are intended to allow the structure of a parse tree node to be
described more precisely in code, to support these use cases:
- Automated checking that the parse tree conforms to the expected
structure. (Added to `Tree::Verify`.)
- Easier reading and understanding of the structure of the parse tree by
toolchain developers. (See `parse/typed_nodes.h`.)
- Easier navigation of the parse tree, for example for tooling uses and
for use when forming diagnostics.
On this last point, an object representing the file may be inspecting
using `Tree::ExtractFile`, as in:
```
auto file = tree->ExtractFile();
for (AnyDeclId decl_id : file.decls) {
// `decl_id` is convertible to a `NodeId`.
if (std::optional<FunctionDecl> fn_decl =
tree->ExtractAs<FunctionDecl>(decl_id)) {
// fn_decl->params is a `TuplePatternId` (which extends `NodeId`)
// that is guaranteed to reference a `TuplePattern`.
std::optional<TuplePattern> params = tree->Extract(fn_decl->params);
// `params` has a value unless there was an error in that node.
} else if (auto class_def = tree->ExtractAs<ClassDefinition>(decl_id)) {
// ...
}
}
```
The `Extract...` functions collect the child nodes into the typed parse
node's fields (internally using a `Tree::SiblingIterator`) for easy
access. However, this is not as fast as directly observing the tree
structure using the postorder strategy being used by the check stage.
These functions rely on using struct reflection on the typed parse node
definitions from `parse/typed_nodes.h` to get the expected structure of
child nodes and then populate them.
Note that validating these in `Tree::Verify` adds significant cost to
it, and is currently included in the parsing stage. Without this change,
a 10 mloc test case of lex & parse takes 4.129 s ± 0.041 s. With this
change, it takes 5.768 s ± 0.036 s.
This builds upon and completes #3393.
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
---------
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
|
||
|
|
6419568142 |
Fully underline parse nodes in diagnostics. (#3442)
Another incremental change to diagnostic formatting. I simply recurse over all the tokens in the subtree of a parse node and construct a `DiagnosticLocation` that covers all of the tokens. I believe it's nicer for the user to be directed at the entire chunk of source where the error is occurring rather then just pointing at the bracketing/terminator tokens, but let me know if you all agree. |
||
|
|
d024403dc4 |
Refactor checking flow to allow for ordering based on import/package. (#3379)
As I was working on this, I noticed `import` and `library` syntax needs to be fixed for how it imports the current package, and for `Main` libraries. This mostly reflects the current state in its testing. Otherwise, this should handle most of the errors I could think of: dependency cycles, redundant imports, etc. It does not actually deal with the nuances of cross-IR references. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> |
||
|
|
cafcd88882 |
Split lexing logic and storage to separate files. (#3365)
Just reorganizing logic a little, trying to mirror the direction we've gone with check, lower, etc. That is, lex.h contains a function `Lex` that is used directly. Note, I'm avoiding making meaningful changes here. It could in theory still affect inlining in benchmarks, but I'm not seeing an impact. Before: ``` ------------------------------------------------------------------------------------------------------ Benchmark Time CPU Iterations UserCounters... ------------------------------------------------------------------------------------------------------ BM_ValidKeywords 2784949 ns 2784867 ns 249 bytes_per_second=214.452M/s tokens_per_second=35.9084M/s BM_ValidKeywordsAsRawIdentifiers 3222597 ns 3222551 ns 210 bytes_per_second=244.513M/s tokens_per_second=31.0313M/s BM_RawIdentifierFocus 5907836 ns 5907518 ns 103 bytes_per_second=264.873M/s tokens_per_second=16.9276M/s BM_ValidIdentifiers<1, 64, false> 6255128 ns 6254297 ns 105 bytes_per_second=235.488M/s tokens_per_second=15.989M/s BM_ValidIdentifiers<1, 1, true> |
||
|
|
d13f76e001 |
Add value store to be shared across compile stages. (#3311)
This updates lexing to use the data. I'll do checking separately, just to split changes. Note the ValueStore structure is also set up such that SemIR::File can use it for other fields. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> |
||
|
|
a112f2e802 |
Validate parse nodes correspond to expected tokens (#3295)
Can specify which tokens are allowed generally, and any additional tokens that only occur when the parse node has an error. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> |
||
|
|
32a1be3690 |
Detect invalid yaml. (#3239)
Previous code could fail silently, only caught with yaml output mismatches. This uses ErrorOr to be explicit about the error detection; recovery isn't supported because we should only print valid yaml (in tests, at least). Also moves the test helpers from //toolchain/base to //toolchain/testing. Switches tree_test to use the matchers so that, on mismatch, gtest prints a path to the mismatch (the straight value means it was just printing "not equal"). Addresses https://github.com/carbon-language/carbon-lang/pull/3217#discussion_r1323586836 |
||
|
|
8aa3d960f5 |
Merge toolchain file_test children in order to improve linking. (#3206)
Specifically this should improve linking by producing one large binary instead of one-per-directory. The inclusion of the driver hits the size issue. Separating out things which have more llvm deps has been discussed, but I'm not doing that here because I think the semantics layer will need to depend on clang for interop, and we'd lose a lot of the benefits that way. Also, having just one place to look seems simpler. Includes supporting changes to file_test infrastructure, the most significant of which is probably passing tests via file instead of a really large args thing, using a custom rule to do that. That's because dealing with the layered filegroups that allow the toolchain setup is more complicated, and this approach scales well. Combined test time is ~9s, so not sharding right now. I wasn't sure if people would prefer having the autoupdate script under testing, so I left it alone for now. |
||
|
|
de6e00c351 |
Fix toolchain glob_sh_run uses. (#3181)
These were missed when changing the command line. Also brings -v output to parity with formatted IR printing, and removes a redundant flush. |
||
|
|
ec182fb00d |
Rename lexer dir to lex (#3179)
Continuing with #3070. Just a dir and file rename (only prefix change is lexer_file_test). Everything in the lex dir should be marked as a move. Note, I think this closes #3070. There may still be further cleanup later, but the organizational changes suggested there are being completed. --------- Co-authored-by: Chandler Carruth <chandlerc@gmail.com> |
||
|
|
c555b39a2c |
Rename parser dir to parse (#3178)
Continuing with #3070. Just a dir and file rename (mostly removing prefixes, although for parse_tree_fuzzer and parse_tree_file_test I'm dropping "tree" instead of "parse"). Everything in the parse dir should be marked as a move. |