Commit Graph
35 Commits
Author SHA1 Message Date
Jon Ross-Perkins 4e1b585fcf clang-tidy --fix (#2577)
Only automatic fixes.
2023-02-02 17:43:10 -08:00
Jon Ross-Perkins 94872ef6da Change TokenKind's Print overload to a format_provider. (#2534)
Fundamentally this `.Print()` is wrong for debug output at present because `.fixed_spelling()` can be empty. It's also inconsistent with other enums to use it. We frequently print tokens for debugging, and it's easy to forget to specify `.name()` there.

Diagnostics use formatv, so we can provide a format_provider and address it in one spot that way. It also makes it harder to just forget to do the right thing.
2023-01-18 12:20:56 -08:00
Jon Ross-PerkinsandChandler Carruth 78ac6cb7d1 Switch TokenKind to EnumBase (#2509)
This shouldn't have any behavior change, it's just using #2504

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2023-01-05 14:11:30 -08:00
Jon Ross-Perkins d42d864e82 Make TokenKind's API closer to toolchain's typical API setup. (#2456)
This is somewhat based on the name vs Name difference, but I figured I'd split it out and just sweep up the API on the whole while looking at a different approach to #2453
2022-12-08 15:31:04 -08:00
Jon Ross-Perkins 04f0288cd2 Bracket the tokenized buffer output. (#2446)
This makes TokenizedBuffer more consistent with ParseTree and SemanticsIR, which also wrap with [] to produce a sequence value.

It also makes it possible in driver.cpp to just prefix the line with a name, so it ends up with:

var_name: [
  (content)
]
Noticed this due to bracketing comments on #2443 and trying to think of better answers. With this, we can also say that the [] bracket a variable.
2022-12-07 11:22:33 -08:00
Chandler CarruthandJon Ross-Perkins 94cf343b05 Update LLVM and switch to std::optional. (#2424)
LLVM's bazel build has changed a bit, so this updates the tree for that.

LLVM is also moving `llvm::Optional` to match the standard API, but it seemed simpler to just switch to `std::optional`.

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2022-12-01 09:22:43 -08:00
Jon Ross-Perkins 8480a0bbd6 Fix a couple debug output issues (#2415)
Make `.run` targets include symbol information, and include token names on CHECK output in more places.
2022-11-30 10:24:09 -08:00
Jon Ross-Perkins 97ed697386 For flyweights, shift from llvm::SmallVector to a type that enforces index types. (#2398)
The intent here is to reduce use of vanilla `int32_t` without a clear indicator of what it's referencing, and to more tightly link references with the underlying types they reference into.

The type name DataIndex doesn't feel great, but I was kind of floundering for a better name.
2022-11-18 14:07:01 -08:00
Jon Ross-Perkins 4c8fdf5124 Start drafting out semantic type checking. (#2406)
This is a first pass at what semantic type checking might look like. Types propagate along nodes, we use an InvalidType object when there's an error, and once there's an InvalidType we stop doing so much type checking.

This adds some RealLiteral handling in order to get type mismatches. I'm cautious about creating some real value for SemanticsIR (since the tokenized buffer version is a bit constrained), so I'm not doing that yet. But I will probably need to in order to maintain SemanticsIR having hermetic copies of its data, without a parse tree dependency.
2022-11-17 13:36:40 -08:00
Jon Ross-Perkins adac572430 Replace some reference members with pointer members. (#2408)
I'd noted this while rewriting the parser and, while I kept it there during conversion, I think switching is consistent with the higher-level desire and the use of references therein was just an oversight.

Discussed at:
https://discord.com/channels/655572317891461132/655578254970716160/1042551242813083840
2022-11-17 08:58:43 -08:00
Jon Ross-Perkins 8e5dcc2588 Enable readability-qualified-auto (#2314)
As suggested on #2310
2022-10-18 19:21:49 -07:00
Kareem Ergawyandergawy 3d44169199 Fix integer literal token printing. (#2050)
Summary:
An `llvm::APInt` is always treated as a signed value by `operator<<`;
check [1]. This resulted in printing incorrect values for tokens that
have their MSB set to 1. For example, a value 9 would be printed as -7
since its `APInt` object would be 4-bits wide. However, integer literals
are always tokenized without the sign character so it is safe to treat
the values as unsigned for printing pruposes.

[1] https://llvm.org/doxygen/APInt_8h_source.html

Co-authored-by: ergawy <kareem.ergawy@guardsquare.com>
2022-08-17 09:29:07 -07:00
Jon Meow 2e776a442d FIXME -> TODO for C++ style (#1282) 2022-05-20 14:28:25 -07:00
Jon Meow af694b97cb Prefix most macro names with CARBON_ (#1232)
I'm doing this to avoid macro name conflicts, following https://google.github.io/styleguide/cppguide.html#Preprocessor_Macros: "If you do export a macro from a header, it must have a globally unique name. To achieve this, it must be named with a prefix consisting of your project's namespace name (but upper case)."

Commands run:

```
sed -i 's/\(DCHECK\|CHECK\|FATAL\|MAKE_UNIQUE_NAME\|MAKE_UNIQUE_NAME_IMPL\|RAW_EXITING_STREAM\|RETURN_IF_ERROR\|RETURN_IF_ERROR_IMPL\|ASSIGN_OR_RETURN\|ASSIGN_OR_RETURN_IMPL\|DIAGNOSTIC_KIND\|RETURN_IF_STACK_LIMITED\)(/CARBON_\1(/g' $(git ls-files *.cpp *.h *.lpp *.ypp *.def ':!third_party')
sed -i 's/#undef DIAGNOSTIC_KIND/#undef CARBON_DIAGNOSTIC_KIND/' toolchain/diagnostics/diagnostic_registry.def
```

Note this isn't *quite* everything, but it's intended to be a large pass at everything:

```
╚╡git grep '#define ' *.cpp *.h *.lpp *.ypp *.def ':!third_party' | grep -v '#define CARBON' | grep -v _H_
explorer/syntax/lexer.lpp:  #define YY_USER_ACTION                                             \
explorer/syntax/lexer.lpp:  #define SIMPLE_TOKEN(name) \
explorer/syntax/lexer.lpp:  #define ARG_TOKEN(name, arg) \
explorer/syntax/parse_and_lex_context.h:#define YY_DECL                                                         \
migrate_cpp/cpp_refactoring/var_decl.cpp:#define ABSTRACT_TYPE(Class, Base)
migrate_cpp/cpp_refactoring/var_decl.cpp:#define TYPE(Class, Base)     \
```

We may in particular want to do a pass to clean up #ifdef guards and make them be CARBON_ rooted.
2022-05-06 15:30:25 -07:00
Jon Meow a24963ea6b Move diagnostic definitions to be closer to where they emit (#1169) 2022-04-04 15:14:43 -07:00
Jon MeowandRichard Smith aaca540a05 Restructure Diagnostic objects to allow late formatting (#1131)
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2022-04-04 10:59:09 -07:00
Jon Meow f9014a6d10 clang-tidy with readability checks (#1148) 2022-03-24 13:22:44 -07:00
Jon Meow 7bbf4fbba3 Accessor renames on diagnostics+source buffer (#1133) 2022-03-15 15:43:53 -07:00
Jon Meow 16c6ba6bd1 Accessor renames on lexer (#1134) 2022-03-15 13:02:44 -07:00
Jon Meow c6c050ccfc Switch radix to an enum for easier formatting (#1130)
Note the use of an enum allows for the operator to be defined for formatting. This reduces the custom formatting required, which is a simplification carrying over to error handling changes I'm working on too.

Also, I kind of feel the code's now clearer about what's supported; I'm dropping CHECKs that seemed redundant with the enum.
2022-03-14 13:40:09 -07:00
Jon Meow f5f02babdd Apply a digit limit for all getAsInteger calls (#1117)
In particular noticed the issue in type literal parsing, but given this has come up once before, trying to address it consistently.
2022-03-03 12:26:04 -08:00
Jon Meow c4c5cdc1f5 Replace is_sorted with comparison (#1104)
* Replace is_sorted with comparison

* Switch to StringRefContainsPointer
2022-03-03 10:21:38 -08:00
Jon MeowandRichard Smith 0a8c0dc271 Adjust string parsing to consume everything until the terminator. (#1111)
Note I've added a few TODOs, particularly that multi-line strings should only consume until the dedent.

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2022-03-03 10:21:16 -08:00
Jon Meow f4f9b23291 Add int and real printing (#1116) 2022-03-03 09:57:41 -08:00
Jon Meow eecf4a34b3 Replace assert with CHECK (#1098) 2022-02-25 08:29:39 -08:00
Jon MeowandRichard Smith 22721da92b Avoid unnecessary relexing of the last line (#999)
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2022-01-25 10:20:42 -08:00
Jon Meow cab5bb0158 Clean up SimpleDiagnostic inheritance (#1004) 2022-01-07 11:01:01 -08:00
Jon Meow fd489b69ad Lexer style cleanup (#1002) 2022-01-05 10:30:19 -08:00
Jon MeowandChandler Carruth a562f872e7 Switch from assert to CHECK (#975)
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2021-12-08 08:38:27 -08:00
Jon MeowandChandler Carruth 652cd8c636 Style updates, mostly _ naming (#970)
There are some declaration order changes, and a few test classes switched from `struct` to `class`. However, this PR is mostly adopting `_` naming of private member variables due to the shift in naming style. None of what's here should have behavior impacts, it should just be style.

Note, there are a lot of things that *look* like they could be accessor-named, but I'm not doing that in this change. Happy to do it separately if you want me to do another PR focused on it.

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2021-12-07 09:46:44 -08:00
Chandler Carruth 5f67029479 Use upstream GoogleTest and add related test utils. (#876)
This moves over to the vanilla upstream GoogleTest pulled in the more
expected manner with Bazel. It also adds Abseil and Google Benchmark
libraries in the same fashion (there are cross dependencies here).

As part of this, also introduce a dependency check test that can enforce
basic layering of dependencies. For example, this lets us ensure that
non-test Carbon code only depends on LLVM and Clang despite having other
libraries available. There remains some cleanup to improve the way these
dependency tests work, but this at least ensures we don't regress.

I've also provided workarounds to allow both Carbon code and LLVM code
to freely be used with GoogleTest (and other `std::ostream` based
output code). This is done by extending the code in
`//common/ostream.h`. One downside is that it requires opening the
`llvm` namespace and adding an ADL_found overload there. I think on
balance this is still a win and doesn't make me too nervous.

The new version of GoogleTest requires printing more often from matchers
and so I've also added several printing routines to types that
previously didn't require them. Otherwise, most of the updates are just
using the more conventional upstream style of including the headers and
adding `ostream.h` where it is needed.

I did consider moving code over to use `std::ostream` instead of LLVM's
`raw_ostream`, but the advantages of not doing virtual dispatch still
seem significant, and it also seems good to retain access to LLVM's
formatting utilities built around `raw_ostream` given that we can't pull
arbitrary dependencies into Carbon code outside of test code.

All of this was slightly motivated by requests for newer features in
GoogleTest, but much more-so by my desire to have access to Google
Benchmark and Abseil when writing benchmarks. For example, using
Abseil's random number generator seems extremely helpful when generating
inputs for benchmarks. The growing dependencies between these packages
further motivated me to just pull them all in and ensure they worked
well.
2021-11-02 20:14:12 -07:00
Richard SmithandChandler Carruth a83c22288f [toolchain] Implement lexing and parsing support for #543. (#693)
Lex [iuf][1-9][0-9]* as a new kind of "sized type literal" token. When
parsing that token, form a literal expression.

Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
2021-08-02 15:37:43 -07:00
Chandler Carruth a857b7ea1a Cleanup or suppress numerous clang-tidy issues. (#577)
This gets us to a nearly clean state across the toolchain. A couple of
these are checks that I don't think we want to try to rigidly use and
I've disabled them completely. Others I've added relevant `NOLINT` style
suppressions or applied the automatic fix suggested by `clang-tidy`.

The implicit conversions that are allowed here with `NOLINT` are
probably worth at least a tiny bit of scrutiny to see if we could
replace the construct with something more direct without undue effort
and no longer need the implicit conversion. But until then, it seemed
fine to suppress.
2021-06-14 19:46:49 -07:00
Richard Smith 1b122924e8 [toolchain] Parse postfix operator * as a type operator. (#576) 2021-06-14 16:12:14 -07:00
Chandler Carruth 8f8ab23a77 Move the toolchain into a top-level directory. (#567)
This should clean up our top level directory and the build patterns.

No non-mechanical edits here. Just injecting `toolchain/` and
`TOOLCHAIN_` and then running formatting tools.
2021-06-08 03:01:37 -07:00