Now that we've centralized on a monorepo, tidy up some lingering
references.
For the `pr_comments.py` script, I've left the flag in place to select
a repository as it seems useful functionality to have available even if
we don't expect to need it in the near term. That said, I can pull it
out if folks prefer.
I haven't updated any of the *proposals* because it seems better to
leave those as-is from when they were written. The references seem
unlikely to be confusing to me. Happy for suggestions if needed there
though.
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
This is often a (much) better option for this code as we have many types
that are just an integer, pointer, or enum of data that would be much
more efficient as a by-value parameter. It is also a cleaner design in
most cases.
One place can't migrate in this way: iterators that use the LLVM CRTP
library for filling out iterator methods. The CRTP dispatch expects
operators to be actual members, and so those are left as-is in this
patch.
This patch tries to make everything adhere to the Google declaration
order. This is tricky as there doesn't appear to be a `clang-tidy` check
that helps us here at all.
One of the complex cases are the `enum`-wrapping classes we use. These
need to have a *public* conversion operator to the nested `enum` type to support
`switch` statements and the like. However, this makes it impossible to
fully respect the Google declaration order. We need to define the enum
type before the public API in order to use it in contexts like the
conversion operator. Even defining out-of-line won't help avoid this. It
is weird to have a type in the public API that is private, but again
this is only intended to be used for implicit conversions within
a `switch` statement or a `case` label. In one case, we had the
non-conforming order of declaration. I've added a comment to explain why
there. In the other case, the conversion operator is actually *private*
rather than public, which doesn't actually work in practice. I've made
this public and moved the declaration order to match with a matching
comment.
Most of these were fixed automatically (including things like adding
`[[nodiscard]]` and such). A number of others required manual edits.
I think all of them were pretty nice improvements.
There were a few places where the issues really stem from external
constraints and I've disabled the checks: GoogleTest macros or the
specific LibFuzzer entry points.
The only other places I disabled are the implicit conversions to
a private `enum` in the classes wrapping those `enum`s. These implicit
conversions are necessarily implicit to serve their only purpose:
enabling their use in `switch` statements and `case` labels. When these
were highlighted, it showed that one of these was actually converting to
an *`int`*. I've switched that to use the private `enum` instead as
doing so is important to enable warnings on non-covering `switch`
statements over than `enum`. And indeed, there is a `switch` that was
was implicitly relying on falling through in this way, so I've added the
explicit documentation of the intentional pattern to address that
warning.
Sorry this is so large, all of this somewhat fell out of enabling the
`clang-tidy` checks. If it is too difficult to review as lump, I can
work on breaking it apart as needed. Just let me know.
The term "constant" in this setting seems to actually mean anything that
is `const` qualified. For Carbon code, anything that should actually use
`CamelCase` should also use `constexpr` which has its own naming setting
and is correct. With this setting change and making a test variable
`constexpr`, clang-tidy is happy with all the names.
The rationale and rule for this was added in #194 to the C++ style guide
we are using for Carbon.
I've applied the automated fixes from running `clang-tidy` over all the
code, and then run `clang-format` afterward.
There are a few places where `clang-format` fixed a formatting issue
that snuck through in prior commits. These were rare enough that it
didn't seem worth splitting them out into a separate change.
This was generated using the same steps as outlined for the lexer:
```
mkdir /tmp/new_corpus
./bazel-bin/parser/parse_tree_fuzzer -jobs=N /tmp/new_corpus
```
After some time, I stopped the fuzzer and then minimized and merged the
corpus:
```
rm parser/fuzzer_corpus/empty
./bazel-bin/parser/parse_tree_fuzzer -merge=1 parser/fuzzer_corpus /tmp/new_corpus
```
I'll add this short version to documentation in another PR.
Only change is to update the path to the fuzzer build extension.
Original main commit message:
> Add an initial parser library. (#30)
>
> This library builds a parse tree, very similar to a concrete syntax
> tree. There are no semantics here, simply introducing the basic
> syntactic structure.
>
> The current focus has been on the APIs and the data structures used to
> represent the parse tree, and not on the actual code doing the
> parsing. The code doing the parsing tries to be reasonably efficient
> and reasonably easy to understand recursive descent parser. But there
> is likely much that can be done to improve this code path. A notable
> area where very little thought has been given yet are emitting good
> diagnostics and doing good recovery in the event of parse errors.
>
> Also, this code does not try to match the current under-discussion
> grammar closely. It is only partial and reflects discussions from some
> time ago. It should be updated incrementally to reflect the current
> expected grammar.
>
> The data structure used for the parse tree is unusual. The first
> constraint is that there is a precise one-to-one correspondence
> between the tokens produced by the lexer and the nodes in the parse
> tree. Every token results in exactly one node. In that way, the parse
> tree can be thought of as merely shaping the token stream into a tree.
>
> Each node is also represented with a fixed set of data that is densely
> packed. Combined with the exact relationship to tokens, this allows us
> to fully allocate the parse tree's storage, and to use a dense array
> rather than a pointer-based tree structure.
>
> The tree structure itself is implicitly defined by tracking the size
> of each subtree rooted at a particular node. See the code comments for
> more details (and I'm happy to add more comments where necessary). The
> goal is to minimize both the allocations (one), the working set size
> of the tree as a whole, and optimize common iteration patterns. The
> tree is stored in postorder. This allows depth-first postorder
> iteration as well as topological iteration by walking in reverse.
>
> Building the parse tree in postorder is a natural consequence of the
> grammar being LR rather than LL, which is a consequence of supporting
> infix operators.
>
> As with the Lexer, the parser supports an API for operating on the
> parse tree, as well as the ability to print the tree in both
> a human-readable and machine-readable format (YAML-based). It includes
> significant unit tests and a fuzz tester. The fuzzer's corpus will be
> in a follow-up commit.
>
> This is the largest chunk of code already written by several of us
> prior to open sourcing. (There are a few more pieces, but they are
> significantly smaller and less interesting.) If there are major things
> that folks would like to see happen here, it may make sense to move
> them into issues for tracking. I have tried to update the code to
> follow the style guidelines, but apologies if I missed anything, just
> let me know. We also have issues #19 and #29 to track things that
> already came up with the lexer.
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
The only change here is to update the fuzzer build extension path.
The main original commit message:
> Add an initial lexer. (#17)
>
> The specific logic here hasn't been updated to track the latest
> discussed changes, much less implement many aspects of things like
> Unicode support.
>
> However, this should lay out a reasonable framework and set of APIs.
> It gives an idea of the overall lexer architecture being proposed. The
> actual lexing algorithm is a relatively boring and naive hand written
> loop. It may make sense to replace this with something generated or
> other more advanced approach in the future, getting the implementation
> right was not the primary goal here. Instead, the focus was entirely
> on the architecture, encapsulation, APIs, and the testing
> infrastructure.
>
> The architecture of the lexer differs from "classical" high
> performance lexers in compilers. A high level summary:
>
> - It is eager rather than lazy, lexing an entire file.
> - Tokens intrinsically know their source location.
> - Grouping lexical symbols are tracked within the lexer.
> - Indentation is tracked within the lexer.
>
> Tracking of grouping and indentation is intended to simplify the
> strategies used for recovery of mismatched grouping tokens, and
> eventually use indentation.
>
> Folding source location into the token itself simplifies the data
> structures significantly, and doesn't lose any fidelity due to the
> absence of a preprocessor with token pasting.
>
> The fact that this is an eager lexer instead of a lazy lexer is
> designed to simplify the implementation and testing of the lexer (and
> subsequent components). There is no reason to expect Carbon to lex so
> many tokens that there are significant locality advantages of lazy
> lexing. Moreover, if we want comparable performance benefits, I think
> pipelining is a much more promising architecture than laziness. For
> now, the simplicity is a huge win.
>
> Being eager also makes it easy for us to use extremely dense memory
> encodings for the information about lexed tokens. Everything is
> created in a dense array, and small indices are used to identify each
> token within the array.
>
> There is a fuzzer included here that we have run extensively over the
> code, but currently toolchain bugs and Bazel limitations prevent it
> from easily building. I'm hoping myself or someone else can push on
> this soon and enable the fuzzer to at least build if not run fuzz
> tests automatically. We have a significant fuzzing corpus that I'll
> add in a subsequent commit as well.
This also includes the fuzzer whose commit message was:
> Add fuzz testing infrastructure and the lexer's fuzzer. (#21)
>
> This adds a fairly simple `cc_fuzz_test` macro that is specialized for
> working with LLVM's LibFuzzer. In addition to building the fuzzer
> binary with the toolchain's `fuzzer` feature, it also sets up the test
> execution to pass the corpus as file arguments which is a simple
> mechanism to enable regression testing against the fuzz corpus.
>
> I've included an initial fuzzer corpus as well. To run the fuzzer in
> an open ended fashion, and build up a larger corpus:
> ```shell
> mkdir /tmp/new_corpus
> cp lexer/fuzzer_corpus/* /tmp/new_corpus
> ./bazel-bin/lexer/tokenized_buffer_fuzzer /tmp/new_corpus
> ```
>
> You can parallelize the fuzzer by adding `-jobs=N` for N threads. For
> more details about running fuzzers, see the documentation:
> http://llvm.org/docs/LibFuzzer.html
>
> To minimize and merge any interesting new inputs:
> ```shell
> ./bazel-bin/lexer/tokenized_buffer_fuzzer -merge=1 \
> lexer/fuzzer_corpus /tmp/new_corpus
> ```
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
Main original commit message:
> Add stubs of a diagnostic emission library. (#13)
>
> This doesn't have much wired up yet, but tries to lay out the most
> primitive API pattern.
We've already started the merge, but this tries to lay out a reasonable
trajectory for things to move along.
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
This library manages buffers of source code, either in-memory or mapped
from the filesystem. At the moment it is a bit simplistic and assumes
`mmap` is available in its implementation details. We should make this
more portable in the future, but this is currently just a direct copy
from the toolchain repository.
The Bazel bits are collected into a directory and given less confusing
names (I hope). Other than names, everything is a direct copy from the
toolchain repository without any edits.
A few files needed to be merged in:
- `.gitignore`
- `.pre-commit-config.yaml`
Subsequent commits will add relevant C++ infrastructure and then the
source code itself.
The jekyll build keeps top-level without expanding them. So the site works, but the publish copy failed because some symlinks to files that aren't visible through the site (like CONTRIBUTING.md, versus the site's CONTRIBUTING.html) didn't work.
One implication here is that both proposals/README.md and the website should use basically the same format for their proposal lists. Another subtle change is that the website sidebar was previously sorted by *filename*, and is now sorted by *title*, which seems better because it's something that readers can see.
Enumerating why files change:
- gen_sidebar.py is the centerpiece, with a couple helpful functions for re-use elsewhere
- Delete the old sidebar html includes
- Modify the Makefile to handle gen_sidebar.py reasonably well (let's be honest... I'm not great at Makefiles)
- Move proposal listing out to its own file (proposals.py) for re-use
- Add a few PYTHONPATH things + `__init__.py` files to get modules importing correctly (may be a better way at this, I've hit a wall though).
This takes #175 and publishes it as interop goals. I've made a few small editorial changes, but the main addition is the "Offer equivalent support for languages other than C++" non-goal which I thought may be useful (in particular, it can be used to clarify why this directory is "interoperability" and not "interoperability_cpp").
This simply requires braces and doesn't allow single-line `if`s. There
is some minor readability loss here, but it seems minor and provides
extremely simple rules which I'd value.
I can also trivially get clang-tidy to both check for this and
automatically fix code to conform.
Just sending this as a code review as it seems fully in the direction of
the style guide approved by the core team and I've heard no real
objections. That said, if anyone is concerned, I'm happy to take it
through the proposal process.
* first draft of proposal for basic syntax
* rename proposal
* formating
* fixing typos
* precendence and associativity
* minor edit
* added abstract syntax
* some revisions based on feedback from meeting today
* optional return type
* oops, not optional for function declarations
* replacing abbreviations with full names
* name changes
* Update proposals/p0162.md
Co-authored-by: Dmitri Gribenko <gribozavr@gmail.com>
* removed = from precedence table
* for fun decl, back to optional return type, shorthand for void return
* more fiddling with return types
* change base case of statement_list to empty
* added pattern non-terminal
* added expression style function definitions
* change && to and, || to or
* change and to have equal precedence
* updating text to match grammar, fix typo in grammar regarding pattern
* adding trailing comma thing to tuples
* removed 'alt' keyword, not necessary
* flipped expression and pattern
* added period for named arguments and documented the reason
* change alternative syntax to use tuple instead of expression
* comment about abstract syntax
* code block language annotations
* added alternative designs, some cleanup for pre-commit
* spell checking
* filled out the TOC
* minor edits
* minor edit
* trying to fix pre-commit error
* changes from pre-commit?
* changes based on meeting today
* describe alternatives regarding methods
* more rationale in discussion of alternatives
* addressing comments
* typo fix, added text about next steps
* removed *, changed a ! to not
* Update proposals/p0162.md
Co-authored-by: Geoff Romer <gromer@google.com>
* edits from feedback
* pre-commit
* added second reason for period in field initializer
* added executable semantics, fixed misunderstanding regarding associativity
* update README
* added a paragraph about tuples and tuple types
* Update proposals/p0162.md
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
* addressing comments from josh11b
* Update executable-semantics/README.md
Co-authored-by: Dmitri Gribenko <gribozavr@gmail.com>
* link to implicit parameters in generics proposal
* added non-terminal for designator as per geoffromer's suggestion
* changed handling of tuples in function call and similar places as per josh11b and zygoloid
* moving code to separate PR
Co-authored-by: Dmitri Gribenko <gribozavr@gmail.com>
Co-authored-by: Geoff Romer <gromer@google.com>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Adopts new copyright and markdown toc checks.
The new toc check generates the toc header, so that's why all the md files changed (this had felt better to me for the long-term, more auto-generated content)
`.tmp` -> `.tmp.md` causes the right copyright to be chosen (new issue due to the check-copyright pre-commit)
The `_get_proposals_dir` changes are to allow pytest to be run from any directory.
When writing C++ code for Carbon, we want to keep all of our code consistent,
easy to learn, and help avoid spending undue code review time arguing about
the same core style and idiomatic issues.
This adopts the Google C++ style guide as a baseline, and then makes minimal,
focused additions and adjustments to it to suit the needs of Carbon.
Co-authored-by: Thomas Köppe <tkoeppe@google.com>
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
Co-authored-by: Geoff Romer <gromer@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: austern <austern@google.com>
Co-authored-by: Dmitri Gribenko <gribozavr@gmail.com>
Factored out of the Lexical Conventions proposal (https://github.com/carbon-language/carbon-lang/pull/173). This proposal covers only the encoding aspect: Carbon text is Unicode, source files are encoded in UTF-8, and we will follow the Unicode Consortium's recommendations to the extent that they make sense to us.
There is limited direct or obvious support for working with a stack of
dependent pull requests with clean code review of each incremental
change.
Carbon needs to support high-latency asynchronous code review due to
timezones, schedule differences (especially in an open source project),
and the delays imposed by our proposal process. To this end we need some
solution for doing stacked pull requests with a reasonable code review
experience. This suggests a compromise flow that seems to minimize the
costs of doing this.
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Co-authored-by: Dmitri Gribenko <gribozavr@gmail.com>
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
Co-authored-by: austern <austern@google.com>
The pull request workflow was supposed to connect to the review guidance
anyways, but the link hadn't been added.
Also adds a TOC to the pull request workflow document as it was missing
one.