Now that LLVM 12 has been released we no longer have any need to
bootstrap LLVM to get the desired featureset. LLVM 12 is available
widely, including in Homebrew across multiple platforms and in the
GitHub action runners.
Sadly, the Linux distribution builds of LLVM-12 are largely broken and
not as useful for us. The Homebrew Linux install was also broken
originally, but I've worked extensively with the Homebrew folks to get
the Linux install into a really good shape. It should now work reliably.
There are two primary bugs in Linux LLVM packages that need to be fixed
before we can just use them:
- https://bugs.llvm.org/show_bug.cgi?id=43604
- https://bugs.llvm.org/show_bug.cgi?id=46321
Once those are addressed and point releases with the fixes widely
available we can further simplify things.
Even with the need to use Homebrew installs, using the released LLVM has
the extra advantage of making it easy to properly support Darwin ARM and
I've added that configuration so that I can test things there.
Last but not least, this will significantly shrink our build outputs
which should allow building much more in continuous integration on
GitHub actions without exceeding the action cache size limits. I've even
added several tweaks and adjustments to the compile and build flags to
improve the build performance and reduce the build output size.
Once this is landed and stable, we can consider adding the refactoring
tooling back to our CI.
One of the biggest downsides of this path is that our CI has to download
and install the LLVM toolchain from Homebrew on each run. This is pretty
slow (takes a couple of minutes). But it is a fixed overhead -- it won't
get worse over time. Eventually, we can either look at a much fancier
action configuration to avoid this or hopefully the Debian packages will
get updated and we can move back to those.
The bootstrapping has served us long enough at this point. We can
resurrect it if we ever find a compelling reason for breaking off of the
latest LLVM release as our host toolchain.
Co-authored-by: Jon Meow <46229924+jonmeow@users.noreply.github.com>
This picks up a newer version of LLVM and the LLVM Bazel integration.
The big change here is that we can configure the LLVM targets that are
built, which allows us to dramatically reduce the build costs by
focusing on a couple of CPUs for the time being.
There are a few API updates needed as well.
This also rotates the cache version so we start with a clean Bazel cache
from here. Otherwise we'd potentially pay the cost of carrying around
stale bits of LLVM endlessly.
What this does:
- Sets up a `migrate_cpp` tool which currently only runs `clang-tidy`.
- This is intended to have more transformations in the future.
- Sets up a `migrate_cpp.sh` script.
- This copies the original woff2 code into a `carbon` directory and runs the `migrate_cpp` tool on it there.
- Adds the initial `carbon` directory of woff2
- To be clear, this is currently only updated via `clang-tidy`.
- More transformations should be expected in the future.
- Minor related edits. For example:
- Adjust pre-commit to skip the `carbon` directory, because it's third-party code and shouldn't be edited in the same way.
- Adds `@brotli_carbon` as a local repository so that we can "build" outputs.
- Makes clang-tidy from the bootstrap toolchain accessible for BUILD dependencies, as it's then used for `migrate_cpp`.
What this does not do:
- Any actual transformation of C++ code to Carbon
Co-authored-by: Matthew Riley <mdriley@gmail.com>
This is to support converting the code to Carbon. My theory with the setup is:
- Have the code available to build in C++ under `third_party/<project>/original`.
- Created a `third_party/<project>/carbon` for the converted version.
Having the existing code building should, I think, make it easier to run analysis on said code. Using a submodule means we should be aiming to keep it pristine, for easy comparison / updates.
Originally, I experimented with special rules for C++ builds of LLVM but
we ended up with a native build of it instead. Loading and using this
was completely unnecessary now and would have needed an update. Just
remove it.
Fixes#400
This makes executable semantics build and pass tests for me without
installing either Bison or Flex. We just use the primitive toolchain
with the existing genrule as the packaged rules don't quite fit how
we're building and organizing the code.
Currently, this points at forks of the upstream rule repositories while
PRs I have sent there are going through, but this should be functional
for now and there doesn't seem to be any reason to wait for those PRs to
go through.
Per chandlerc's comment:
First, this builds inside the Bazel tree rather than in the source
repository. This ensures we start in a clean directory each time the
repository rule is run again. We don't need to detect an existing build
with this, and it will only be rebuilt when the workspace file or the
repository rule implementation is changed. This can be forced by using:
```
bazel sync --configure
```
Second, we use an implicit dependency on the `WORKSPACE` file to locate
the workspace directory automatically, and the `HEAD` file from the
`llvm-project` submodule to trigger a rebuild if the submodule is
updated.
Third, teach the CMake script to try to use a system-installed `clang`
if installed and not overridden by the `CC` environment variable. This
is very different from the prior logic -- this is only used with the
CMake build, and so should work with any system C++ compiler that can
build Clang and LLVM.
Lastly, this tweaks the CMake options to tune this build given that we
now fully control it and it will only be used in this context. This
still results in a 1.5gb build for me. =/ But its as small as I can make
it really. It's a frustrating long list, but I couldn't find a more
brief way of representing this.
Co-authored-by: Chandler Carruth <chandlerc@gmail.com>
The Bazel bits are collected into a directory and given less confusing
names (I hope). Other than names, everything is a direct copy from the
toolchain repository without any edits.
A few files needed to be merged in:
- `.gitignore`
- `.pre-commit-config.yaml`
Subsequent commits will add relevant C++ infrastructure and then the
source code itself.