## Summary ## Restructures the lexer to first scan the entire source text for newlines and create all the line structures needed. Doing this up-front makes it easy to produce an optimized version with minimal complexity. Currently, it leverages the system `memchr`, but even when expanded to handle more complex cases like CR+LF line endings, being isolated in this way will result in a significantly simpler implementation. This change improves the lexing of comment lines significantly by skipping their contents immediately. The overhead of the pre-scan is unmeasurable in all realistic benchmarks, and 10-30% in benchmarks consisting almost entirely of blank lines or comments. The improvement of comment lexing with average length comment lines mixed with code starts at 20% and goes up. Regressing blank line handling for non-empty comments seems like the right tradeoff (by far). ## Background and details ## One weak point in the lexer implementation were large runs of comments. While those aren't terribly common, they shouldn't present a hazard to the lexer performance. A bit more common is a pattern of comments like the following: ```carbon // Some method comment here. fn SomeMethodName(...) -> ...; // Some other method comment here. fn SomeOtherMethodName(...) -> ...; ``` Here, the lexer spends an inordinate amount of time getting from the `\n` after the first semicolon to the `fn` token. It has to skip a blank line, scan a line, find the `//` comment start, then scan to find the next `\n`, and then scan horizontal whitespace, etc. It is tempting to build a scanner *exactly* for this. In fact, I built one, and I can publish it in a PR if folks are interested in what it looks like. For x86-64, the PSHUFB trick used for scanning identifiers technically works. But it is *complicated*. Amazingly so. 150 lines of very subtle code with subtle performance pitfalls at ever turn. I felt very uncomfortable submitting it, but we can always go back to it. Nothing I've come up with quite matches it for sheer speed. However, most of the complexity and time is spent walking from a `//` to the end of the line. And *that* is something we can do very simply. In fact, there is a tuned function for that in libc: `memchr`. Using this we can build a very fast and much simpler scanner to split lines up-front. This PR uses that and a carefully crafted fast loop to first build up all the line info we need. Getting this to be as fast as possible required some other subtle changes, for example always creating a line structure that goes from the last `\n` and the end of the file. We then back up the EOF token to avoid surfacing this to users. The nice thing is the EOF token isn't part of any hot loop, and so this removes branches everywhere else at modest complexity. Once we have that, the rest of the lexer just needs to keep track of its current line in order to record column offsets. I've taken some care to try and optimize the lexer's usage of the line structures but there are more opportunities here I suspect. Combined, this gets much but not all of the performance of a huge SIMD scanner for newline-through-to-next-token. For extreme cases (100s of blank lines or empty comment lines between tokens) the holistic scanner is of course still much faster, but those don't seem nearly worth the cost. I was initially worried about the overhead of taking two passes over the source text, but in practice I've not been able to measure any appreciable cost to this with realistic source files. In some cases benchmarks with no newlines get *faster* because we use a much more efficient approach to fetch the source text into cache as a happenstance. And that in turn makes the byte-wise dispatched loop run faster as it stalls less. I'm particularly happy with this approach because it seems very clear how to extend this to support CR+LF, bare CR, and even complex mixtures without any significant speed cost. That wasn't at all true for the other approaches explored. I may try some further PRs to smooth out the last bits of slowness here, but already this is working excellent for me in practice. My 10mloc test case is down to 2.3s to lex. ## Raw benchmark data Using a tool that runs benchmarks before and after and analyzes the results, the following summarizes the CPU-time impact, each of these for lexing 100k tokens: ``` BM_ValidKeywords 2.57ms ± 1% 2.58ms ± 0% ~ (p=0.190 n=5+4) BM_ValidIdentifiers<1, 64, false> 9.24ms ± 4% 9.31ms ± 4% ~ (p=0.421 n=5+5) BM_ValidIdentifiers<1, 1, true> 3.05ms ± 4% 3.11ms ± 4% ~ (p=0.222 n=5+5) BM_ValidIdentifiers<3, 5, true> 10.9ms ± 0% 11.1ms ± 1% +1.76% (p=0.016 n=4+5) BM_ValidIdentifiers<3, 16, true> 11.1ms ± 7% 11.0ms ± 1% ~ (p=0.310 n=5+5) BM_ValidIdentifiers<12, 64, true> 12.2ms ± 1% 12.3ms ± 2% ~ (p=0.111 n=4+5) BM_HorizontalWhitespace/1 11.2ms ± 6% 11.1ms ± 2% ~ (p=0.841 n=5+5) BM_HorizontalWhitespace/4 12.0ms ± 3% 12.0ms ± 2% ~ (p=0.548 n=5+5) BM_HorizontalWhitespace/16 16.2ms ± 6% 15.9ms ± 8% ~ (p=0.690 n=5+5) BM_HorizontalWhitespace/64 27.7ms ± 3% 28.4ms ± 3% ~ (p=0.151 n=5+5) BM_HorizontalWhitespace/128 44.3ms ± 1% 45.6ms ± 6% +3.15% (p=0.032 n=5+5) BM_RandomSource 7.75ms ± 2% 7.72ms ± 1% ~ (p=1.000 n=5+5) BM_BlankLines/1 11.7ms ± 1% 12.1ms ± 1% +3.46% (p=0.008 n=5+5) BM_BlankLines/4 14.0ms ± 2% 15.2ms ± 3% +8.12% (p=0.008 n=5+5) BM_BlankLines/16 23.5ms ± 2% 31.1ms ± 4% +32.26% (p=0.008 n=5+5) BM_BlankLines/64 75.3ms ± 1% 81.2ms ± 3% +7.83% (p=0.008 n=5+5) BM_BlankLines/128 133ms ± 3% 150ms ± 2% +12.74% (p=0.008 n=5+5) BM_CommentLines/1/0/0 13.1ms ± 0% 13.7ms ± 1% +5.11% (p=0.008 n=5+5) BM_CommentLines/4/0/0 16.6ms ± 1% 18.2ms ± 4% +9.56% (p=0.008 n=5+5) BM_CommentLines/128/0/0 169ms ± 4% 182ms ± 1% +7.24% (p=0.008 n=5+5) BM_CommentLines/1/30/0 18.7ms ± 5% 14.1ms ± 0% -24.84% (p=0.008 n=5+5) BM_CommentLines/4/30/0 36.5ms ± 6% 20.6ms ± 3% -43.59% (p=0.008 n=5+5) BM_CommentLines/128/30/0 525ms ± 4% 198ms ± 1% -62.38% (p=0.008 n=5+5) BM_CommentLines/1/70/0 23.4ms ± 6% 14.7ms ± 2% -37.15% (p=0.008 n=5+5) BM_CommentLines/4/70/0 53.3ms ± 7% 22.4ms ± 4% -57.99% (p=0.008 n=5+5) BM_CommentLines/128/70/0 1.05s ± 4% 0.21s ± 2% -80.31% (p=0.008 n=5+5) BM_CommentLines/1/0/2 14.1ms ± 6% 14.3ms ± 1% ~ (p=0.151 n=5+5) BM_CommentLines/4/0/2 19.4ms ± 5% 20.1ms ± 1% ~ (p=0.151 n=5+5) BM_CommentLines/128/0/2 238ms ± 8% 229ms ± 0% ~ (p=0.151 n=5+5) BM_CommentLines/1/30/2 19.2ms ± 7% 14.6ms ± 1% -23.87% (p=0.008 n=5+5) BM_CommentLines/4/30/2 40.3ms ±13% 22.3ms ± 4% -44.63% (p=0.008 n=5+5) BM_CommentLines/128/30/2 568ms ± 7% 254ms ± 3% -55.28% (p=0.008 n=5+5) BM_CommentLines/1/70/2 23.3ms ± 1% 15.0ms ± 3% -35.61% (p=0.016 n=4+5) BM_CommentLines/4/70/2 57.2ms ± 9% 24.1ms ± 2% -57.81% (p=0.008 n=5+5) BM_CommentLines/128/70/2 1.07s ± 0% 0.26s ± 2% -75.51% (p=0.016 n=4+5) BM_CommentLines/1/0/8 15.9ms ± 7% 16.0ms ± 1% ~ (p=0.151 n=5+5) BM_CommentLines/4/0/8 24.2ms ± 6% 27.9ms ± 2% +15.36% (p=0.008 n=5+5) BM_CommentLines/128/0/8 386ms ± 4% 445ms ± 1% +15.28% (p=0.008 n=5+5) BM_CommentLines/1/30/8 20.6ms ± 5% 16.3ms ± 1% -20.95% (p=0.008 n=5+5) BM_CommentLines/4/30/8 45.3ms ± 6% 30.3ms ± 3% -32.98% (p=0.008 n=5+5) BM_CommentLines/128/30/8 699ms ± 3% 477ms ± 3% -31.83% (p=0.008 n=5+5) BM_CommentLines/1/70/8 25.7ms ± 5% 16.8ms ± 2% -34.67% (p=0.008 n=5+5) BM_CommentLines/4/70/8 62.0ms ± 4% 31.6ms ± 2% -49.10% (p=0.008 n=5+5) BM_CommentLines/128/70/8 1.20s ± 2% 0.48s ± 4% -59.60% (p=0.008 n=5+5) ``` The horizontal whitespace benchmark (and all of the non-line-oriented ones) are noisier than they appear here but do show some improvements (surprisingly). My guess is that it has a lot to do with system load, as the advantage is that we're using a vectorized loop to scan the text first and then doing the byte-dispatched loop. So when the cache is a bit slower to populate, the vectorized version starts to be faster. --------- Co-authored-by: Richard Smith <richard@metafoo.co.uk> Co-authored-by: josh11b <josh11b@users.noreply.github.com>
Carbon Language:
An experimental successor to C++
Why? | Goals | Status | Getting started | Join us
See our announcement video from CppNorth. Note that Carbon is not ready for use.
Fast and works with C++
- Performance matching C++ using LLVM, with low-level access to bits and addresses
- Interoperate with your existing C++ code, from inheritance to templates
- Fast and scalable builds that work with your existing C++ build systems
Modern and evolving
- Solid language foundations that are easy to learn, especially if you have used C++
- Easy, tool-based upgrades between Carbon versions
- Safer fundamentals, and an incremental path towards a memory-safe subset
Welcoming open-source community
- Clear goals and priorities with robust governance
- Community that works to be welcoming, inclusive, and friendly
- Batteries-included approach: compiler, libraries, docs, tools, package manager, and more
Why build Carbon?
C++ remains the dominant programming language for performance-critical software, with massive and growing codebases and investments. However, it is struggling to improve and meet developers' needs, as outlined above, in no small part due to accumulating decades of technical debt. Incrementally improving C++ is extremely difficult, both due to the technical debt itself and challenges with its evolution process. The best way to address these problems is to avoid inheriting the legacy of C or C++ directly, and instead start with solid language foundations like modern generics system, modular code organization, and consistent, simple syntax.
Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should. Unfortunately, the designs of these languages present significant barriers to adoption and migration from C++. These barriers range from changes in the idiomatic design of software to performance overhead.
Carbon is fundamentally a successor language approach, rather than an attempt to incrementally evolve C++. It is designed around interoperability with C++ as well as large-scale adoption and migration for existing C++ codebases and developers. A successor language for C++ requires:
- Performance matching C++, an essential property for our developers.
- Seamless, bidirectional interoperability with C++, such that a library anywhere in an existing C++ stack can adopt Carbon without porting the rest.
- A gentle learning curve with reasonable familiarity for C++ developers.
- Comparable expressivity and support for existing software's design and architecture.
- Scalable migration, with some level of source-to-source translation for idiomatic C++ code.
With this approach, we can build on top of C++'s existing ecosystem, and bring along existing investments, codebases, and developer populations. There are a few languages that have followed this model for other ecosystems, and Carbon aims to fill an analogous role for C++:
- JavaScript → TypeScript
- Java → Kotlin
- C++ → Carbon
Language Goals
We are designing Carbon to support:
- Performance-critical software
- Software and language evolution
- Code that is easy to read, understand, and write
- Practical safety and testing mechanisms
- Fast and scalable development
- Modern OS platforms, hardware architectures, and environments
- Interoperability with and migration from existing C++ code
While many languages share subsets of these goals, what distinguishes Carbon is their combination.
We also have explicit non-goals for Carbon, notably including:
- A stable application binary interface (ABI) for the entire language and library
- Perfect backwards or forwards compatibility
Our detailed goals document fleshes out these ideas and provides a deeper view into our goals for the Carbon project and language.
Project status
Carbon Language is currently an experimental project. There is no working compiler or toolchain. You can see the demo interpreter for Carbon on compiler-explorer.com.
We want to better understand whether we can build a language that meets our successor language criteria, and whether the resulting language can gather a critical mass of interest within the larger C++ industry and community.
Currently, we have fleshed out several core aspects of both Carbon the project and the language:
- The strategy of the Carbon Language and project.
- An open-source project structure, governance model, and evolution process.
- Critical and foundational aspects of the language design informed by our
experience with C++ and the most difficult challenges we anticipate. This
includes designs for:
- Generics
- Class types
- Inheritance
- Operator overloading
- Lexical and syntactic structure
- Code organization and modular structure
- A prototype interpreter demo that can both run isolated examples and gives a detailed analysis of the specific semantic model and abstract machine of Carbon. We call this the Carbon Explorer.
If you're interested in contributing, we would love help completing the 0.1 language designs, and completing the Carbon Explorer implementation of this design. We are also currently working to get more broad feedback and participation from the C++ community. Beyond that, we plan to prioritize C++ interoperability and a realistic toolchain that implements the 0.1 language and can be used to evaluate Carbon in more detail.
You can see our full roadmap for more details.
Carbon and C++
If you're already a C++ developer, Carbon should have a gentle learning curve. It is built out of a consistent set of language constructs that should feel familiar and be easy to read and understand.
C++ code like this:
corresponds to this Carbon code:
You can call Carbon from C++ without overhead and the other way around. This means you migrate a single C++ library to Carbon within an application, or write new Carbon on top of your existing C++ investment. For example:
Read more about C++ interop in Carbon.
Beyond interoperability between Carbon and C++, we're also planning to support migration tools that will mechanically translate idiomatic C++ code into Carbon code to help you switch an existing C++ codebase to Carbon.
Generics
Carbon provides a modern generics system with checked definitions, while still supporting opt-in templates for seamless C++ interop. Checked generics provide several advantages compared to C++ templates:
- Generic definitions are fully type-checked, removing the need to
instantiate to check for errors and giving greater confidence in code.
- Avoids the compile-time cost of re-checking the definition for every instantiation.
- When using a definition-checked generic, usage error messages are clearer, directly showing which requirements are not met.
- Enables automatic, opt-in type erasure and dynamic dispatch without a separate implementation. This can reduce the binary size and enables constructs like heterogeneous containers.
- Strong, checked interfaces mean fewer accidental dependencies on implementation details and a clearer contract for consumers.
Without sacrificing these advantages, Carbon generics support specialization, ensuring it can fully address performance-critical use cases of C++ templates. For more details about Carbon's generics, see their design.
In addition to easy and powerful interop with C++, Carbon templates can be constrained and incrementally migrated to checked generics at a fine granularity and with a smooth evolutionary path.
Memory safety
Safety, and especially memory safety, remains a key challenge for C++ and something a successor language needs to address. Our initial priority and focus is on immediately addressing important, low-hanging fruit in the safety space:
- Tracking uninitialized states better, increased enforcement of initialization, and systematically providing hardening against initialization bugs when desired.
- Designing fundamental APIs and idioms to support dynamic bounds checks in debug and hardened builds.
- Having a default debug build mode that is both cheaper and more comprehensive than existing C++ build modes even when combined with Address Sanitizer.
Once we can migrate code into Carbon, we will have a simplified language with room in the design space to add any necessary annotations or features, and infrastructure like generics to support safer design patterns. Longer term, we will build on this to introduce a safe Carbon subset. This will be a large and complex undertaking, and won't be in the 0.1 design. Meanwhile, we are closely watching and learning from efforts to add memory safe semantics onto C++ such as Rust-inspired lifetime annotations.
Getting started
As there is no compiler yet, to try out Carbon, you can use the Carbon explorer to interpret Carbon code and print its output. You can try it out immediately at compiler-explorer.com.
To build the Carbon explorer yourself, you'll need to install dependencies (Bazel, Clang, libc++), and then you can run:
# Download Carbon's code.
$ git clone https://github.com/carbon-language/carbon-lang
$ cd carbon-lang
# Build and run the explorer.
$ bazel run //explorer -- ./explorer/testdata/print/format_only.carbon
For complete instructions, including installing dependencies, see our contribution tools documentation.
Learn more about the Carbon project:
Conference talks
Past Carbon focused talks from the community:
2022
- Carbon Language: An experimental successor to C++, CppNorth
- Carbon Language: Syntax and trade-offs, Core C++
2023
- Carbon’s Successor Strategy: From C++ interop to memory safety, C++Now
- Definition-Checked Generics (Part 1, Part 2), C++Now
- Modernizing Compiler Design for Carbon’s Toolchain, C++Now
Join us
We'd love to have folks join us and contribute to the project. Carbon is committed to a welcoming and inclusive environment where everyone can contribute.
- Most of Carbon's design discussions occur on Discord.
- Carbon is a Google Summer of Code 2023 organization.
- To watch for major release announcements, subscribe to our Carbon release post on GitHub and star carbon-lang.
- See our code of conduct and contributing guidelines for information about the Carbon development community.
Contributing
You can also directly:
- Contribute to the language design: feedback on design, new design proposal
- Contribute to the language implementation
- Carbon Explorer: bug report, bug fix, language feature implementation
- Carbon Toolchain, and project infrastructure
You can check out some
"good first issues",
or join the #contributing-help channel on
Discord. See our full
CONTRIBUTING documentation for more details.
