6ba8712fbd Predetermine all the line splits in the lexer. (#3278)
## Summary ##

Restructures the lexer to first scan the entire source text for newlines
and create all the line structures needed. Doing this up-front makes it
easy to produce an optimized version with minimal complexity. Currently,
it leverages the system `memchr`, but even when expanded to handle more
complex cases like CR+LF line endings, being isolated in this way will
result in a significantly simpler implementation. This change improves
the lexing of comment lines significantly by skipping their contents
immediately. The overhead of the pre-scan is unmeasurable in all
realistic benchmarks, and 10-30% in benchmarks consisting almost
entirely of blank lines or comments. The improvement of comment lexing
with average length comment lines mixed with code starts at 20% and goes
up. Regressing blank line handling for non-empty comments seems like the
right tradeoff (by far).

## Background and details ##

One weak point in the lexer implementation were large runs of comments.
While those aren't terribly common, they shouldn't present a hazard to
the lexer performance.

A bit more common is a pattern of comments like the following:

```carbon
  // Some method comment here.
  fn SomeMethodName(...) -> ...;

  // Some other method comment here.
  fn SomeOtherMethodName(...) -> ...;
```

Here, the lexer spends an inordinate amount of time getting from the
`\n` after the first semicolon to the `fn` token. It has to skip a blank
line, scan a line, find the `//` comment start, then scan to find the
next `\n`, and then scan horizontal whitespace, etc.

It is tempting to build a scanner *exactly* for this. In fact, I built
one, and I can publish it in a PR if folks are interested in what it
looks like. For x86-64, the PSHUFB trick used for scanning identifiers
technically works. But it is *complicated*. Amazingly so. 150 lines of
very subtle code with subtle performance pitfalls at ever turn. I felt
very uncomfortable submitting it, but we can always go back to it.
Nothing I've come up with quite matches it for sheer speed.

However, most of the complexity and time is spent walking from a `//` to
the end of the line. And *that* is something we can do very simply. In
fact, there is a tuned function for that in libc: `memchr`. Using this
we can build a very fast and much simpler scanner to split lines
up-front. This PR uses that and a carefully crafted fast loop to first
build up all the line info we need. Getting this to be as fast as
possible required some other subtle changes, for example always creating
a line structure that goes from the last `\n` and the end of the file.
We then back up the EOF token to avoid surfacing this to users. The nice
thing is the EOF token isn't part of any hot loop, and so this removes
branches everywhere else at modest complexity.

Once we have that, the rest of the lexer just needs to keep track of its
current line in order to record column offsets. I've taken some care to
try and optimize the lexer's usage of the line structures but there are
more opportunities here I suspect.

Combined, this gets much but not all of the performance of a huge SIMD
scanner for newline-through-to-next-token. For extreme cases (100s of
blank lines or empty comment lines between tokens) the holistic scanner
is of course still much faster, but those don't seem nearly worth the
cost.

I was initially worried about the overhead of taking two passes over the
source text, but in practice I've not been able to measure any
appreciable cost to this with realistic source files. In some cases
benchmarks with no newlines get *faster* because we use a much more
efficient approach to fetch the source text into cache as
a happenstance. And that in turn makes the byte-wise dispatched loop run
faster as it stalls less.

I'm particularly happy with this approach because it seems very clear
how to extend this to support CR+LF, bare CR, and even complex mixtures
without any significant speed cost. That wasn't at all true for the
other approaches explored.

I may try some further PRs to smooth out the last bits of slowness here,
but already this is working excellent for me in practice. My 10mloc test
case is down to 2.3s to lex.

## Raw benchmark data

Using a tool that runs benchmarks before and after and analyzes the
results, the following summarizes the CPU-time impact, each of these for
lexing 100k tokens:

```
BM_ValidKeywords                               2.57ms ± 1%  2.58ms ± 0%     ~     (p=0.190 n=5+4)
BM_ValidIdentifiers<1, 64, false>              9.24ms ± 4%  9.31ms ± 4%     ~     (p=0.421 n=5+5)
BM_ValidIdentifiers<1, 1, true>                3.05ms ± 4%  3.11ms ± 4%     ~     (p=0.222 n=5+5)
BM_ValidIdentifiers<3, 5, true>                10.9ms ± 0%  11.1ms ± 1%   +1.76%  (p=0.016 n=4+5)
BM_ValidIdentifiers<3, 16, true>               11.1ms ± 7%  11.0ms ± 1%     ~     (p=0.310 n=5+5)
BM_ValidIdentifiers<12, 64, true>              12.2ms ± 1%  12.3ms ± 2%     ~     (p=0.111 n=4+5)
BM_HorizontalWhitespace/1                      11.2ms ± 6%  11.1ms ± 2%     ~     (p=0.841 n=5+5)
BM_HorizontalWhitespace/4                      12.0ms ± 3%  12.0ms ± 2%     ~     (p=0.548 n=5+5)
BM_HorizontalWhitespace/16                     16.2ms ± 6%  15.9ms ± 8%     ~     (p=0.690 n=5+5)
BM_HorizontalWhitespace/64                     27.7ms ± 3%  28.4ms ± 3%     ~     (p=0.151 n=5+5)
BM_HorizontalWhitespace/128                    44.3ms ± 1%  45.6ms ± 6%   +3.15%  (p=0.032 n=5+5)
BM_RandomSource                                7.75ms ± 2%  7.72ms ± 1%     ~     (p=1.000 n=5+5)
BM_BlankLines/1                                11.7ms ± 1%  12.1ms ± 1%   +3.46%  (p=0.008 n=5+5)
BM_BlankLines/4                                14.0ms ± 2%  15.2ms ± 3%   +8.12%  (p=0.008 n=5+5)
BM_BlankLines/16                               23.5ms ± 2%  31.1ms ± 4%  +32.26%  (p=0.008 n=5+5)
BM_BlankLines/64                               75.3ms ± 1%  81.2ms ± 3%   +7.83%  (p=0.008 n=5+5)
BM_BlankLines/128                               133ms ± 3%   150ms ± 2%  +12.74%  (p=0.008 n=5+5)
BM_CommentLines/1/0/0                          13.1ms ± 0%  13.7ms ± 1%   +5.11%  (p=0.008 n=5+5)
BM_CommentLines/4/0/0                          16.6ms ± 1%  18.2ms ± 4%   +9.56%  (p=0.008 n=5+5)
BM_CommentLines/128/0/0                         169ms ± 4%   182ms ± 1%   +7.24%  (p=0.008 n=5+5)
BM_CommentLines/1/30/0                         18.7ms ± 5%  14.1ms ± 0%  -24.84%  (p=0.008 n=5+5)
BM_CommentLines/4/30/0                         36.5ms ± 6%  20.6ms ± 3%  -43.59%  (p=0.008 n=5+5)
BM_CommentLines/128/30/0                        525ms ± 4%   198ms ± 1%  -62.38%  (p=0.008 n=5+5)
BM_CommentLines/1/70/0                         23.4ms ± 6%  14.7ms ± 2%  -37.15%  (p=0.008 n=5+5)
BM_CommentLines/4/70/0                         53.3ms ± 7%  22.4ms ± 4%  -57.99%  (p=0.008 n=5+5)
BM_CommentLines/128/70/0                        1.05s ± 4%   0.21s ± 2%  -80.31%  (p=0.008 n=5+5)
BM_CommentLines/1/0/2                          14.1ms ± 6%  14.3ms ± 1%     ~     (p=0.151 n=5+5)
BM_CommentLines/4/0/2                          19.4ms ± 5%  20.1ms ± 1%     ~     (p=0.151 n=5+5)
BM_CommentLines/128/0/2                         238ms ± 8%   229ms ± 0%     ~     (p=0.151 n=5+5)
BM_CommentLines/1/30/2                         19.2ms ± 7%  14.6ms ± 1%  -23.87%  (p=0.008 n=5+5)
BM_CommentLines/4/30/2                         40.3ms ±13%  22.3ms ± 4%  -44.63%  (p=0.008 n=5+5)
BM_CommentLines/128/30/2                        568ms ± 7%   254ms ± 3%  -55.28%  (p=0.008 n=5+5)
BM_CommentLines/1/70/2                         23.3ms ± 1%  15.0ms ± 3%  -35.61%  (p=0.016 n=4+5)
BM_CommentLines/4/70/2                         57.2ms ± 9%  24.1ms ± 2%  -57.81%  (p=0.008 n=5+5)
BM_CommentLines/128/70/2                        1.07s ± 0%   0.26s ± 2%  -75.51%  (p=0.016 n=4+5)
BM_CommentLines/1/0/8                          15.9ms ± 7%  16.0ms ± 1%     ~     (p=0.151 n=5+5)
BM_CommentLines/4/0/8                          24.2ms ± 6%  27.9ms ± 2%  +15.36%  (p=0.008 n=5+5)
BM_CommentLines/128/0/8                         386ms ± 4%   445ms ± 1%  +15.28%  (p=0.008 n=5+5)
BM_CommentLines/1/30/8                         20.6ms ± 5%  16.3ms ± 1%  -20.95%  (p=0.008 n=5+5)
BM_CommentLines/4/30/8                         45.3ms ± 6%  30.3ms ± 3%  -32.98%  (p=0.008 n=5+5)
BM_CommentLines/128/30/8                        699ms ± 3%   477ms ± 3%  -31.83%  (p=0.008 n=5+5)
BM_CommentLines/1/70/8                         25.7ms ± 5%  16.8ms ± 2%  -34.67%  (p=0.008 n=5+5)
BM_CommentLines/4/70/8                         62.0ms ± 4%  31.6ms ± 2%  -49.10%  (p=0.008 n=5+5)
BM_CommentLines/128/70/8                        1.20s ± 2%   0.48s ± 4%  -59.60%  (p=0.008 n=5+5)
```

The horizontal whitespace benchmark (and all of the non-line-oriented
ones) are noisier than they appear here but do show some improvements
(surprisingly). My guess is that it has a lot to do with system load, as
the advantage is that we're using a vectorized loop to scan the text
first and then doing the byte-dispatched loop. So when the cache is
a bit slower to populate, the vectorized version starts to be faster.

---------

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
Co-authored-by: josh11b <josh11b@users.noreply.github.com>
2023-10-12 07:13:44 +00:00
2023-10-11 05:39:59 +00:00
2023-06-16 11:11:46 -07:00
2021-08-25 14:15:01 -07:00
2020-06-16 17:58:17 -07:00

Carbon Language:
An experimental successor to C++

Why? | Goals | Status | Getting started | Join us

See our announcement video from CppNorth. Note that Carbon is not ready for use.

Quicksort code in Carbon. Follow the link to read more.

Fast and works with C++

  • Performance matching C++ using LLVM, with low-level access to bits and addresses
  • Interoperate with your existing C++ code, from inheritance to templates
  • Fast and scalable builds that work with your existing C++ build systems

Modern and evolving

  • Solid language foundations that are easy to learn, especially if you have used C++
  • Easy, tool-based upgrades between Carbon versions
  • Safer fundamentals, and an incremental path towards a memory-safe subset

Welcoming open-source community

  • Clear goals and priorities with robust governance
  • Community that works to be welcoming, inclusive, and friendly
  • Batteries-included approach: compiler, libraries, docs, tools, package manager, and more

Why build Carbon?

C++ remains the dominant programming language for performance-critical software, with massive and growing codebases and investments. However, it is struggling to improve and meet developers' needs, as outlined above, in no small part due to accumulating decades of technical debt. Incrementally improving C++ is extremely difficult, both due to the technical debt itself and challenges with its evolution process. The best way to address these problems is to avoid inheriting the legacy of C or C++ directly, and instead start with solid language foundations like modern generics system, modular code organization, and consistent, simple syntax.

Existing modern languages already provide an excellent developer experience: Go, Swift, Kotlin, Rust, and many more. Developers that can use one of these existing languages should. Unfortunately, the designs of these languages present significant barriers to adoption and migration from C++. These barriers range from changes in the idiomatic design of software to performance overhead.

Carbon is fundamentally a successor language approach, rather than an attempt to incrementally evolve C++. It is designed around interoperability with C++ as well as large-scale adoption and migration for existing C++ codebases and developers. A successor language for C++ requires:

  • Performance matching C++, an essential property for our developers.
  • Seamless, bidirectional interoperability with C++, such that a library anywhere in an existing C++ stack can adopt Carbon without porting the rest.
  • A gentle learning curve with reasonable familiarity for C++ developers.
  • Comparable expressivity and support for existing software's design and architecture.
  • Scalable migration, with some level of source-to-source translation for idiomatic C++ code.

With this approach, we can build on top of C++'s existing ecosystem, and bring along existing investments, codebases, and developer populations. There are a few languages that have followed this model for other ecosystems, and Carbon aims to fill an analogous role for C++:

  • JavaScript → TypeScript
  • Java → Kotlin
  • C++ → Carbon

Language Goals

We are designing Carbon to support:

  • Performance-critical software
  • Software and language evolution
  • Code that is easy to read, understand, and write
  • Practical safety and testing mechanisms
  • Fast and scalable development
  • Modern OS platforms, hardware architectures, and environments
  • Interoperability with and migration from existing C++ code

While many languages share subsets of these goals, what distinguishes Carbon is their combination.

We also have explicit non-goals for Carbon, notably including:

Our detailed goals document fleshes out these ideas and provides a deeper view into our goals for the Carbon project and language.

Project status

Carbon Language is currently an experimental project. There is no working compiler or toolchain. You can see the demo interpreter for Carbon on compiler-explorer.com.

We want to better understand whether we can build a language that meets our successor language criteria, and whether the resulting language can gather a critical mass of interest within the larger C++ industry and community.

Currently, we have fleshed out several core aspects of both Carbon the project and the language:

  • The strategy of the Carbon Language and project.
  • An open-source project structure, governance model, and evolution process.
  • Critical and foundational aspects of the language design informed by our experience with C++ and the most difficult challenges we anticipate. This includes designs for:
    • Generics
    • Class types
    • Inheritance
    • Operator overloading
    • Lexical and syntactic structure
    • Code organization and modular structure
  • A prototype interpreter demo that can both run isolated examples and gives a detailed analysis of the specific semantic model and abstract machine of Carbon. We call this the Carbon Explorer.

If you're interested in contributing, we would love help completing the 0.1 language designs, and completing the Carbon Explorer implementation of this design. We are also currently working to get more broad feedback and participation from the C++ community. Beyond that, we plan to prioritize C++ interoperability and a realistic toolchain that implements the 0.1 language and can be used to evaluate Carbon in more detail.

You can see our full roadmap for more details.

Carbon and C++

If you're already a C++ developer, Carbon should have a gentle learning curve. It is built out of a consistent set of language constructs that should feel familiar and be easy to read and understand.

C++ code like this:

A snippet of C++ code. Follow the link to read it.

corresponds to this Carbon code:

A snippet of converted Carbon code. Follow the link to read it.

You can call Carbon from C++ without overhead and the other way around. This means you migrate a single C++ library to Carbon within an application, or write new Carbon on top of your existing C++ investment. For example:

A snippet of mixed Carbon and C++ code. Follow the link to read it.

Read more about C++ interop in Carbon.

Beyond interoperability between Carbon and C++, we're also planning to support migration tools that will mechanically translate idiomatic C++ code into Carbon code to help you switch an existing C++ codebase to Carbon.

Generics

Carbon provides a modern generics system with checked definitions, while still supporting opt-in templates for seamless C++ interop. Checked generics provide several advantages compared to C++ templates:

  • Generic definitions are fully type-checked, removing the need to instantiate to check for errors and giving greater confidence in code.
    • Avoids the compile-time cost of re-checking the definition for every instantiation.
    • When using a definition-checked generic, usage error messages are clearer, directly showing which requirements are not met.
  • Enables automatic, opt-in type erasure and dynamic dispatch without a separate implementation. This can reduce the binary size and enables constructs like heterogeneous containers.
  • Strong, checked interfaces mean fewer accidental dependencies on implementation details and a clearer contract for consumers.

Without sacrificing these advantages, Carbon generics support specialization, ensuring it can fully address performance-critical use cases of C++ templates. For more details about Carbon's generics, see their design.

In addition to easy and powerful interop with C++, Carbon templates can be constrained and incrementally migrated to checked generics at a fine granularity and with a smooth evolutionary path.

Memory safety

Safety, and especially memory safety, remains a key challenge for C++ and something a successor language needs to address. Our initial priority and focus is on immediately addressing important, low-hanging fruit in the safety space:

  • Tracking uninitialized states better, increased enforcement of initialization, and systematically providing hardening against initialization bugs when desired.
  • Designing fundamental APIs and idioms to support dynamic bounds checks in debug and hardened builds.
  • Having a default debug build mode that is both cheaper and more comprehensive than existing C++ build modes even when combined with Address Sanitizer.

Once we can migrate code into Carbon, we will have a simplified language with room in the design space to add any necessary annotations or features, and infrastructure like generics to support safer design patterns. Longer term, we will build on this to introduce a safe Carbon subset. This will be a large and complex undertaking, and won't be in the 0.1 design. Meanwhile, we are closely watching and learning from efforts to add memory safe semantics onto C++ such as Rust-inspired lifetime annotations.

Getting started

As there is no compiler yet, to try out Carbon, you can use the Carbon explorer to interpret Carbon code and print its output. You can try it out immediately at compiler-explorer.com.

To build the Carbon explorer yourself, you'll need to install dependencies (Bazel, Clang, libc++), and then you can run:

# Download Carbon's code.
$ git clone https://github.com/carbon-language/carbon-lang
$ cd carbon-lang

# Build and run the explorer.
$ bazel run //explorer -- ./explorer/testdata/print/format_only.carbon

For complete instructions, including installing dependencies, see our contribution tools documentation.

Learn more about the Carbon project:

Conference talks

Past Carbon focused talks from the community:

2022

2023

Join us

We'd love to have folks join us and contribute to the project. Carbon is committed to a welcoming and inclusive environment where everyone can contribute.

Contributing

You can also directly:

You can check out some "good first issues", or join the #contributing-help channel on Discord. See our full CONTRIBUTING documentation for more details.

S
Description
Carbon Language's main repository: documents, design, implementation, and related tools. (NOTE: Carbon Language is experimental; see README)
Readme Apache-2.0
2.7 GiB
Languages
C++ 89.5%
Starlark 4.4%
Python 3.6%
Carbon 1.5%
JavaScript 0.4%
Other 0.3%