Files
carbon-lang/testing/base/source_gen.h
T
a9c815c9f4 Introduce a source generator and end-to-end compile benchmarks (#4124)
The big addition here is a very, very rough and very early skeleton of a
source code generator framework. This builds upon the lexers identifier
synthesis logic, improving on its framework and wiring it up with the
most rudimentary of source file generation. This is just enough to
roughly replicate my "big API file" source code benchmarks.

The source generation works *very* hard to both vary the structure and
content of the source as much as possible while ensuring the same
*total* amount of each construct is in use, from bytes in identifiers to
line breaks, parameters, etc. This lets us generate randomly structure
inputs that should consistently take the exact same amount of total work
to compile.

The complex identifier synthesis logic from the lexer's benchmark is
moved over here and the lexer uses APIs in the source generator for
identifiers. The other source synthesis in the lexer's benchmark isn't
yet moved over, but should likely be slowly absorbed here as it can be
refactored into a more principled and re-usable form. Some bits may stay
of course if they're just too lexer-specific.

Next, this adds a simple end-to-end compile benchmark for the driver
that directly and much more clearly reproduces all the measurements I've
done manually up until now. It should also be easy to extend to more
patterns over time as we add support to the source generator to produce
those patterns.

Last but not least, I've added a tiny CLI to the source generator so
that you can generate source code manually. This is especially nice for
generating demo source code to actually run through the driver or look
at in an editor. The CLI can also generate C++ source code which lets us
do some minimal comparative benchmarking between Carbon and C++/Clang.

There are huge number of TODOs in the source generation framework. This
is going to be a large ongoing effort I suspect.

There are also a bunch of rough edges I've left to try and get this out
for review sooner. I've left TODOs for refactorings that really need to
be done here, but hoping these can maybe be follow-ups. If not, please
flag and I'll try to layer them on here.

Sample compile benchmark output, nicely showing where we are w.r.t. our
goal speeds (2x behind on lex and check, 5x on parse) at least on a
recent AMD server CPU:
```
------------------------------------------------------------------------------------------------------
Benchmark                                                 Time             CPU   Iterations      Lines
------------------------------------------------------------------------------------------------------
BM_CompileAPIFileDenseDecls<Phase::Lex>/256           29420 ns        29419 ns        22860 6.62847M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/1024         146130 ns       146128 ns         4840 6.69959M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/4096         601584 ns       601577 ns         1020 6.69573M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/16384       2547578 ns      2547313 ns          280   6.404M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/65536      10816591 ns     10816389 ns           80 6.05193M/s
BM_CompileAPIFileDenseDecls<Phase::Lex>/262144     52191320 ns     52189828 ns           20 5.02261M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/256        101706 ns       101698 ns         6900 1.91745M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/1024       512161 ns       512162 ns         1380  1.9115M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/4096      2078426 ns      2078430 ns          340   1.938M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/16384     8795786 ns      8795583 ns          100 1.85468M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/65536    35073596 ns     35072973 ns           20 1.86639M/s
BM_CompileAPIFileDenseDecls<Phase::Parse>/262144  151100688 ns    151097370 ns           20 1.73483M/s
BM_CompileAPIFileDenseDecls<Phase::Check>/256        957059 ns       957049 ns          740 203.751k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/1024      1956134 ns      1955985 ns          360 500.515k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/4096      5797864 ns      5797417 ns          120 694.792k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/16384    21219608 ns     21217584 ns           40 768.843k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/65536    96311116 ns     96302334 ns           20 679.734k/s
BM_CompileAPIFileDenseDecls<Phase::Check>/262144  371637963 ns    371609964 ns           20 705.387k/s
```

Lest someone think this is *bad*, the fact that we're already within 2x
of our rather audacious goals makes me quite happy. =D

---------

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2024-08-13 19:37:55 +00:00

248 lines
10 KiB
C++

// Part of the Carbon Language project, under the Apache License v2.0 with LLVM
// Exceptions. See /LICENSE for license information.
// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
#ifndef CARBON_TESTING_BASE_SOURCE_GEN_H_
#define CARBON_TESTING_BASE_SOURCE_GEN_H_
#include <string>
#include "absl/random/random.h"
#include "common/map.h"
#include "common/set.h"
#include "llvm/ADT/ArrayRef.h"
#include "llvm/ADT/StringRef.h"
#include "llvm/Support/Allocator.h"
namespace Carbon::Testing {
// Provides source code generation facilities.
//
// This class works to generate valid but random & meaningless source code in
// interesting patterns for benchmarking. It is very incomplete. A high level
// set of long-term goals:
//
// - Generate interesting patterns and structures of code that have emerged as
// toolchain performance bottlenecks in practice in C++ codebases.
// - Generate code that includes most Carbon language features (and whatever
// reasonable C++ analogs could be used for comparative purposes):
// - Functions
// - Classes with class functions, methods, and fields
// - Interfaces
// - Checked generics and templates
// - Nested and unnested impls
// - Nested classes
// - Inline and out-of-line function and method definitions
// - Imports and exports
// - API files and impl files.
// - Be random but deterministic. The goal is benchmarking and so while this
// code should strive for not producing trivially predictable patterns, it
// should also strive to be consistent and suitable for benchmarking. Wherever
// possible, it should permute the order and content without randomizing the
// total count, size, or complexity.
//
// Note that the default and primary generation target is interesting Carbon
// source code. We have a best-effort to alternatively generate comparable C++
// constructs to the Carbon ones for comparative benchmarking, but there is no
// goal to cover all the interesting C++ patterns we might want to benchmark,
// and we don't aim for perfectly synthesizing C++ analogs. We can always drop
// fidelity for the C++ code path if needed for simplicity.
//
// TODO: There are numerous places where we hard code a fixed quantity. Instead,
// we should build a rich but general system to easily encode a discrete
// distribution that is sampled. We have a specialized version of this for
// identifiers that should be generalized.
class SourceGen {
public:
enum class Language {
Carbon,
Cpp,
};
struct FunctionDeclParams {
// TODD: Arbitrary default, should switch to a distribution from data.
int max_params = 4;
};
struct MethodDeclParams {
// TODD: Arbitrary default, should switch to a distribution from data.
int max_params = 4;
};
// Parameters used to generate a class in a generated file.
//
// Currently, this uses a fixed number of each kind of declaration, with
// arbitrary defaults chosen. The defaults currently skew towards large
// classes with lots of nested declarations.
// TODO: Switch these to distributions based on data.
//
// TODO: Add support for generating definitions and parameters to control
// them.
struct ClassParams {
int public_function_decls = 4;
FunctionDeclParams public_function_decl_params = {.max_params = 8};
int public_method_decls = 10;
MethodDeclParams public_method_decl_params;
int private_function_decls = 2;
FunctionDeclParams private_function_decl_params = {.max_params = 6};
int private_method_decls = 8;
MethodDeclParams private_method_decl_params = {.max_params = 6};
int private_field_decls = 6;
};
// Parameters used to generate a file with dense declarations.
struct DenseDeclParams {
// TODO: Add more parameters to control generating top-level constructs
// other than class definitions.
// Parameters used when generating class definitions.
ClassParams class_params = {};
};
// Access a global instance of this type to generate Carbon code for
// benchmarks, tests, or other places where sharing a common instance is
// useful. Note that there is nothing thread safe about this instance or type.
static auto Global() -> SourceGen&;
// Construct a source generator for the provided language, by default Carbon.
explicit SourceGen(Language language = Language::Carbon);
// Generate an API file with dense classes containing function forward
// declarations.
//
// Accepts a number of `target_lines` for the resulting source code. This is a
// rough approximation used to scale all the other constructs up and down
// accordingly. For C++ source generation, we work to generate the same number
// of constructs as Carbon would for the given line count over keeping the
// actual line count close to the target.
//
// TODO: Currently, the formatting and line breaks of generating code are
// extremely rough still, and those are a large factor in adherence to
// `target_lines`. Long term, the goal is to get as close as we can to any
// automatically formatted code while still keeping the stability of
// benchmarking.
auto GenAPIFileDenseDecls(int target_lines, DenseDeclParams params)
-> std::string;
// Get some number of randomly shuffled identifiers.
//
// The identifiers start with a character [A-Za-z], other characters may also
// include [0-9_]. Both Carbon and C++ keywords are excluded along with any
// other non-identifier syntaxes that overlap to ensure all of these can be
// used as identifiers.
//
// The order will be different for each call to this function, but the
// specific identifiers may remain the same in order to reduce the cost of
// repeated calls. However, the sum of the identifier sizes returned is
// guaranteed to be the same for every call with the same number of
// identifiers so that benchmarking all of these identifiers has predictable
// and stable cost.
//
// Optionally, callers can request a minimum and maximum length. By default,
// the length distribution used across the identifiers will mirror the
// observed distribution of identifiers in C++ source code and our expectation
// of them in Carbon source code. The maximum length in this default
// distribution cannot be more than 64.
//
// Callers can request a uniform distribution across [min_length, max_length],
// and when it is requested there is no limit on `max_length`.
auto GetShuffledIdentifiers(int number, int min_length = 1,
int max_length = 64, bool uniform = false)
-> llvm::SmallVector<llvm::StringRef>;
// Same as `GetShuffledIdentifiers`, but ensures there are no collisions.
auto GetShuffledUniqueIdentifiers(int number, int min_length = 4,
int max_length = 64, bool uniform = false)
-> llvm::SmallVector<llvm::StringRef>;
// Returns a collection of un-shuffled identifiers, otherwise the same as
// `GetShuffledIdentifiers`.
//
// Usually, benchmarks should use the shuffled version. However, this is
// useful when there is already a post-processing step to shuffle things as it
// is *dramatically* more efficient, especially in debug builds.
auto GetIdentifiers(int number, int min_length = 1, int max_length = 64,
bool uniform = false)
-> llvm::SmallVector<llvm::StringRef>;
// Returns a collection of un-shuffled unique identifiers, otherwise the same
// as `GetShuffledUniqueIdentifiers`.
//
// Usually, benchmarks should use the shuffled version. However, this is
// useful when there is already a post-processing step to shuffle things.
auto GetUniqueIdentifiers(int number, int min_length = 1, int max_length = 64,
bool uniform = false)
-> llvm::SmallVector<llvm::StringRef>;
// Returns a shared collection of random identifiers of a specific length.
//
// For a single, exact length, we have an even cheaper routine to return
// access to a shared collection of identifiers. The order of these is a
// single fixed random order for a given execution. The returned array
// reference is only valid until the next call any method on this generator.
auto GetSingleLengthIdentifiers(int length, int number)
-> llvm::ArrayRef<llvm::StringRef>;
private:
// The shuffled state used to generate some number of classes.
//
// This state encodes all the shuffled entropy used for generating a number of
// class definitions. While generating definitions, the state here will be
// consumed until empty.
struct ClassGenState {
llvm::SmallVector<int> public_function_param_counts;
llvm::SmallVector<int> public_method_param_counts;
llvm::SmallVector<int> private_function_param_counts;
llvm::SmallVector<int> private_method_param_counts;
llvm::SmallVector<llvm::StringRef> class_names;
llvm::SmallVector<llvm::StringRef> member_names;
llvm::SmallVector<llvm::StringRef> param_names;
};
class UniqueIdentifierPopper;
friend UniqueIdentifierPopper;
using AppendFn = auto(int length, int number,
llvm::SmallVectorImpl<llvm::StringRef>& dest) -> void;
auto IsCpp() -> bool { return language_ == Language::Cpp; }
auto GenerateRandomIdentifier(llvm::MutableArrayRef<char> dest_storage)
-> void;
auto AppendUniqueIdentifiers(int length, int number,
llvm::SmallVectorImpl<llvm::StringRef>& dest)
-> void;
auto GetIdentifiersImpl(int number, int min_length, int max_length,
bool uniform, llvm::function_ref<AppendFn> append)
-> llvm::SmallVector<llvm::StringRef>;
auto GetShuffledInts(int number, int min, int max) -> llvm::SmallVector<int>;
auto GetClassGenState(int number, ClassParams params) -> ClassGenState;
auto GenerateFunctionDecl(llvm::StringRef name, bool is_private,
bool is_method, int param_count,
llvm::StringRef indent,
llvm::SmallVectorImpl<llvm::StringRef>& param_names,
llvm::raw_ostream& os) -> void;
auto GenerateClassDef(const ClassParams& params, ClassGenState& state,
llvm::raw_ostream& os) -> void;
absl::BitGen rng_;
llvm::BumpPtrAllocator storage_;
Map<int, llvm::SmallVector<llvm::StringRef>> identifiers_by_length_;
Map<int, std::pair<int, Set<llvm::StringRef>>> unique_identifiers_by_length_;
Language language_;
};
} // namespace Carbon::Testing
#endif // CARBON_TESTING_BASE_SOURCE_GEN_H_