mirror of
https://github.com/carbon-language/carbon-lang.git
synced 2026-09-29 08:44:57 +01:00
The big addition here is a very, very rough and very early skeleton of a source code generator framework. This builds upon the lexers identifier synthesis logic, improving on its framework and wiring it up with the most rudimentary of source file generation. This is just enough to roughly replicate my "big API file" source code benchmarks. The source generation works *very* hard to both vary the structure and content of the source as much as possible while ensuring the same *total* amount of each construct is in use, from bytes in identifiers to line breaks, parameters, etc. This lets us generate randomly structure inputs that should consistently take the exact same amount of total work to compile. The complex identifier synthesis logic from the lexer's benchmark is moved over here and the lexer uses APIs in the source generator for identifiers. The other source synthesis in the lexer's benchmark isn't yet moved over, but should likely be slowly absorbed here as it can be refactored into a more principled and re-usable form. Some bits may stay of course if they're just too lexer-specific. Next, this adds a simple end-to-end compile benchmark for the driver that directly and much more clearly reproduces all the measurements I've done manually up until now. It should also be easy to extend to more patterns over time as we add support to the source generator to produce those patterns. Last but not least, I've added a tiny CLI to the source generator so that you can generate source code manually. This is especially nice for generating demo source code to actually run through the driver or look at in an editor. The CLI can also generate C++ source code which lets us do some minimal comparative benchmarking between Carbon and C++/Clang. There are huge number of TODOs in the source generation framework. This is going to be a large ongoing effort I suspect. There are also a bunch of rough edges I've left to try and get this out for review sooner. I've left TODOs for refactorings that really need to be done here, but hoping these can maybe be follow-ups. If not, please flag and I'll try to layer them on here. Sample compile benchmark output, nicely showing where we are w.r.t. our goal speeds (2x behind on lex and check, 5x on parse) at least on a recent AMD server CPU: ``` ------------------------------------------------------------------------------------------------------ Benchmark Time CPU Iterations Lines ------------------------------------------------------------------------------------------------------ BM_CompileAPIFileDenseDecls<Phase::Lex>/256 29420 ns 29419 ns 22860 6.62847M/s BM_CompileAPIFileDenseDecls<Phase::Lex>/1024 146130 ns 146128 ns 4840 6.69959M/s BM_CompileAPIFileDenseDecls<Phase::Lex>/4096 601584 ns 601577 ns 1020 6.69573M/s BM_CompileAPIFileDenseDecls<Phase::Lex>/16384 2547578 ns 2547313 ns 280 6.404M/s BM_CompileAPIFileDenseDecls<Phase::Lex>/65536 10816591 ns 10816389 ns 80 6.05193M/s BM_CompileAPIFileDenseDecls<Phase::Lex>/262144 52191320 ns 52189828 ns 20 5.02261M/s BM_CompileAPIFileDenseDecls<Phase::Parse>/256 101706 ns 101698 ns 6900 1.91745M/s BM_CompileAPIFileDenseDecls<Phase::Parse>/1024 512161 ns 512162 ns 1380 1.9115M/s BM_CompileAPIFileDenseDecls<Phase::Parse>/4096 2078426 ns 2078430 ns 340 1.938M/s BM_CompileAPIFileDenseDecls<Phase::Parse>/16384 8795786 ns 8795583 ns 100 1.85468M/s BM_CompileAPIFileDenseDecls<Phase::Parse>/65536 35073596 ns 35072973 ns 20 1.86639M/s BM_CompileAPIFileDenseDecls<Phase::Parse>/262144 151100688 ns 151097370 ns 20 1.73483M/s BM_CompileAPIFileDenseDecls<Phase::Check>/256 957059 ns 957049 ns 740 203.751k/s BM_CompileAPIFileDenseDecls<Phase::Check>/1024 1956134 ns 1955985 ns 360 500.515k/s BM_CompileAPIFileDenseDecls<Phase::Check>/4096 5797864 ns 5797417 ns 120 694.792k/s BM_CompileAPIFileDenseDecls<Phase::Check>/16384 21219608 ns 21217584 ns 40 768.843k/s BM_CompileAPIFileDenseDecls<Phase::Check>/65536 96311116 ns 96302334 ns 20 679.734k/s BM_CompileAPIFileDenseDecls<Phase::Check>/262144 371637963 ns 371609964 ns 20 705.387k/s ``` Lest someone think this is *bad*, the fact that we're already within 2x of our rather audacious goals makes me quite happy. =D --------- Co-authored-by: Jon Ross-Perkins <jperkins@google.com> Co-authored-by: Richard Smith <richard@metafoo.co.uk>
248 lines
10 KiB
C++
248 lines
10 KiB
C++
// Part of the Carbon Language project, under the Apache License v2.0 with LLVM
|
|
// Exceptions. See /LICENSE for license information.
|
|
// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
|
|
|
|
#ifndef CARBON_TESTING_BASE_SOURCE_GEN_H_
|
|
#define CARBON_TESTING_BASE_SOURCE_GEN_H_
|
|
|
|
#include <string>
|
|
|
|
#include "absl/random/random.h"
|
|
#include "common/map.h"
|
|
#include "common/set.h"
|
|
#include "llvm/ADT/ArrayRef.h"
|
|
#include "llvm/ADT/StringRef.h"
|
|
#include "llvm/Support/Allocator.h"
|
|
|
|
namespace Carbon::Testing {
|
|
|
|
// Provides source code generation facilities.
|
|
//
|
|
// This class works to generate valid but random & meaningless source code in
|
|
// interesting patterns for benchmarking. It is very incomplete. A high level
|
|
// set of long-term goals:
|
|
//
|
|
// - Generate interesting patterns and structures of code that have emerged as
|
|
// toolchain performance bottlenecks in practice in C++ codebases.
|
|
// - Generate code that includes most Carbon language features (and whatever
|
|
// reasonable C++ analogs could be used for comparative purposes):
|
|
// - Functions
|
|
// - Classes with class functions, methods, and fields
|
|
// - Interfaces
|
|
// - Checked generics and templates
|
|
// - Nested and unnested impls
|
|
// - Nested classes
|
|
// - Inline and out-of-line function and method definitions
|
|
// - Imports and exports
|
|
// - API files and impl files.
|
|
// - Be random but deterministic. The goal is benchmarking and so while this
|
|
// code should strive for not producing trivially predictable patterns, it
|
|
// should also strive to be consistent and suitable for benchmarking. Wherever
|
|
// possible, it should permute the order and content without randomizing the
|
|
// total count, size, or complexity.
|
|
//
|
|
// Note that the default and primary generation target is interesting Carbon
|
|
// source code. We have a best-effort to alternatively generate comparable C++
|
|
// constructs to the Carbon ones for comparative benchmarking, but there is no
|
|
// goal to cover all the interesting C++ patterns we might want to benchmark,
|
|
// and we don't aim for perfectly synthesizing C++ analogs. We can always drop
|
|
// fidelity for the C++ code path if needed for simplicity.
|
|
//
|
|
// TODO: There are numerous places where we hard code a fixed quantity. Instead,
|
|
// we should build a rich but general system to easily encode a discrete
|
|
// distribution that is sampled. We have a specialized version of this for
|
|
// identifiers that should be generalized.
|
|
class SourceGen {
|
|
public:
|
|
enum class Language {
|
|
Carbon,
|
|
Cpp,
|
|
};
|
|
|
|
struct FunctionDeclParams {
|
|
// TODD: Arbitrary default, should switch to a distribution from data.
|
|
int max_params = 4;
|
|
};
|
|
|
|
struct MethodDeclParams {
|
|
// TODD: Arbitrary default, should switch to a distribution from data.
|
|
int max_params = 4;
|
|
};
|
|
|
|
// Parameters used to generate a class in a generated file.
|
|
//
|
|
// Currently, this uses a fixed number of each kind of declaration, with
|
|
// arbitrary defaults chosen. The defaults currently skew towards large
|
|
// classes with lots of nested declarations.
|
|
// TODO: Switch these to distributions based on data.
|
|
//
|
|
// TODO: Add support for generating definitions and parameters to control
|
|
// them.
|
|
struct ClassParams {
|
|
int public_function_decls = 4;
|
|
FunctionDeclParams public_function_decl_params = {.max_params = 8};
|
|
|
|
int public_method_decls = 10;
|
|
MethodDeclParams public_method_decl_params;
|
|
|
|
int private_function_decls = 2;
|
|
FunctionDeclParams private_function_decl_params = {.max_params = 6};
|
|
|
|
int private_method_decls = 8;
|
|
MethodDeclParams private_method_decl_params = {.max_params = 6};
|
|
|
|
int private_field_decls = 6;
|
|
};
|
|
|
|
// Parameters used to generate a file with dense declarations.
|
|
struct DenseDeclParams {
|
|
// TODO: Add more parameters to control generating top-level constructs
|
|
// other than class definitions.
|
|
|
|
// Parameters used when generating class definitions.
|
|
ClassParams class_params = {};
|
|
};
|
|
|
|
// Access a global instance of this type to generate Carbon code for
|
|
// benchmarks, tests, or other places where sharing a common instance is
|
|
// useful. Note that there is nothing thread safe about this instance or type.
|
|
static auto Global() -> SourceGen&;
|
|
|
|
// Construct a source generator for the provided language, by default Carbon.
|
|
explicit SourceGen(Language language = Language::Carbon);
|
|
|
|
// Generate an API file with dense classes containing function forward
|
|
// declarations.
|
|
//
|
|
// Accepts a number of `target_lines` for the resulting source code. This is a
|
|
// rough approximation used to scale all the other constructs up and down
|
|
// accordingly. For C++ source generation, we work to generate the same number
|
|
// of constructs as Carbon would for the given line count over keeping the
|
|
// actual line count close to the target.
|
|
//
|
|
// TODO: Currently, the formatting and line breaks of generating code are
|
|
// extremely rough still, and those are a large factor in adherence to
|
|
// `target_lines`. Long term, the goal is to get as close as we can to any
|
|
// automatically formatted code while still keeping the stability of
|
|
// benchmarking.
|
|
auto GenAPIFileDenseDecls(int target_lines, DenseDeclParams params)
|
|
-> std::string;
|
|
|
|
// Get some number of randomly shuffled identifiers.
|
|
//
|
|
// The identifiers start with a character [A-Za-z], other characters may also
|
|
// include [0-9_]. Both Carbon and C++ keywords are excluded along with any
|
|
// other non-identifier syntaxes that overlap to ensure all of these can be
|
|
// used as identifiers.
|
|
//
|
|
// The order will be different for each call to this function, but the
|
|
// specific identifiers may remain the same in order to reduce the cost of
|
|
// repeated calls. However, the sum of the identifier sizes returned is
|
|
// guaranteed to be the same for every call with the same number of
|
|
// identifiers so that benchmarking all of these identifiers has predictable
|
|
// and stable cost.
|
|
//
|
|
// Optionally, callers can request a minimum and maximum length. By default,
|
|
// the length distribution used across the identifiers will mirror the
|
|
// observed distribution of identifiers in C++ source code and our expectation
|
|
// of them in Carbon source code. The maximum length in this default
|
|
// distribution cannot be more than 64.
|
|
//
|
|
// Callers can request a uniform distribution across [min_length, max_length],
|
|
// and when it is requested there is no limit on `max_length`.
|
|
auto GetShuffledIdentifiers(int number, int min_length = 1,
|
|
int max_length = 64, bool uniform = false)
|
|
-> llvm::SmallVector<llvm::StringRef>;
|
|
|
|
// Same as `GetShuffledIdentifiers`, but ensures there are no collisions.
|
|
auto GetShuffledUniqueIdentifiers(int number, int min_length = 4,
|
|
int max_length = 64, bool uniform = false)
|
|
-> llvm::SmallVector<llvm::StringRef>;
|
|
|
|
// Returns a collection of un-shuffled identifiers, otherwise the same as
|
|
// `GetShuffledIdentifiers`.
|
|
//
|
|
// Usually, benchmarks should use the shuffled version. However, this is
|
|
// useful when there is already a post-processing step to shuffle things as it
|
|
// is *dramatically* more efficient, especially in debug builds.
|
|
auto GetIdentifiers(int number, int min_length = 1, int max_length = 64,
|
|
bool uniform = false)
|
|
-> llvm::SmallVector<llvm::StringRef>;
|
|
|
|
// Returns a collection of un-shuffled unique identifiers, otherwise the same
|
|
// as `GetShuffledUniqueIdentifiers`.
|
|
//
|
|
// Usually, benchmarks should use the shuffled version. However, this is
|
|
// useful when there is already a post-processing step to shuffle things.
|
|
auto GetUniqueIdentifiers(int number, int min_length = 1, int max_length = 64,
|
|
bool uniform = false)
|
|
-> llvm::SmallVector<llvm::StringRef>;
|
|
|
|
// Returns a shared collection of random identifiers of a specific length.
|
|
//
|
|
// For a single, exact length, we have an even cheaper routine to return
|
|
// access to a shared collection of identifiers. The order of these is a
|
|
// single fixed random order for a given execution. The returned array
|
|
// reference is only valid until the next call any method on this generator.
|
|
auto GetSingleLengthIdentifiers(int length, int number)
|
|
-> llvm::ArrayRef<llvm::StringRef>;
|
|
|
|
private:
|
|
// The shuffled state used to generate some number of classes.
|
|
//
|
|
// This state encodes all the shuffled entropy used for generating a number of
|
|
// class definitions. While generating definitions, the state here will be
|
|
// consumed until empty.
|
|
struct ClassGenState {
|
|
llvm::SmallVector<int> public_function_param_counts;
|
|
llvm::SmallVector<int> public_method_param_counts;
|
|
llvm::SmallVector<int> private_function_param_counts;
|
|
llvm::SmallVector<int> private_method_param_counts;
|
|
|
|
llvm::SmallVector<llvm::StringRef> class_names;
|
|
llvm::SmallVector<llvm::StringRef> member_names;
|
|
llvm::SmallVector<llvm::StringRef> param_names;
|
|
};
|
|
|
|
class UniqueIdentifierPopper;
|
|
friend UniqueIdentifierPopper;
|
|
|
|
using AppendFn = auto(int length, int number,
|
|
llvm::SmallVectorImpl<llvm::StringRef>& dest) -> void;
|
|
|
|
auto IsCpp() -> bool { return language_ == Language::Cpp; }
|
|
|
|
auto GenerateRandomIdentifier(llvm::MutableArrayRef<char> dest_storage)
|
|
-> void;
|
|
auto AppendUniqueIdentifiers(int length, int number,
|
|
llvm::SmallVectorImpl<llvm::StringRef>& dest)
|
|
-> void;
|
|
auto GetIdentifiersImpl(int number, int min_length, int max_length,
|
|
bool uniform, llvm::function_ref<AppendFn> append)
|
|
-> llvm::SmallVector<llvm::StringRef>;
|
|
|
|
auto GetShuffledInts(int number, int min, int max) -> llvm::SmallVector<int>;
|
|
|
|
auto GetClassGenState(int number, ClassParams params) -> ClassGenState;
|
|
|
|
auto GenerateFunctionDecl(llvm::StringRef name, bool is_private,
|
|
bool is_method, int param_count,
|
|
llvm::StringRef indent,
|
|
llvm::SmallVectorImpl<llvm::StringRef>& param_names,
|
|
llvm::raw_ostream& os) -> void;
|
|
auto GenerateClassDef(const ClassParams& params, ClassGenState& state,
|
|
llvm::raw_ostream& os) -> void;
|
|
|
|
absl::BitGen rng_;
|
|
llvm::BumpPtrAllocator storage_;
|
|
|
|
Map<int, llvm::SmallVector<llvm::StringRef>> identifiers_by_length_;
|
|
Map<int, std::pair<int, Set<llvm::StringRef>>> unique_identifiers_by_length_;
|
|
|
|
Language language_;
|
|
};
|
|
|
|
} // namespace Carbon::Testing
|
|
|
|
#endif // CARBON_TESTING_BASE_SOURCE_GEN_H_
|