Our testing had started to take more noticable time so I looked at where the time went to see if it could be improved easily. Almost all of the time was spent building regex matchers. That code path allocates a lot of memory and does a lot of processing in order to make the regex fast to run over the input. While that's not really a great tradeoff for how we use these matchers, the bigger issue is that we don't need a regex is the vast majority of cases. This change inspects the pattern more deeply to avoid most of the work. First, it looks for the cases where the pattern doesn't require any adjustment at all and directly builds a matcher from that if it can. This avoids an extra string allocation entirely. This is probably only a minor improvement, but it is also very easy. The larger change is to process the string in two phases. First, we expand the keywords and check for a regex region. If we find no regex region, we can directly build a string equality matcher for the expanded string. This still allocates an extra copy of the string but is *dramatically* cheaper that building a regex. Finally, if we *do* find a regex, we re-process the string to escape everything and transform the regex sequence into its valid form before building the matcher. The result for me is an over 5x reduction in test time. Profiling afterward shows a lot more opportunities for optimization here if things get slow again. The actual toolchain code isn't really visible in the profile yet. Note, I was profiling the normal build, which has asserts and ASan and such. I've not looked at the optimized build.
file_test
BUILD
A typical BUILD target will look like:
load("rules.bzl", "file_test")
file_test(
name = "my_file_test",
srcs = ["my_file_test.cpp"],
tests = glob(["testdata/**"]),
deps = [
":my_lib",
"//testing/file_test:file_test_base",
"@googletest//:gtest",
"@llvm-project//llvm:Support",
],
)
Implementation
A typical implementation will look like:
#include "my_library.h"
#include "testing/file_test/file_test_base.h"
namespace Carbon::Testing {
namespace {
class MyFileTest : public FileTestBase {
public:
using FileTestBase::FileTestBase;
// Called as part of individual test executions.
auto Run(const llvm::SmallVector<llvm::StringRef>& test_args,
const llvm::SmallVector<TestFile>& test_files,
llvm::raw_pwrite_stream& stdout, llvm::raw_pwrite_stream& stderr)
-> ErrorOr<RunResult> override {
return MyFunctionality(test_args, stdout, stderr);
}
// Provides arguments which are used in tests that don't provide ARGS.
auto GetDefaultArgs() -> llvm::SmallVector<std::string> override {
return {"default_args", "%s"};
}
};
} // namespace
// Registers for the framework to construct the tests.
CARBON_FILE_TEST_FACTORY(MyFileTest);
} // namespace Carbon::Testing
Filename fail_ prefixes
When a run fails, information about what pieces failed are returned on
RunResult. This affects whether a fail_ prefix on the file is required,
including in combination with split-file tests (using the // --- <filename>
comment marker).
The main test file and any split-files must have a fail_ prefix if and only if
they have an associated error. An exception is that the main test file may omit
fail_ when it contains split-files that have a fail_ prefix.
Comment markers
Settings in files are provided in comments, similar to FileCheck syntax.
bazel run :file_test -- --autoupdate automatically constructs compatible
CHECK:STDOUT: and CHECK:STDERR: lines.
Supported comment markers are:
-
// AUTOUDPATE // NOAUTOUPDATEControls whether the checks in the file will be autoupdated if --autoupdate is passed. Exactly one of these two markers must be present. If the file uses splits, AUTOUPDATE must currently be before any splits.
When autoupdating, CHECKs will be inserted starting below AUTOUPDATE. When a CHECK has line information, autoupdate will try to insert the CHECK immediately next to the line it's associated with, with stderr CHECKs preceding the line and stdout CHECKs following the line. When that happens, any subsequent CHECK lines without line information, or that refer to lines appearing earlier, will immediately follow. As an exception, if no STDOUT check line refers to any line in the test, all STDOUT check lines are placed at the end of the file instead of immediately after AUTOUPDATE.
-
// ARGS: <arguments>Provides a space-separated list of arguments, which will be passed to RunWithFiles as test_args. These are intended for use by the command as arguments.
Supported replacements within arguments are:
-
%sReplaced with the list of files. Currently only allowed as a standalone argument, not a substring.
-
%tReplaced with
${TEST_TMPDIR}/temp_file.
ARGS can be specified at most once. If not provided, the FileTestBase child is responsible for providing default arguments.
-
-
// SET-CHECK-SUBSETBy default, all lines of output must have a CHECK match. Adding this as a option sets it so that non-matching lines are ignored. All provided CHECK:STDOUT: and CHECK:STDERR: lines must still have a match in output.
SET-CHECK-SUBSET can be specified at most once.
-
// --- <filename>By default, all file content is provided to the test as a single file in test_files. Using this marker allows the file to be split into multiple files which will all be passed to test_files.
Files are not created on disk; it's expected the child will create an InMemoryFilesystem if needed.
-
// CHECK:STDOUT: <output line> // CHECK:STDERR: <output line>These provide a match for output from the command. See
SET-CHECK-SUBSETfor how to change from full to subset matching of output.Output line matchers may contain
[[@LINE+offset]and{{regex}}syntaxes, similar toFileCheck.