Commit Graph
299 Commits
Author SHA1 Message Date
Richard Smith 41e357492b Fix crash if no input file is given to carbon dump objcode (#2997)
If the final argument was `--target_triple=...`, we tried to read off
the end of the arguments list. Found by fuzzer.
2023-07-19 16:39:45 +00:00
Jon Ross-Perkins 9e149a21c1 Use lit to provide a temp for driver output files. (#2993)
Specifying output.o causes issues if it's not writable, which it may not
be in some test environments.
2023-07-17 21:56:13 +00:00
Jon Ross-Perkins 446b0ce4ae Refactor declaration name context logic to its own class. (#2989)
This started with cleaning up the remaining Name/expression type punning
in the node stack, and grew. I'm factoring out a class because we've
previously expressed the desire to factor logic out of SemanticsContext
where possible, and this seemed like a reasonable cut.

NameExpression as the first node as a QualifiedExpression allows the
qualifier handling to consider Name in one less spot, an incremental
simplification. However, the additional complexity caused by this makes
me split ApplyNameQualifier/ApplyExpressionQualifier in order to avoid
repeat checks of the parse node's kind. The logic is still largely
shared, thus a couple helper functions. I think this is all fairly well
structured in the isolated class.

I can see that we may want to avoid passing SemanticsContext as an
argument in the future if it elides a step of lookup.
2023-07-17 21:55:29 +00:00
Jon Ross-Perkins 43065a1257 Finish refactoring Push/Pop for stronger type handling. (#2987)
This adds a distinction between Unused and SoloParseNode, rather than
equating the two. This is intended to help identify nodes which are
getting pushed but maybe don't need to be.

Not totally done because I want to adjust declaration name handling due
to a quirk with how it mixes Name with Expression, but almost done. Once
that's done the type punning will be completely gone.
2023-07-14 23:16:16 +00:00
Jordan Rupprecht 6a2b9684fb Add a name to glob_lit_tests (#2988)
Adding `name` to `glob_lit_tests` will make it conform with other
implementations of `glob_lit_tests` out there. If someone uses this repo
while providing a different version of `glob_lit_tests` that requires a
name, those build rules will become invalid.
2023-07-14 20:22:57 +00:00
Jon Ross-Perkins 9751b4701d Start node stack push/pop setting IdT based on ParseNodeKind. (#2985)
I think there's more we can do here, but this seemed like a good
checkpoint to make sure the path I'm going down is roughly what you
expected. There's one actual edit in if expression structure to match
the increased enforcement.
2023-07-12 23:26:53 +00:00
037196f69f Creates object file from the module. (#2955)
This pr creates the object file for the carbon code that returns either
0 or 1. The command to convert the object code to binary is `clang
<object_file_name> -o <binary_name>`

---------

Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-07-07 22:16:41 +00:00
Richard Smith 123662f5b1 Basic support for non-defining declarations of functions. (#2977)
This doesn't support redeclaration, so there's no way to provide a
definition in Carbon for a forward-declared function yet.
2023-07-07 01:08:02 +00:00
Jon Ross-Perkins 239cdcc457 Finish splitting out semantics_handle.cpp (#2976)
I was running into some difficult merges in #2940, so I'm wanting to
finish splitting this last file to minimize the chance of future issues.
2023-07-06 22:56:36 +00:00
Jon Ross-Perkins bc84f109fe Rename semantics InvalidType to Error (#2975)
Following up on zygoloid's request [on
#2940](https://github.com/carbon-language/carbon-lang/pull/2940#discussion_r1253522564)
2023-07-06 22:30:56 +00:00
Jon Ross-Perkins 918c089e03 Add namespace support. (#2940)
This handles namespacing of functions. Parsing and semantics are changed
significantly, while lowering works without changes. Variables can't be
namespaced yet because they're dealing with patterns, and I didn't dig
through that code.

Most of the logic is done through the new name declaration stack, which
is necessary because semantics isn't quite sure where the declaration
name ends. It'd be complex for parsing to send a signal about this,
probably involving node variants and rewrites of the tree, and this
solution seems to work well. Unfortunately this means a new stack, but
that may be inevitable due to the extra information needing to be
tracked.

Note this doesn't deal with scoped lookups of non-namespace things,
which we'll need for generics. That'll probably involve pushing resolved
scopes onto a stack (or maybe just setting a singleton value?) to affect
contextual name lookup. But, I think the basics are there to make it
work when we can test the behavior.

This renames "designator expression" to "qualified expression" and adds
"qualified declaration" in order to use terminology more consistent with
C++.

Namespaces will probably need to be considered for name mangling down
the line, but this still uses the basic name.
2023-07-06 20:43:40 +00:00
Richard Smith f992d4d960 Don't return SemanticsFunction by value. (#2968)
It contains a vector, so it's not cheap to copy. Return by const
reference instead.

Thanks to @fasiddique for spotting this!
2023-07-05 22:09:50 +00:00
Richard Smith cb16a1ffca Update toolchain keyword list to match design. (#2961)
- `xor` keyword is removed, with `^` used in its place.
- For consistency, also replaced unary `~` with unary `^`, as those
changes come from the same proposal.
- `type` keyword is added.
- To keep existing tests working, parsing support for `type` literal is
added too.
- `is` keyword is removed; we'd already added the `impls` keyword to the
list.
- Several other missing keywords added.
2023-06-29 04:11:00 +00:00
Richard Smith 288ad9f8e5 Split statement-specific parts of semantics_handle.cpp into separate files. (#2949)
`semantics_handle.cpp` is a little on the large side, and is going to grow as we add new nodes. Some of the statement-specific parts have already been split into their own files. Split out the remaining such parts.
2023-06-27 09:26:59 -07:00
Jon Ross-Perkins b2084ea15d Shift Parser from 'Identifier' to 'Name' naming (#2947)
This PR renames parse nodes on a Name/NameExpression taxonomy. NameExpressions occur in a name context. The difference is that in non-expression contexts it's useful to return the identifier / string ID for adding to name lookup, whereas in expression contexts it's useful to return the resolved node ID for consistency with other expressions.

In the code, I do note SelfValueName is returned in the expression context: I'd expect this to change, as `self` in `[self: Self]` versus `self.Foo()` will probably be best handled similarly to the above. That means that, in the proposed taxonomy, both `SelfValueName` and `SelfValueNameExpression` will exist in order to assist semantics.

To contrast choices:

Original | Current | [zygoloid suggestion](https://discord.com/channels/655572317891461132/655578254970716160/1121581663399464970) | [This PR](https://discord.com/channels/655572317891461132/655578254970716160/1121814551789318215)
--- | --- | --- | ---
DeclaredName/DesignatedName | Identifier | NameComponent | Name
NameReference | NameReference | NameReference | NameExpression
SelfValueIdentifier | SelfValueIdentifier | SelfValueReference | SelfValueName
SelfTypeIdentifier | SelfTypeIdentiifer | SelfTypeReference | SelfTypeNameExpression
2023-06-26 15:49:57 -07:00
Richard Smith 3a9bc01ec4 Flatten two terminator kind macros into one. (#2948)
As requested in #2942.
2023-06-26 15:44:46 -07:00
Richard Smith 4b69264cb1 Add implied return; at end of non-value-returning functions. (#2942)
Add validation that every code block in a function is terminated by a sequence of terminating instructions, and that terminators don't appear anywhere else in code blocks.

This required tracking whether we're in a reachable code block. That's done on the fly when we create a new code block; the new `SemanticsNodeBlockId::Unreachable` is used to represent the case where we're not actually creating a code block because we're in unreachable code.
2023-06-26 15:17:47 -07:00
Farzana Ahmed SiddiqueandFarzana Ahmed Siddique aad4ed2083 Codegen: Given a carbon file prints the assembly to the stdout (#2944)
Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
2023-06-23 15:23:39 -07:00
Richard Smith b908c6e274 Track the list of blocks that form the body of a function. (#2941)
Use that list for lowering in lexical order, instead of rediscovering
the list based on which blocks are referenced as branch targets.
2023-06-23 08:36:13 -07:00
Jon Ross-Perkins 5a90f660b9 Unify DeclaredName and DesignatedName as just Identifier (#2939)
This is just a simplification: I think these different forms are getting in the way more than they're helping, particularly as I was looking into namespace functionality. The handling in semantics can be identical, providing a more uniform behavior.
2023-06-22 16:20:06 -07:00
0b49ae32de Switched from raw_ostream to pwrite_stream in the driver.cpp (#2937)
Co-authored-by: Farzana Ahmed Siddique <fasiddique@google.com>
Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-06-22 13:55:46 -07:00
josh11b 78444fa7b1 Remove and rename keywords from toolchain to reflect #2760 (#2921)
- Renames `extends` -> `extend`
- Removes `external`
- Adds `require`

None of these seem to be in use by the parser yet, so no other changes are needed.
2023-06-19 08:37:03 -07:00
Richard Smith 1ea123cd94 Semantics IR building for if statements. (#2920)
This follows the same structure as `if` expressions, except that no
result value is needed.

Semantics IR building for code blocks is also added. Rather than popping
all the node stack entries we push for statements within a code block,
change statements and declarations to not push themselves onto the
stack. We're not notionally performing work recursively within prior
statements, and we don't need their value for anything, so it seems
cleaner to not push them. This also allows statements and declarations
to determine what syntactic context they're in by peeking at the top of
the stack, though that's not used in this patch.
2023-06-16 20:10:33 -07:00
Jon Ross-Perkins e2b1511a0d Clean up lit tests and config to reflect current uses. (#2913)
Overall, cleaning up remaining lit uses.

#2851 had removed FileCheck invocations from some of the explorer tests; this starts as just restoring that. But, now that we have far fewer `lit` tests, it seems best to refine `lit.cfg.py` to focus on providing fewer commands (not all were even used).

Also, adding testing of an error to explorer (which I noticed due to a change that would've broken that) made me notice that autoupdate_lit_test's for_lit logic didn't actually work, so I'm just cutting it and going to manual updates. Really, we might want to just remove lit autoupdate support altogether since it's only a couple tests using it, but I'm not ready to make that change right now.

Fixes #2912
2023-06-16 11:52:11 -07:00
Richard SmithandJon Ross-Perkins 0d4d392d12 Lowering of Branch / BranchIf / BranchWithArg / BlockArg. (#2904)
This gives us complete lowering of `if` expressions plus `and` and `or`.

Co-authored-by: Jon Ross-Perkins <jperkins@google.com>
2023-06-15 17:53:07 -07:00
Jon Ross-Perkins 87dbc4bd8a Fix missing forward_list include (#2911)
Noticed due to some build issues not caught by github actions.
2023-06-15 11:33:51 -07:00
Richard Smith f7f73d42ca Split out a per-function lowering context from LoweringContext. (#2901)
This is being done in preparation for adding more state to per-function
lowering, to handle lowering functions containing more than one block.
2023-06-14 15:27:49 -07:00
Richard Smith 70c6199496 Semantics handling for grouping parentheses. (#2900) 2023-06-14 14:30:19 -07:00
Richard Smith 06ce3b0161 Parsing, semantic analysis, and lowering for and, or, not. (#2897)
Lowering for `and` and `or` is not yet complete because `Branch` lowering isn't done yet.
2023-06-14 13:03:13 -07:00
Jon Ross-Perkins 8b1e820848 Migrate some lexer tests to file tests. (#2892)
These tests are doing string comparisons on output that don't seem to be meaningfully different from a file_test.

I'm tempted to migrate lexer tests in general, but I'm not doing that here since others may find more value in the current approach.
2023-06-14 09:50:48 -07:00
Jon Ross-Perkins d18c1347d7 Migrate compatible uses to TestRawOstream. (#2891)
Replacing direct raw_string_ostream uses. I figure the wrapper should be used more consistently.

There are still remaining raw_string_ostream uses that weren't compatible -- I'm continuing to look at those, but felt it was cleaner to have this on its own.
2023-06-14 09:31:56 -07:00
Richard Smith aa40e2b8a9 true and false support, and lowering for bool type. (#2896) 2023-06-13 16:43:03 -07:00
Jon Ross-Perkins 8e940d9724 Migrate //common test libraries to //testing/util. (#2890)
This is just a cleanup. Since we now have a testing directory, I think this is a better home for testonly libraries than //common. (I was thinking about this when I was considering adding more test_raw_ostream deps)
2023-06-13 16:38:08 -07:00
Richard Smith 202d3f5993 Semantic analysis for if expressions (#2893)
Add semantic analysis and semantics IR building for `if` expressions, and add the first parts of control flow handling to semantics IR. After discussion with @chandlerc, use [block arguments](https://en.wikipedia.org/wiki/Static_single-assignment_form#Block_arguments) to convey values from the two arms of the `if` to the result. For now, only a single block argument is supported, but we should revisit this as we explore more of the requirements of the Semantics IR form.

Functions can now contain multiple code blocks, so grab the entry block up-front instead of assuming the entry block will be at the top of the block stack when we reach the end of function emission.

Add trivial support for `bool` type literal, because without it we can't write testcases.
2023-06-13 14:38:27 -07:00
Jon Ross-Perkins 9606ce2127 Switch SemanticsIRTest to just use the driver. (#2889)
Now that the driver uses vfs, there's less reason for tests to do their own flow. Switch SemanticsIRTest to use the driver directly as an example simplification.
2023-06-12 15:25:45 -07:00
Jon Ross-PerkinsandRichard Smith a93e621488 Add vfs support to toolchain. (#2888)
This adds vfs support to the toolchain, allowing Driver to take in-memory inputs in tests. As a consequence, I'm simplifying SourceBuffer: rather than allowing tests to pass in their own memory buffer, I'm using InMemoryFileSystem to push for greater consistency with production code. This does hit a quirk where I need to be careful about null terminator handling because fuzzer imports don't always have one, but that's probably more robust anyways.

Co-authored-by: Richard Smith <richard@metafoo.co.uk>
2023-06-12 13:19:57 -07:00
Richard Smith 2c45cb3be8 Parsing support for if expressions. (#2883)
We model `if a then b else` as a prefix operator for parsing precedence purposes. The rule that a statement starting with `if` is never an `if` expression is handled implicitly because the statement parser never invokes the expression parser for a statement starting with `if`.

This exposed a bug in our diagnosis of the whitespace rule for prefix operators, which was incorrectly being applied to non-symbolic operators in some cases, and was producing a bogus second diagnostic in some cases, which is also fixed here.
2023-06-09 16:41:10 -07:00
Jon Ross-Perkins c43839e1b1 Switch FileTest to use StringRefs instead of files. (#2885)
In explorer, we already support parsing a string_view, so use that. In toolchain, we need to build support, probably using vfs, so that's a todo.

bazel test //explorer:file_test --runs_per_test=5

- branch: Stats over 250 runs: max = 18.3s, min = 5.2s, avg = 11.2s, dev = 2.9s
- trunk: Stats over 250 runs: max = 22.3s, min = 5.9s, avg = 12.1s, dev = 2.8s

Not a dramatic improvement, but maybe more effective long-term, and this'd been requested on #2876
2023-06-09 09:05:27 -07:00
Jon Ross-Perkins 00232846f8 Refactor toolchain tests so that driver API changes affect fewer files. (#2884)
This is really just about limiting the impact of changes, since I'm considering a driver API change to add vfs logic.
2023-06-08 16:09:54 -07:00
Jon Ross-Perkins 6586179c8d Add support for splitting a test file. (#2876)
I'm looking at this as I start thinking about handling `import`. Syntax is based on llvm's `split-file` tool.

The `std::vector` -> `llvm::SmallVector` switch is minor, I'm doing it here because I had to touch everything anyways and I think for tests I'll lean slightly more towards the toolchain's way of doing things versus explorer's.
2023-06-07 08:56:20 -07:00
Jon Ross-Perkins 39f7aae98d Sorting out how bindings are added to name lookup. (#2869)
This shifts logic so that bindings are added to name lookup only after the scope is complete, removing logic around adding/removing/re-adding names in certain scopes.

This does mean that things like a function's forward declaration will need to go through an extra hoop for name conflict checks, because under this approach a function definition does conflict checking when it adds names for the body's use. But, that seems easy to address, and better than the current hoops.
2023-06-01 16:50:34 -07:00
Jon Ross-Perkins a1a7251716 Refactor NodeStack APIs to better handle the increase of distinct ID types. (#2864)
I'm looking at adding CallableId, so figured I'd do this API refactoring which should reduce code duplication. This increases the amount of cross-calls between APIs because it may not be an issue for performance after optimizations, and should simplify reading of the API.

Also note the prior static_asserts on layout were missing a couple types, which is why I'm moving them into Entry where it's going to be more obvious when something's added.
2023-06-01 16:39:02 -07:00
Jon Ross-Perkins c83eea3ce2 Switch nodes to a map, and relabel as locals. (#2861)
The "locals" naming is intended to reflect that I haven't really thought through how globals work, but suspect they're going to end up at least partially separated. Stepping away from the "node" naming feels like it'll reduce confusion, although maybe "values" might be better? (but values seems imperfect in the presence of globals)

The map is the more important part here; I want to avoid anchoring on the prior vector approach.
2023-06-01 15:10:46 -07:00
Jon Ross-Perkins 4152cf720a Add loads for data. (#2860)
Start inserting loads when getting a value. This is done conditionally, but in theory we could start tracking nodes which need loading differently if the `isa` is considered cumbersome.
2023-06-01 14:44:57 -07:00
Jon Ross-Perkins db5629024e Drop "Lowered" from various LoweringContext APIs. (#2859)
As I've worked on this more, it's just felt verbosely redundant -- lowered should be the default assumption, needing no further explanation. This does mean there's `context.GetType()` and `context.semantics_ir().GetType()`, but I feel like that's still reasonably clear on reading.
2023-06-01 14:34:02 -07:00
Jon Ross-Perkins dce67e4062 Rename semantics to semantics_ir (#2858)
This is a rename of a commonly used accessor. I tend to prefer shorter names, but it's semantics_ir in LoweringContext, and it feels odder to shorten to semantics there than to lengthen to semantics_ir here. I already have builtins_ir around here, too. I think it's helpful to use consistent names for semantics_ir since they're dealing with essentially the same thing.
2023-06-01 14:32:22 -07:00
Jon Ross-Perkins 2e4beaf8f0 Canonicalize struct types. (#2855)
This adds canonicalization of struct types based on their type fields. It obsoletes the current CanImplicitAsStruct because the type ids should now be identical when they're structurally identical; there's only a reason to implicit CanImplicitAsStruct to detect _compatible_ conversions.

The type fields themselves aren't canonicalized because it would need to be done during the first parse, and could yield name conflicts being associated with the wrong location. i.e.:

```
var x: {a: i32, a: i32};
var y: {a: i32, b: i32, a: i32};
```

This should yield two separate name conflict diagnostics pointing at the type fields for each respective line, but if struct type fields were canonicalized then both would point at the first `a: i32` field definition. This isn't expected to be an issue for types because I'm trying to print those, but we may also end up with a "first defined at" situation in some cases (still, less confusing because the type should match). Regardless, I think individual fields gets much more awkward.
2023-05-26 14:42:45 -07:00
Jon Ross-Perkins 1497e1333d Switch types to a SemanticsTypeId. (#2854)
This switches types to using SemanticsTypeId instead of SemanticsNodeId, and lowering pre-builds its list of types. The empty tuple type is special-cased because we don't want to emit it unless it's in-use, but as the implicit return for functions, it's frequently used. Callables use invalid to indicate the implicit return, and that seems undesirable to change due to the size increase.
2023-05-26 11:19:49 -07:00
Jon Ross-Perkins 709412ca97 Remove the builtin empty tuple value (not type). (#2850)
We do use the empty tuple type for function returns, so I don't want to get rid of it, but the value is unused. With this change, builtins are all types. Long-term the empty tuple value should have a representation similar to structs; just a tuple value that's empty, not a built-in.
2023-05-26 11:07:35 -07:00
Jon Ross-Perkins 76151d1fff Stop special-casing empty structs. (#2849)
This removes special-casing of empty structs, handling them as just a regular value instead of a builtin. Note the `{} as Type` is still special-cased.

In lowering, removes the test of calling a function using `{}` because it's missing the proper load/store. This setup notices that error whereas the prior worked due to said special-casing. Fixing this will need to be done as part of generally adding loads for variable uses.
2023-05-26 10:33:17 -07:00