Repository navigation
Migrate parser to the new span combining scheme #126763
Description
Activity
- addedneeds-triageThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triagingThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triaging
on Jun 20, 2024 - addedA-diagnosticsArea: Messages for errors, warnings, and lintsArea: Messages for errors, warnings, and lintsA-parserArea: The lexing & parsing of Rust source code to an ASTArea: The lexing & parsing of Rust source code to an ASTE-mentorCall for participation: This issue has a mentor. Use #t-compiler/help on Zulip for discussion.Call for participation: This issue has a mentor. Use #t-compiler/help on Zulip for discussion.A-macrosArea: All kinds of macros (custom derive, macro_rules!, proc macros, ..)Area: All kinds of macros (custom derive, macro_rules!, proc macros, ..)T-compilerRelevant to the compiler team, which will review and decide on the PR/issue.Relevant to the compiler team, which will review and decide on the PR/issue.
on Jun 20, 2024 It may also be possible to introduce a smarter
n-ary span combining operationto(a1, ..., aN)instead of the current binaryto, but it may be more expensive and will make spans better only in very rare cases.
So combining using a chain of binarytooperations should be fine for now.- addedC-cleanupCategory: PRs that clean code up or issues documenting cleanup.Category: PRs that clean code up or issues documenting cleanup.and removedneeds-triageThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triagingThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triaging
on Jun 20, 2024 either an individual token
The "token" here is a "token tree" really, despite Rust parser working mostly with flattened token sequences at the moment.
Sospan(#[word])should ideally bespan(#) to span([word])and notspan(#) to span([) to span(word) to(span(]).Also span combining for delimited groups can ignore all their internal tokens and only consider their delimiters.
Opening and closing delimiters always have the same context, and cannot be passed to macros separately in any way.
Sospan([word])should bespan([) to span(])and notspan([) to span(word) to span(]).@danik292 Yes
I suggest implementing this for some specific AST nodes that are not too complex, e.g. for paths maybe (
fn parse_pathinrustc_parse), or for something in types (compiler\rustc_parse\src\parser\ty.rs).
The current (and/or previous) node span may need to be kept instruct Parser.I'd also suggest starting with some testing infra, like a new attribute for showing spans in an AST node.
#[rustc_show_spans(expr)] fn foo() { let x = a + b; }
should show a diagnostic (probably a note) similar to this
#[rustc_show_spans(expr)] fn foo() { let x = a + b; ^ ^ ^^^^^ }See e.g.
rustc_effective_visibility(in./compiler) for an example of adding a new internal attribute.
The logic for showing spans should likely live in therustc_ast_passescrate.Is the main issue that something like
$a::$bwould have an incorrect span? I haven't fully understood the problem yet, but if the span is like a start, a length, and a context, how does the problem come from ignoring the middle tokens?
Summary:
I
a,b, ...,zare multiple consecutive span nodes that need to be combined into a single new span node, then its combined span should be"Span node" is either an individual token (every token has a span), or some larger AST node having a
Spanfield.Examples:
The combined span for these pieces of code will be
Status quo:
Currently the resulting span is typically built like
span_first_token.to(span_last_token).So anything in the middle and the internal structure (i.e. nodes, as opposed to tokens) are ignored.
Why we need to change it:
The
tooperation will automatically take macro variables into account, and will try to put the resulting span into the best suitable macro context (this was implemented in #119673).E.g. in
$tt + 5the combined expression span will be put into the context of the macro using$ttas a macro parameter.Note, that the same thing often happens in the current parser as well, but not consistently, e.g.
$a::$bwill produce an incorrect resulting path span because the::in the middle is not considered.This will also give us some single relatively well predictable rule for combining AST spans.
Implementation:
This work should be parallelizable relatively well (but may require a one time initial setup).
I'll review PRs doing this, they can be assigned to me.
It may be convenient to have a rolling value in the
Parserstructure for span of the current (or previous?) span node.The parser has a lot of bespoke diagnostic logic (including snapshotting) that stands in the way of any systematic improvements like this.
How this can be tested:
Make a macro that emits complex nodes using tokens from different contexts, e.g.
and emit some diagnostic using those nodes' spans (maybe can add a special internal diagnostic for this testing).