bug: the lexer has no error channel, so unclosed strings, tab indentation and invalid characters parse clean #1
Labels
No labels
accessibility
bug
documentation
duplicate
enhancement
good first issue
help wanted
invalid
question
wontfix
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
barrettruth/starlark-cst#1
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Version
0.1.1
Dialect
Dialect::Bazel, reproduced onBUILDand.bzl. Not dialect-specific.Which guarantee broke
Neither of the two documented invariants: round-trip holds byte for byte and
nothing panics. This is the inverse of the template's "Parse error on a file
Bazel accepts" — no parse error on a file Bazel rejects. The bug report
template has no option for that class, and all three symptoms below are in it.
Observed
parse(..).errors()x = "hello\ny = 1\nunclosed string literalunfinished string literalx = "hello(EOF)unclosed string literalunfinished string literalx = """hello\nunclosed string literalunfinished string literaldef f():\n\tpass\nTab characters are not allowed for indentation. Use spaces instead.tabs are not allowedx = ?\nexpected an expressioninvalid character: '?'invalid input `?`A consumer cannot distinguish these from clean files. For a language server
that means no squiggle on an unterminated string — the single most common
transient state while typing a label.
Root cause
One structural gap, three symptoms: the lexer has no error output.
There is nowhere for a lexical diagnostic to go, so none is produced.
fn string(src/lexer.rs:287) terminates the literal onNone(
src/lexer.rs:297) or on an unescaped newline outside a triple quote(
src/lexer.rs:318), then unconditionally pushesSTRING/BYTES(
src/lexer.rs:327). Whether the closing quote was ever found is notrecorded.
ERROR_TOKENcarrying no message(
src/lexer.rs:275,src/lexer.rs:497). The parser then reports whatevergeneric syntactic expectation it happened to hold —
expected an expression— and the lexical fact is lost.Precedent
Bazel's
Lexer.javais the normative implementation, and it reports the errorwhile still emitting the token:
Four further sites do the same:
Lexer.java:294(backslash at EOF),Lexer.java:421(triple-quote EOF),Lexer.java:452(raw string newline),Lexer.java:501(raw string EOF). Tabs areLexer.java:213; invalidcharacters are
Lexer.java:874.That shape matters here: because Bazel keeps the token and only adds a
diagnostic, adopting it changes no byte range, so
to_string() == srcisunaffected. Both invariants survive.
LexerTest.javapins the range convention — the caret sits on the openingquote (
literalStartPos), not at the truncation point, and theSTRINGtoken still spans the truncated literal:
starlark-rust spans the whole truncated literal rather than a single point,
which is the better rendering of the same anchor:
Proposed behaviour
Give the lexer an error channel and keep the token stream exactly as it is
today. Messages verbatim from
Lexer.java, since those are the ones usershave already read from Bazel itself:
unclosed string literalr/bprefix, through the end of the truncated literal\tin leading indentationTab characters are not allowed for indentation. Use spaces instead.invalid character: '<c>'ERROR_TOKENwidthAnchoring on the opening quote rather than on the newline is the useful choice
and the one both reference implementations make: it points at the quote that
needs closing.
For
x = ?the lexical message should replace the parser'sexpected an expression, not stack with it — one byte of bad input is one diagnostic.Minimal reproduction
All four currently report zero errors while round-tripping correctly.
closed by referenced commit