Repository navigation
Fix lineage through LATERAL FLATTEN, reused derived aliases, predicate subqueries and star options - #75
Closed
IL-William wants to merge 8 commits into
Closed
Fix lineage through LATERAL FLATTEN, reused derived aliases, predicate subqueries and star options#75IL-William wants to merge 8 commits into
IL-William wants to merge 8 commits into
Conversation
The parser builds a chain of binary operators left-deep: `a || b || c`
is `(a || b) || c`. The expression walkers recursed into both operands
and counted every operator as one level of nesting, so a flat chain of
more than about 100 operators passed MAX_RECURSION_DEPTH. The walkers
then gave up at the guard, which sits at the leftmost end of the chain,
and the statement was reported as APPROXIMATE_LINEAGE.
Generated surrogate keys reach that length quickly, because every field
is wrapped and joined with a separator:
SELECT MD5(
COALESCE(CAST(field_000 AS VARCHAR), '') || '-' ||
COALESCE(CAST(field_001 AS VARCHAR), '') || '-' ||
...
COALESCE(CAST(field_119 AS VARCHAR), '')
) AS row_key
FROM events
Over 120 fields, row_key was derived from the last 49 only: field_000
to field_070 were missing from its lineage. A key over 50 fields
already loses its first operands.
Walk the left spine of a BinaryOp chain in a loop and visit its operands
left to right, each one level below the chain. The length of a chain no
longer counts as depth, while nesting in right operands, function
arguments and parentheses still does, so the guard keeps protecting the
stack. The four walkers that share MAX_RECURSION_DEPTH use it:
visit_expression_for_subqueries, collect_column_refs,
find_aggregate_function and collect_simple_identifiers. Operands are
still visited in source order.
Tests: the 120-field key above keeps every field and raises no
APPROXIMATE_LINEAGE; unit tests pin the operand order and check that
nesting on the right still trips the guard.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sqlparser gives TRIM, SUBSTRING, POSITION, CEIL and FLOOR, AT TIME ZONE,
IS [NOT] DISTINCT FROM, the `col:path` and `col[i]` accessors, COLLATE,
OVERLAY, SIMILAR TO, RLIKE, ANY and ALL, and a few more, their own `Expr`
variants rather than function calls. collect_column_refs matched none of
them and ended in a catch-all arm, so a column inside one of these forms
was never a source, and no issue said so:
SELECT o.id, TRIM(c.name) AS customer_name
FROM orders o
JOIN customers c ON o.customer_id = c.id
customer_name came out with no source column. In
`SUBSTR(code, 1, 2) || suffix` only suffix was read, and
`WHERE TRIM(c.status) = 'active'` was not attached to customers.
collect_simple_identifiers, which drives lateral column alias
resolution, already walked SUBSTRING, CEIL and POSITION but not TRIM, so
a key hashed from earlier aliases of the same SELECT list lost every
source:
SELECT
MD5(UPPER(TRIM(CAST(order_id AS VARCHAR)))) AS order_key,
MD5(UPPER(TRIM(CAST(customer_id AS VARCHAR)))) AS customer_key,
MD5(CONCAT(UPPER(TRIM(CAST(order_key AS VARCHAR))), '||',
UPPER(TRIM(CAST(customer_key AS VARCHAR))))) AS link_key
FROM orders
Make the match in collect_column_refs exhaustive. Every form descends
into the operands it evaluates in the enclosing scope, including named
arguments written `name => value`, the tested value of
`x IN (SELECT ...)`, subscripts and bracket keys. Field names in a path
step and lambda parameters are not read, and subqueries keep their own
scope. The parser keeps `o.items[1]` as the root `o` followed by the
steps `.items` and `[1]`, so the leading dot steps are folded back into
the name: the column read is `o.items`, as without the subscript, and
not a column `o` of orders. With no catch-all arm left, a variant added
to sqlparser is a compile error rather than a column dropped from
lineage.
collect_simple_identifiers walks the same forms, so an alias hidden in
one of them is replaced by its sources instead of being resolved as a
column of a table in scope.
The postgres_array_slicing snapshot changes accordingly: `a[:]`,
`b[:1]`, `c[2:]` and `d[2:3]` now derive from the columns a, b, c and d
rather than from the table node.
Tests: the queries above, SUBSTR, CEIL, POSITION and `:` operands in
Snowflake, the left operand of IN (subquery), a DuckDB lambda and a
subscript of a qualified column in Postgres at the analysis level; a
table of dedicated forms pins what collect_column_refs reads, and checks
that extract_simple_identifiers sees the same bare names.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sqlparser gives `LATERAL FLATTEN(...)` its own table factor,
`TableFactor::Function`, which the lineage visitor did not match. The
alias was never registered and no issue was raised, so a column read
through it lost its source:
SELECT o.order_id, f.value::STRING AS tag
FROM orders o,
LATERAL FLATTEN(input => o.tags) f
`f` was taken for a table name: tag came out derived from a column
`value` of a table `F` that does not exist, the implied schema listed
`F` beside ORDERS, and orders.tags was read neither for lineage nor for
the implied schema. An unqualified `value` was reported as ambiguous
across the tables in scope, and a FLATTEN over the value of another had
nothing to follow.
Register FLATTEN under its alias, or under its name when it has none,
as a relation with the columns Snowflake documents: SEQ, KEY, PATH,
INDEX, VALUE and THIS. VALUE and THIS carry the element, so they derive
from the columns of the `INPUT =>` argument, or of the first positional
one, with the call as their expression. SEQ, KEY, PATH and INDEX say
where an element sits and read nothing. They are added without the
relation-level edge a source-less projection gets, which would have made
`'n' || f.index` depend on every base relation of the FROM clause. Each
FLATTEN gets its own node, keyed by its scope, since CTEs commonly reuse
one alias such as `f`. A star over the scope now includes the six
columns.
The alias is also recorded as a subquery alias, as a derived table's
is, so that what is read through it stays out of the implied schema,
and the qualified columns of the input are recorded there as projected
columns are: the query above now implies ORDERS(order_id, tags) and
nothing else. With column lineage disabled only that alias is recorded:
at table level the input is a column of a relation already in the FROM
clause, so a FLATTEN node would have no edge, and no column is emitted.
Those six columns are all a FLATTEN has, so the scope records it as a
relation with fixed columns. A bare name it lacks, with no schema to
place it, still resolves to the one other relation of the FROM clause,
as it did while FLATTEN went unregistered: `id` in `SELECT id FROM
orders o, LATERAL FLATTEN(input => o.tags) f` stays orders.id instead of
turning ambiguous, and in `FROM customers, LATERAL FLATTEN(input =>
addresses) a, LATERAL FLATTEN(input => phones) p` the second input is
read from customers. A warning for a name that is ambiguous between
tables no longer lists the FLATTEN beside them. VALUE and THIS read the
same input, so an input that cannot be resolved is reported once.
Any other `LATERAL <function>(...)` now raises UNSUPPORTED_SYNTAX, as
`TABLE(<function>(...))` already does, instead of being dropped
silently.
The snowflake_lateral_flatten snapshot changes accordingly: `value AS
p_id` now derives from b.cool_ids through the FLATTEN node, B's implied
columns gain cool_ids, and the ambiguity warning for `value` is gone.
Tests: the query above with a decoy `value` column on orders, its
implied schema without a schema given, chained flattens back to the
first input, the position columns reading nothing under a CTE joined to
another table, a literal input, one alias reused by two CTEs with named
and positional input, a star over FLATTEN, a bare column and a bare
filter beside a FLATTEN with no schema, a bare input to a second
FLATTEN, an ambiguous input reported once, SPLIT_TO_TABLE reported as
unsupported, and no FLATTEN node or column with column lineage
disabled.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A derived table's node was keyed by statement and alias. Two derived
tables of one statement sharing an alias in different scopes, as a
generated point in time query writes one `(...) AS v` per satellite,
therefore shared one node and one set of column nodes, and each consumer
received every producer's lineage:
WITH a_intervals AS (
SELECT k, ts FROM (SELECT DISTINCT k, ts FROM sat_a) AS v
),
b_intervals AS (
SELECT k, ts FROM (SELECT DISTINCT k, ts FROM sat_b) AS v
)
SELECT a.ts AS a_ts, b.ts AS b_ts
FROM a_intervals a JOIN b_intervals b ON a.k = b.k
a_ts came out derived from sat_a.ts and sat_b.ts both, and so did b_ts,
with no issue raised.
Count the derived tables of each alias per statement. The first keeps
the key it always had, so nothing changes for a statement that does not
repeat an alias, and each later one is keyed `alias#n`, getting a node
and column nodes of its own. The node keeps the alias as its label. Two
derived tables of one alias in the same scope are invalid SQL, and
lookups already go through the scope, so a per statement count is
enough.
A CTE name defined again in a nested WITH merges the same way, but CTEs
are also looked up by name once the statement has been walked, so that
takes more than a key and is left alone here.
Tests: the query above, each output from its own satellite and two
nodes labelled v; the same alias over twelve satellites, each output
from its own. Both fail without the change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… out of the outputs
A subquery a predicate reads was analysed with no target node, so its
select list fell back to whatever the enclosing select targets: the
statement's output at the top, and inside a CTE that CTE's column list,
which a star over it then returned.
SELECT k FROM t WHERE x >= (SELECT MAX(y) AS floor_y FROM u)
returned k and floor_y. The same held for IN, for EXISTS, whose
`SELECT 1` came back as col_0, and for a subquery in HAVING, a join
condition or a grouping expression. No issue said so.
Such a subquery is now analysed against a node of its own, labelled
`(subquery)` and keyed per occurrence, which owns its columns and their
lineage, and what it projects is taken back from the projection buffer
once it has been read. A scalar subquery in the select list is not
visited there and is returned as before.
A derived table without an alias, which Snowflake and others accept,
had no node at all: its projection hung on the enclosing target in the
same way, and the outer select found no table in scope and reported so.
It now gets a node like an aliased one, registered in its scope under a
name no query can write, `(derived)`, so that a star over it expands and
a bare column resolves.
Tests: a scalar comparison, IN, a correlated NOT EXISTS, HAVING and a
join condition each return only k; inside a CTE the subquery's column
stays on its own node and out of the star; a select list scalar
subquery is still returned; an unaliased derived table reads like an
aliased one, with no issue. All but the select list case fail without
the change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
expand_wildcard emitted every column of the relation whatever options
the wildcard carried: `SELECT * EXCLUDE (b) FROM p` returned b, and
`SELECT * RENAME (b AS bee) FROM p` returned b under its old name.
The common idiom of replacing a column fared worse:
SELECT * EXCLUDE (b), UPPER(b) AS b FROM p
The star copy of b and the computed b shared one output node, the copy
came first, and the UPPER(b) derivation was merged away: b came out as
a plain copy of p.b.
The wildcard's options now reach the expansion. A column named by
EXCLUDE, or by EXCEPT where the dialect writes it so, is not expanded; a
column named by RENAME is expanded under its new name, still read from
the old one. Names are compared as the expanded columns' own names are.
REPLACE and ILIKE are still not read. A plain `SELECT *, f(x) AS x`,
which does write two columns named x, is left as it was.
Tests: EXCLUDE with and without parentheses, qualified and inside a CTE;
EXCEPT; RENAME with and without parentheses, read from the old column;
the idiom above, whose b carries UPPER(b). All fail without the change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Builds on #74, whose three commits come first here until it lands: the tests
below use the
expected_sourceshelper it adds. The last five commits, fourfixes and their changelog entry, are this PR's.
Summary
Four ways a relation or a scope was misread, each giving wrong column lineage
without raising an issue.
LATERAL FLATTENwas not matched by the lineage visitor, so its alias wastaken for a table name: in
SELECT f.value::STRING AS tag FROM orders o, LATERAL FLATTEN(input => o.tags) f,tagcame from a columnvalueof a tableFthat does not exist, andorders.tagswas never read.generated point in time query writes one
(SELECT DISTINCT k, ts FROM sat_x) AS vper satellite, in separate CTEs, and every consumer then received everysatellite's lineage.
list fell back to whatever the enclosing select targets:
SELECT k FROM t WHERE x >= (SELECT MAX(y) AS floor_y FROM u)returnedkandfloor_y, anEXISTS (SELECT 1 ...)returnedcol_0, and inside a CTE the column joinedthe CTE's list, which a star over it then returned. A derived table without an
alias had no node at all, and the outer select found no table in scope.
* EXCLUDE,* EXCEPTand* RENAMEwere ignored in wildcard expansion.With the common idiom
SELECT * EXCLUDE (b), UPPER(b) AS b FROM p, the starcopy of
band the computedbshared one output node, the copy came first,and the
UPPER(b)derivation was lost.Changes
One commit per fix, each with its tests, after those of #74:
fix(core): read LATERAL FLATTEN as a relation with its own columnsregisters FLATTEN under its alias as a relation with the six columns
Snowflake documents; VALUE and THIS derive from the input, SEQ, KEY, PATH and
INDEX read nothing. Each FLATTEN gets a node keyed by its scope, since CTEs
commonly reuse one alias. Any other
LATERAL <function>(...)now raisesUNSUPPORTED_SYNTAX. Thesnowflake_lateral_flattensnapshot changesaccordingly.
fix(core): give each derived table its own node when an alias repeatscounts derived tables per alias and statement: the first keeps the key it
always had, so statements that do not repeat an alias are unchanged, and each
later one is keyed
alias#n.fix(core): keep a predicate's subquery and an unaliased derived table out of the outputsanalyses a subquery read by a predicate, a join condition or agrouping expression against a node of its own, labelled
(subquery), andtakes what it projects back from the projection buffer. A scalar subquery in
the select list is not visited there and is returned as before. A derived
table without an alias gets a node like an aliased one, registered in scope
under
(derived). This one builds on the previous commit's per occurrence key.fix(core): honour EXCLUDE, EXCEPT and RENAME in a wildcard's expansionpasses the wildcard's options to
expand_wildcard: excluded columns are notexpanded, renamed ones are expanded under their new name, read from the old
one. REPLACE and ILIKE are still not read, and a plain
SELECT *, f(x) AS xis left as it was.
A CTE name defined again in a nested WITH merges the same way the derived
tables did, but CTEs are also looked up by name after the statement has been
walked, so that takes more than a key and is not attempted here.
Each commit message has the details and examples. There is no API or schema
change.
Validation
cargo test --workspace --locked,cargo clippy --workspace --locked -- -D warnings,cargo fmt --all -- --checkand the schema guard pass on thisbranch.
pinning that a select list scalar subquery is still returned, which passes
either way by design.
On a large Snowflake dbt project: 105 of 111 FLATTEN value pairs had been
missing, 8 point in time models had their satellites crossed, 33 models
gained a column their SQL does not return, and 16 edges ran through columns an
EXCLUDEhad removed.The committed browser WASM is not rebuilt here, since CI builds its own.
🤖 Generated with Claude Code