Repository navigation
Improve substrait NameTracker so it doesn't require uuids #17508
Description
Activity
do we want to make it possible for two columns in the output schema to have the same exact qualified name? Postgresql does allow this:
SELECT data.a, CAST(data.a as string) from data;would yield 2 columns, both nameddata.a.This is not a great feature, and I understand forbidding it. But this PR aligns DF on the behavior of Posgres & spark
I would also like to match the behavior of Postgres and Spark where CAST does not appear in the field names in the schema
I think a first step towards a solution to this issue would either:
- making
CAST(B.C as Utf8)have qualified name ("B", "C") so that the name tracker can detect columnB.Chas an already seen name - OR making it so that unaliased CASTs do appear in the field name in the schema, something like
B.C::TEXT
- making
making CAST(B.C as Utf8) have qualified name ("B", "C")
IIRC two columns with the same qualified name and same field name, might result in an schema error as well during logical planning (see here).
OR making it so that unaliased CASTs do appear in the field name in the schema, something like B.C::TEXT
if we had
SELECT CAST(B.C as TEXT), CAST(B.C as TEXT) FROM tablecould it still conflict?
- IIUC the name tracker should go over it and we'd have B.C and B.C__temp__0, or am I not understanding properly? Is is a hard requirement, even before the substrait parsing? I'm not sure then why we'd accept a query like
SELECT data.a, CAST(data.a as string) from data;
SELECT B.C, CAST(B.C as TEXT) FROM table;we'd get columns "B.C" and "B.C"
- I think your example is a bit of a mix between my two proposals; I was thinking of making CASTs appear in the case where we still don't want duplicate col names, so something like
SELECT B.C, CAST(B.C as TEXT) FROM table;we'd get columns "B.C" and "B.C::TEXT"
Reacted by Lía Adriana- IIUC the name tracker should go over it and we'd have B.C and B.C__temp__0, or am I not understanding properly? Is is a hard requirement, even before the substrait parsing? I'm not sure then why we'd accept a query like
Sorry I think I misunderstood the issue, I graphed this to better picture this, please let me know if I'm understanding this wrong
SELECT B.C, CAST(B.C as TEXT) FROM table;yields to the following structure╔═══════════════════╦══════════════╦══════════════╦═══════════════╗ ║ Expression ║ Qualifier ║ qualified_ ║ schema_name() ║ ║ ║ .0 ║ name().1 ║ ║ ╠═══════════════════╬══════════════╬══════════════╬═══════════════╣ ║ Column(B.C) ║ Some("B") ║ "C" ║ "B.C" ║ ║ ║ ↑ ║ ↑ ║ ↑ ║ ║ ║ Has table! ║ Field only ║ Combined ║ ╠═══════════════════╬══════════════╬══════════════╬═══════════════╣ ║ CAST(B.C as Utf8) ║ None ║ "B.C" ║ "B.C" ║ ║ ║ ↑ ║ ↑ ║ ↑ ║ ║ ║ No table! ║ Full string ║ Same string ║ ╚═══════════════════╩══════════════╩══════════════╩═══════════════╝The null literal situation from the issue description yields to the following structure (without the uuid workaround):
╔═══════════════════════════╦═══════════════╦═══════════════╦═══════════════════╗ ║ Expression ║ Qualifier ║ qualified_ ║ schema_name() ║ ║ ║ .0 ║ name().1 ║ (name_for_alias) ║ ╠═══════════════════════════╬═══════════════╬═══════════════╬═══════════════════╣ ║ lit(NULL) ║ None ║ "UTF8(NULL)" ║ "UTF8(NULL)" ║ ║ (new literal in project) ║ ║ ║ ║ ╠═══════════════════════════╬═══════════════╬═══════════════╬═══════════════════╣ ║ Column(left.UTF8(NULL)) ║ Some("left") ║ "UTF8(NULL)" ║ "left.UTF8(NULL)" ║ ║ (from input after join) ║ ║ ║ ║ ╚═══════════════════════════╩═══════════════╩═══════════════╩═══════════════════╝The solutions proposed in the issue description -> using qualified_name().1 instead of name_for_alias() to detect when to rename would directly fix the null literal case wihtout needing to use uuids. But for the cast situation the name tracker would never trigger a rename since it sees a different qualified_name().1, and because the schema_name are the same they fail later in
validate_unique_namesiiuc we want to solve both situations while avoiding the uuid workaround.
do we want to make it possible for two columns in the output schema to have the same exact qualified name?
Unless I'm missing something, I don't see the harm in allowing it if Postgres already does it.
making CAST(B.C as Utf8) have qualified name ("B", "C") so that the name tracker can detect column B.C has an already seen name
You mean doing this for a fix where the name tracker tracks both schema_name and qualified_name().1 as the issue says? Do you know why right now the name tracker is not triggering a rename if both schema names are the same? (which iiuc is what
name_for_aliasreturns)🤔 I think if you alias the cast right now together with the Literal :
let maybe_apply_alias = match e { lit @ Expr::Literal(_, _) => lit.alias(uuid::Uuid::new_v4().to_string()), cast @Expr::Cast(_) => cast.alias(name), _ => e, };Even if this fixes the issue for the Cast, would it be possible that there are other expressins that fall under this same behaviour as well and would need aliasing also?
Do you know why right now the name tracker is not triggering a rename if both schema names are the same? (which iiuc is what name_for_alias returns)
Right now, the name tracker does! I was unclear, sorry! I meant for the suggested solution from @alamb where we use
qualified_name().1. Then, we'd need the cast to have the same qualified name as the column selection.I'm not sure the issue that the alias fixes? aliasing with the qualified name? As far as I could tell, casts seem to be the only edge case to address.
take
- added a commit that references this issue
on Feb 24, 2026 - added a commit that references this issue
on Feb 25, 2026 - added a commit that references this issue
on Feb 25, 2026 - added a commit that references this issue
on Feb 25, 2026 - added a commit that references this issue
on Feb 26, 2026 - added a commit that references this issue
on Mar 31, 2026
the following PR adds uuids to certain substrait identifiers to disambiguate them, but this may make the plans non reproducable. @Blizzara has some ideas how how we can avoid the UUIDs
FWIW, I looked a bit at what it'd take to fix the tracker. I think a core of the issue is that DF checks name ambiguity in two ways: there's the AmbiguousColumn exception you're running into, and then there is a
validate_unique_names()function which gets called on the creation of the Project. The former needs unique non-qualified names, while the latter needs unique schema names (which can be qualified).An easy fix for the former would be to change
name_for_alias()intoqualified_name()._1heredatafusion/datafusion/substrait/src/logical_plan/consumer/utils.rs
Line 398 in 1d9e138
CAST(B.C as Utf8)with a qualified name ([no qualifier], "B.C") and a schema name "B.C", as well as a reference to the original columnB.Cwith a qualified name ("B", "C") and also schema name "B.C". As the qualified name's name parts are different, it wouldn't be renamed (after the change I propose), and then it'd fail thevalidate_unique_names()check. So maybe for a proper fix, NameTracker would need to track both the schema name and the name-part of the qualified name, and rename until both are unique.(A simple example of the behavior of the CAST and validate_unique_names() is that
SELECT data.a, CAST(data.a as string) from data;also fails in datafusion-cli.)Originally posted by @Blizzara in #17299 (comment)