We should switch to using positional/ordinal column references for resolved columns in the near future. I want to really beef up our join selection/ordering optimizations once I get non-equi joins in, but many of those join optimizations will need some heavy refactors if we build them with name references and then switch to ordinals.
It will be a pretty hefty change though. Wondering if people had thoughts about it.
Kevin Wang
03/01/2025, 1:37 AM
We need ordinals for full SQL as well as Spark support, since they both allow duplicate column names. It will also allow us to support Susbstrait much easier, and make certain logic simpler such as joins.
r
Robert Howell
03/01/2025, 1:45 AM
this is the way
Robert Howell
03/01/2025, 1:48 AM
Do you think you could do this iteratively so that code dependent on field names need not change?
k
Kevin Wang
03/01/2025, 1:51 AM
Maybe we could. I'd like to maybe pair with someone to draft up a plan for doing this