As I build my unresolved/resolved column types I f...
# daft-dev
k
As I build my unresolved/resolved column types I feel an increasing need for a strict, typed boundary between the stages that that operate on unresolved columns vs the ones that operate on the resolved ones. I'm already running into bugs with my code from conflating the two and I am pretty worried that we will accidentally commit code that will do so in the future. We should perhaps have different types or generics for these two column types to get the compiler to catch these issues, instead of just sticking them into the same enum. However I'm not sure what a good approach for this is, especially considering we also use
col
to represent both in the Python side, and also because it may be a pretty significant refactor. Wondering if anyone has any thoughts about this.
c
wondering if we could take a similar approach to polars and have an expr IR?
k
could you elaborate on what polars does?
c
So polars actually has a few different stages of an expr. • Expr (similar to ours, but without any resolved columns or types) • ExprIR (an intermediate repr used during optimization) • PhysicalExpr (fully resolved expression)
FWIW, datafusion doesnt have the IR for optimization, but it does distinguish between a logical and a physical expr.
Just curious, are there specific operations that are causing issues? like is it wildcard expansion?
I also realize now that I misread your question. You were suggesting different variants, not different datatypes for unresolved/resolved. So disregard what I said about polars & datafusion.
k
Basically there is sort of an implicit boundary between when we have unresolved vs resolved columns. They should be unresolved until an optimization pass resolves them. So things like the sql planner, logical plan builder, and maybe unnest subqueries will used unresolved columns, but other optimizer rules as well as translation will all use resolved columns with my current design.
There are things like helper functions that are used in both sides of this boundary, for example
col
and
replace_columns_with_expressions
, and it's not immediately clear what side of the boundary they are in.
c
oh got it. In datafusion, this is distinguished by having a
Rewriter
and an
Optimizer
. It's assumed that all of the rewrite rules are applied before the optimizer. It's also assumed that the expr could be in an unresolved state before rewriting. This something that I wanted to add to our planner as well. A lot of the column resolution logic just happens wherever. But if it were all done within rewrite rules before optimization, I think it would make the boundary between resolved/unresolved a lot more explicit.
k
Yeah this is a similar thought to that. The thing I want more specifically is some sort of typed enforcement that certain rewriter rules (as well as everything before that) only uses unresolved columns, whereas the remaining rewrite rules, optimizer rules, and so-on only use resolved columns.
🤔 1
Alternatively, we do not have this concept at all in the logical plan, and implement resolution in each of the frontends separately. I am kind of leaning towards that tbh. It will prevent us from doing correlated subqueries in dataframe, but other engines don't allow that anyway.
c
i think that would give us a lot more control, but at the same time, it creates a much larger burden on developing the different frontends.
k
Yeah, that is why I wanted to move that into a common path. However it might honestly be less of a burden overall to do them individually
Let me try another approach actually, where the logical plan builder is actually the boundary between unresolved and resolved columns. That way the logical plan will only worked with resolved columns and the only place where you'd need to be wary of the differences is in the frontends, which will make it a little easier to check.
🙌 1
@Sammy Sidhu