<@U071FMQUR7W> we had a discussion in the past abo...
# daft-dev
k
@Cory Grinstead we had a discussion in the past about having all dataframe expression strings go through the sql planner. Something like if you have
df.where("a > 10")
be by default interpreted as sql. You made some good points about why not to do it at this time. Do you remember what they were?
🤔 1
c
I do remember that it can get confusing especially with select what happens if you have a column called
a as b
Copy code
df.select("a", "a as b")
the implicit coercion so sql can be confusing if you are unaware of the behavior. In my experience, explicit is usually better and less error prone in the long term. we added the
sql_expr
function that does this. Which makes it very clear to the reader of the code what exactly is happening.
df.where(sql_expr("a > 10"))
k
An idea I had recently is to try to find it as a column name first, and if it does not match any columns, to then parse it as a sql expr. What do you think about that?
c
I think there's a lot more edge cases to think about. another one is
df.select("a, b, c")
would we allow expr expansion inside a single string? IMO it adds a lot of complexity and diverging ways to write queries. I really don't want to end up like pandas where there are usually 10+ ways of doing every operation. What problem are we trying to solve by adding it in to the dataframe?
k
I mostly just want to get rid of the struct get syntactic sugar stuff we have right now and just use SQL behavior for that kind of thing
c
if people want to write sql, i'd prefer to add a
df.sql
method that could be used instead
Copy code
df.sql("select * from self where a > 10")
r
There are scenarios where the dataframe column resolution would not match the SQL resolution, since today dataframe column resolution is exact-cast whereas SQL column resolution is typically case-insensitive (it's more complicated than that though).
c
I mostly just want to get rid of the struct get syntactic sugar stuff we have right now and just use SQL behavior for that kind of thing
hmm. i wonder if there's a more python-esque way of doing struct access.
r
^ is there an example of what you'd like to do? It could be quite interesting to have a path operator and chain
__getitem__
to have a nice path dsl
k
Yeah so currently
col("a.b")
can be used as syntactic sugar for
col("a").struct.get("b")
. I'm honestly not opposed to removing that. We could maybe just allow
col("a")["b"]
🔥 1
r
An idea I had recently is to try to find it as a column name first, and if it does not match any columns, to then parse it as a sql expr
I was referring to this for the exact-case vs case-insensitive. If you failed to find the column for "A" when the column is named "a" – then parsing as an SQL expression will find it because "a" matches "A" in SQL. This might be surprising coming from python. I agree with Cory about explicit vs implicit – and keep SQL oddities contained to SQL.
Yeah Kev, that's exactly what I was thinking! Would be super cool to have a slick DSL via getitem
c
seems like polars does this via the
struct
namespace
col("my_struct").struct["b"]
https://docs.pola.rs/api/python/stable/reference/expressions/api/polars.Expr.struct.field.html#polars.Expr.struct.field i like the idea of just flattening it into
col
though.
col("my_struct")["b"]
k
We want to move away from namespaces right?
r
I believe it's just a method to enter the path DSL scope. I'm down for
.path
and some special chars for wildcards and unpivoting.
c
We want to move away from namespaces right?
yes i believe so.
r
But I think Kevin is onto something where we might actually be able to drop it altogether
c
you can even get fancy with it and have it work on lists too
col('list')[0:2]
-> list.slice(0,2)
col('list')[-1]
-> list.tail()
col('list')[0]
-> list.get(0)
🔥 1
k
that's true
Yeah ok why don't we do that instead of the string sugar that we have. I think we would still want to keep wildcard but the others all feel more Pythonic to do as
__getitem__
daft party 1
r
Copy code
col("a")."*"
col("a").["*"]
k
@Sammy Sidhu I think you were the one to suggest adding the struct get syntactic sugar. Do you have any input on this?