<@U071FMQUR7W> I'm thinking of doing some work to ...
# daft-dev
k
@Cory Grinstead I'm thinking of doing some work to always parse expressions using our sql expr parser, which would enable string sql expressions everywhere that takes in an expression. this would include moving our expression resolving logic to the sql parser (things like parsing
a.b
to
col("a").struct.get("b")
and parsing
a.*
to a list of expressions). any thoughts about that?
👍 1
🤔 1
c
i think that could be interesting, but could also make for some weird/unwanted flexibility such as:
col("a as b")
when we'd likely want people to just use
sql_expr("a as b")
why not just keep
col
the way it is since we also have
sql_expr
k
Hm, imo that might be fine? If someone wants to do
col("a as b")
I would be okay with it. What are your thoughts? The main motivations for doing this is to 1. unify column name parsing for syntactic sugar such as struct getters and wildcards (right now we have something that is honestly kind of hard to understand on the non-sql side) 2. bring the sql expr integration to places other than .where
c
hmm, I'm leaning towards explicit usage as i think there's a lot of edge cases with using the sql parser for column expansion. Also, the sql parser relies on the column expr, (the rust implementation) so you may run into a circular dependency as well.
k
What kind of edge cases are you thinking about?
c
most sql dialects are case insensitive with columns, unless wrapped in quotes.
select NAME
is the same is
select name
, but different from
select "Name"
, but i'm wondering how that'd all work in dataframes as
col
is case sensitive. so
col("NAME")
,
col("Name")
and
col("name")
refer to 3 distinct columns. I suppose you could require the user to do
col('"Name"')
but that feels weird to me, and would likely introduce a ton of backwards incompatible changes.
also FWIW, our current sql implementation does not follow this convention and is case sensitive (the same as the dataframe api),
select NAME
and
select name
are currently treated as
"NAME"
and
"name"
when they both should be treated as
"name"
k
I see. That makes sense, thanks for the explanation
I noticed that pyspark by default is similar to SQL in that it is case insensitive. Would this be something that we should potentially change to in our dataframe API in a future minor version?
🤔 1
Would it make sense for the behavior to be consistent between the two APIs?
j
I did have a lengthy conversation about case sensitivity with Robert Howell (Amazon) Even within the SQL world this is apparently not standardized. I think we just have to take an opinionated stance and go with it there.
k
I think most engines have some global parameter that configures whether or not it's case sensitive. we can likely do the same
👍 1
The question I have is if that should be applied to dataframe ops as well