list type column iteration/processing :thread:
# daft-dev
g
list type column iteration/processing 🧵
I have a dataframe that has a column containing a list of structs. i need to validate that the data in the list of structs is non null and discard the row that are. from the daft expression api for lists. i see a get() function. but have no way to iterate the list to access the structs. the lists per row can be of different sizes. my current thought is to explode the dataframe and then later re constitute with a group by + agg. This seems like excessive wasteful processing. Is there a better way?
r
qq, is your data like,
Copy code
col
--------------
[ {..}, {..} ]   # row 1.
[ ]              # row 2.
[ None, {..} ]   # row 3.
[ {..}, None ]   # row 4.
[ None ]         # row 2.
And you want to return?
Copy code
col
--------------
[ {..}, {..} ]   # row 1.
[ ]              # row 2.
g
its like this:
Copy code
col
__________________
[{message:"msg", "author": "ai"}]
[{message: "", "author": "ai"}]
[{message: null, "author": "ai"}]
and i want to return rows where the message is not null (and possibly not empty)
and these are arrays of many messages.. just 1 message in the example for illustration tho
c
so it sounds like you'd want a
list.filter
function? something like this?
Copy code
message_key = daft.element().struct.get("message")
df.select(
  daft.col("col").list.filter(
    ~message_key.is_null() and message_key != lit("")
  )
)
g
ah list.filter sounds right, hmm dint see this in the daft docs
c
we don't have a filter on the list functions yet, but we do want to add filter/map functions. In the short term, you will likely need to create a UDF to handle this filtering for you.
g
right ok.. ill try this w a udf for now.
thanks for the direction.
c
here's an example UDF that should work for your use case
Copy code
df = daft.from_pydict({
    "col": [
        {"message": None, "author": "ai"},
        {"message": "Hello", "author": "user"},
        {"message": "", "author": "user"}
    ]
})

@daft.udf(return_dtype = daft.DataType.bool())
def filter_col_by_message(col: daft.Series) -> daft.Series:

    return [item["message"] is not None and item["message"] != "" for item in col.to_pylist()]

df.filter(filter_col_by_message(df["col"])).collect()
👀 1
❤️ 1