Hi everyone! I have an usecase where I have large data set like 500M rows in a table. I need to do row transformation basically need to take all columns and construct a JSON out of all columns data. The logic of creating JSON is complex. Today I use spark to do this. I create a Python UDF to construct the JSON and the whole thing runs in a cluster.
It's an ETL pipleline I query source table convert the data into kind of JSON model using structs and arrays of structs then again persist in iceberg table.
How can I approach something similar with Daft? Can someone gives me hint on how to solve this kind scenario? I want to minimize the requirements of memory do as much parallelism as possible.