Hi everyone! I have an usecase where I have large ...
# general
k
Hi everyone! I have an usecase where I have large data set like 500M rows in a table. I need to do row transformation basically need to take all columns and construct a JSON out of all columns data. The logic of creating JSON is complex. Today I use spark to do this. I create a Python UDF to construct the JSON and the whole thing runs in a cluster. It's an ETL pipleline I query source table convert the data into kind of JSON model using structs and arrays of structs then again persist in iceberg table. How can I approach something similar with Daft? Can someone gives me hint on how to solve this kind scenario? I want to minimize the requirements of memory do as much parallelism as possible.
j
So you want to convert the entire row into a single JSON (string) column?
k
technically it's not single JSON column but more or less completely different shape. The input columns and output columns are completely different. No one-to-one mapping all columns in transformations are compound and compute
c