Hey Daft team and users of daft! What is the large...
# general
p
Hey Daft team and users of daft! What is the largest tabular dataset you have handled successfully with daft ? what kind of a job was it.
j
Hey there! We regularly work with datasets in the range of 1 - 30TB. Mostly Parquet-based. There are known issues with shuffles on large datasets (e.g. joins) especially once you have to go up to 1000*1000 partition shuffles. @Colin Ho actively works on shuffles and we have some really good solutions coming down the pipeline there. Other than that, it should be fairly stable. I know you've been facing some OOM issues -- we'd be happy to take a look if you're open to hopping on a call? Unfortunately without a closer look at your plan and execution it's difficult to debug 😬
👍 1