CR
01/30/2025, 12:11 AMCR
01/30/2025, 12:12 AMCR
01/30/2025, 12:12 AMCR
01/30/2025, 12:12 AMCR
01/30/2025, 12:12 AMCR
01/30/2025, 12:15 AMjay
01/30/2025, 12:18 AMDesmond Cheong
01/30/2025, 12:19 AM.count() I'm curious to compare the performance on a .collect() for daft vs spark. I feel like the performance difference is either in this aggregation or in data skipping on the scans, and .collect() will tell us which it isDesmond Cheong
01/30/2025, 12:20 AM.sum() on a single column will also be a pretty interesting aggregation to look atCR
01/30/2025, 12:34 AMcollect - about the same for each. to be be clear the same as their respective previous runs.
sum - same for spark and for daft it improved quite a bit. instead of 55s it was 28s (about hafl)jay
01/30/2025, 9:27 PMcount(lit(1)) since our engine should technically support micropartitions with no columns but that have a row size now โ @Desmond Cheong perhaps we could try to knock that out quickly?jay
01/30/2025, 9:29 PMDesmond Cheong
01/30/2025, 9:32 PMjay
01/30/2025, 9:38 PMjay
01/30/2025, 9:38 PMjay
01/30/2025, 9:39 PMCR
01/30/2025, 11:05 PMjay
01/30/2025, 11:10 PMDAFT_DEV_ENABLE_EXPLAIN_ANALYZE=1
This will dump a file that contains a little graph that we can then visualize!CR
01/30/2025, 11:20 PMjay
01/30/2025, 11:21 PMCR
01/30/2025, 11:22 PMCR
01/30/2025, 11:23 PMCR
01/30/2025, 11:25 PMColin Ho
01/30/2025, 11:26 PMjay
01/30/2025, 11:26 PMdbfs filesystem, but Daft is reading directly from S3.
The actual compute in this query is really fast (just a few ms)
We'll have to do some more experiments on our end to see if we can maybe special-case I/O operations when run from inside a databricks environment to hit dbfs.jay
01/30/2025, 11:27 PMsum(lit(1)) I think, which should work.CR
01/30/2025, 11:36 PMread_deltalake. Is there an issue with delta-rs not scanning/reading in parallel? I could see reading/scanning each parquet in sequence would be slow.jay
01/30/2025, 11:39 PMCR
01/31/2025, 12:27 AMCR
01/31/2025, 5:42 PMaccess_key in AzureConfig the performance increased significantly. I went from consistently 1m+/- runtimes to closer to 15s. Still slower than PySpark but huge improvement.CR
01/31/2025, 5:54 PMcollect() takes about 12s consistently for PySpark (vs 15s for Daft using access_key). @jay - I think your theory on caching (at least for subsequent runs has some merit).Desmond Cheong
01/31/2025, 5:57 PMspark.databricks.io.cache.enabledCR
01/31/2025, 6:00 PMCR
01/31/2025, 6:00 PMspark.conf.set("<http://spark.databricks.io|spark.databricks.io>.cache.enabled", False)
spark.catalog.clearCache()CR
01/31/2025, 6:01 PMCR
02/04/2025, 8:12 PMcol() ). If you have suggestions on how to improve performance on Databricks, I'm happy to give it a shot. Thanks.Desmond Cheong
02/04/2025, 8:15 PMIf you have suggestions on how to improve performance on DatabricksOne thing that we might be able to do on our side is take advantage of databrick's DBIO cache as well. Need to prototype it a little but I'll keep you in the loop!
jay
02/04/2025, 8:17 PMCR
02/04/2025, 8:21 PMintegration with Spark to give Daft a PySpark-compatible APIThat's interesting. Is there more info on this? A github issue or ?
jay
02/04/2025, 8:22 PMCR
02/04/2025, 8:23 PMCR
02/04/2025, 8:30 PMCory Grinstead
02/05/2025, 3:18 PMThat's interesting. Is there more info on this? A github issue or ?we have an issue label for spark related work. https://github.com/Eventual-Inc/Daft/issues?q=is%3Aissue%20state%3Aopen%20label%3Adaft-connect feel free to look at open/closed issues to get an idea of the current state of "daft-connect" (our spark connect implementation)
Cory Grinstead
02/05/2025, 3:19 PM