I ran the h2o groupby dataset queries with Polars,...
# daft-dev
m
I ran the h2o groupby dataset queries with Polars, DataFusion, and Daft locally on the 100 million row dataset (3 GB Parquet file). Daft's performance looks good 😍 I couldn't get some of the queries to run. See them here. Queries 3, 6, 7, 8, and 9 didn't run. It's possible some of them would run with the Python API. Could be cool to get these all running and then add Daft to the h2o benchmarks that are maintained by DuckDB now.
❤️ 4
s
That’s amazing! @Matthew Powers! I am also curious if you try out the results using our new execution engine (Swordfish) that will be the default soon. On the SQL side, we cranking out the expressions, I believe @Kevin Wang and @Cory Grinstead are working on count distinct right now
m
Sweet, yea, will be happy to update these any time. I don't really like some parts of the h2o methodology, but including it in the official benchmarks would probably be a great way for Daft to get exposure. Once all the queries can be run, then I say we add it. I will try to figure out if some of these missing queries can be run with the Python API.
🔥 2