I ran the h2o groupby dataset queries with Polars, DataFusion, and Daft locally on the 100 million row dataset (3 GB Parquet file). Daft's performance looks good 😍
I couldn't get some of the queries to run. See
them here. Queries 3, 6, 7, 8, and 9 didn't run. It's possible some of them would run with the Python API.
Could be cool to get these all running and then add Daft to the
h2o benchmarks that are maintained by DuckDB now.