Scott Routledge
03/13/2025, 5:38 PMScott Routledge
03/13/2025, 5:39 PMCory Grinstead
03/13/2025, 6:41 PMdef get_time_bucket(t):
bucket = "other"
if t in (8, 9, 10):
bucket = "morning"
elif t in (11, 12, 13, 14, 15):
bucket = "midday"
elif t in (16, 17, 18):
bucket = "afternoon"
elif t in (19, 20, 21):
bucket = "evening"
return bucket
IMO, the easiest way would be to use a daft.sql_expr here.
daft.sql_expr("""
CASE
WHEN hour IN (8, 9, 10) THEN 'morning'
WHEN hour IN (11, 12, 13, 14, 15) THEN 'midday'
WHEN hour IN (16, 17, 18) THEN 'afternoon'
WHEN hour IN (19, 20, 21) THEN 'evening'
ELSE 'other'
END
""")Scott Routledge
03/13/2025, 7:40 PMsql_expr API but that's a great suggestion. Although part of the goal of this benchmark was to stay in Python as much as possible and utilize UDFs, this does make that part of the workload quite a bit faster, so it would be an interesting comparison.Cory Grinstead
03/13/2025, 9:12 PMif_else expression if you wanted it to be pure python.
The sql_expr provided above gets parsed into several if_else statements under the hood, so they are logically equivalent. I just think the syntax for sql is more intuitive for this specific kind of operation.Robert Howell
03/14/2025, 7:17 PM