Garrett Weaver
03/03/2025, 11:21 PMgreatest exist (not finding in docs) https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.functions.greatest.html?Garrett Weaver
03/03/2025, 11:33 PM@daft.udf(return_dtype=daft.DataType.float64())
def greatest(*cols: daft.Series) -> list[float | None]:
"""Gets the max value"""
results: list[float | None] = []
for values in zip(*[col.to_arrow() for col in cols]):
if all(not value.is_valid for value in values):
results.append(None)
continue
results.append(max(value.as_py() for value in values if value.is_valid))
return resultsColin Ho
03/04/2025, 6:32 PMimport daft
df = (
daft.from_pydict(
{
"x": [1, 2, 3],
"y": [4, 5, 6],
}
)
.with_column("greatest", daft.list_("x", "y").list.max())
.collect()
)
print(df)
Use daft.list_("x", "y") to zip the columns together, then use list.max() to do get the max per row