Colin Ho
01/14/2025, 1:32 AMdf.iter_rows on a dataframe with nested dtypes like list or tensor can be slow because Daft converts the columns into python lists (of python objects), and then iterate over each of the lists, see issue: https://github.com/Eventual-Inc/Daft/issues/3634. It is much faster to preserve the nested dtypes as arrow, or convert to a NumPy view.
What do yall think letting the user choose how Daft should deal with nested datatypes during iteration?
For example we can choose to preserve as arrow, iter_rows over a df with list column would yield tuples in the form of {"list": <pyarrow.LargeListScalar: [1, 2, 3]>, ...} instead of {"list": [1, 2, 3], ...}jay
01/14/2025, 8:16 AMColin Ho
01/14/2025, 5:48 PMdef iter_rows(
self, results_buffer_size: Union[Optional[int], Literal["num_cpus"]] = "num_cpus", column_format: Union[Literal["python"], Literal["arrow"], Literal["numpy"]] = "python"
) -> Iterator[Dict[str, Any]]:
I'm thinking along the lines of thisjay
01/14/2025, 7:50 PM@overloads(column_format: Literal["python"]) -> Iterator[Dict[str, Any]]
@overloads(column_format: Literal["arrow"]) -> Iterator[Dict[str, pa.Array]]
@overloads(column_format: Literal["numpy"]) -> Iterator[Dict[str, np.ndarray]]Colin Ho
01/14/2025, 9:44 PMIterator[Dict[str, Any]], lets say we had a simple int column, the output item will be Iterator[Dict[str, int]] if format is python or numpy, and Iterator[Dict[str, pa.Int64Scalar]] if format is arrowjay
01/16/2025, 1:04 AM