`df.iter_rows` on a dataframe with nested dtypes l...
# daft-dev
c
df.iter_rows
on a dataframe with nested dtypes like
list
or
tensor
can be slow because Daft converts the columns into python lists (of python objects), and then iterate over each of the lists, see issue: https://github.com/Eventual-Inc/Daft/issues/3634. It is much faster to preserve the nested dtypes as arrow, or convert to a NumPy view. What do yall think letting the user choose how Daft should deal with nested datatypes during iteration? For example we can choose to preserve as arrow,
iter_rows
over a df with list column would yield tuples in the form of
{"list": <pyarrow.LargeListScalar: [1, 2, 3]>, ...}
instead of
{"list": [1, 2, 3], ...}
👍 1
j
Hmm what would be the API?
c
Copy code
def iter_rows(
        self, results_buffer_size: Union[Optional[int], Literal["num_cpus"]] = "num_cpus", column_format: Union[Literal["python"], Literal["arrow"], Literal["numpy"]] = "python"
    ) -> Iterator[Dict[str, Any]]:
I'm thinking along the lines of this
j
Ok yeah could be good. I think also what we can do is add typing overloads (not sure if I have the syntax right here, but it should look something like this)
Copy code
@overloads(column_format: Literal["python"]) -> Iterator[Dict[str, Any]]
@overloads(column_format: Literal["arrow"]) -> Iterator[Dict[str, pa.Array]]
@overloads(column_format: Literal["numpy"]) -> Iterator[Dict[str, np.ndarray]]
c
I think the return type should still be
Iterator[Dict[str, Any]]
, lets say we had a simple int column, the output item will be
Iterator[Dict[str, int]]
if format is
python
or
numpy
, and
Iterator[Dict[str, pa.Int64Scalar]]
if format is
arrow
j
Gotcha, yes that makes sense