<https://github.com/Eventual-Inc/Daft/blob/43a351e...
# daft-dev
a
https://github.com/Eventual-Inc/Daft/blob/43a351e24d84f01e0448099a093e153adefc09e9[…]src/daft-connect/src/translation/logical_plan/local_relation.rs somehow I need to convert IPC
StreamReader
into a
LogicalPlanBuilder
any thoughts on how best to do this? @Sammy Sidhu @Colin Ho
c
so it looks like the relation already gives you
Vec<u8>
which you can convert to
Chunk<Box<dyn Array>>
via StreamReader, and then you can convert to
Series
and then to
Table
. You can make a new ScanOperator called IPC scan, https://github.com/Eventual-Inc/Daft/tree/main/src/daft-scan/src
a
I'm a little bit confused. Why would it be called an IPC scan, given that we would be passing in tables? Wouldn't it be more proper to just add a scan operator that can take in tables? Or are you thinking about something different? Also, I want to make sure this would work in a distributed environment. So would we actually want to integrate it w/ Python so that we can pass around the tables in Ray? @Colin Ho
@Kevin Wang @Desmond Cheong @Sammy Sidhu are either of you familiar in this area? Would appreciate if I could book like 15 minutes of your time today to decide how to best do this. btw related to https://github.com/Eventual-Inc/Daft/pull/3363/files#diff-b40ffd54503f8d381809e7431d239271b4753103358517855759da6f62050072
k
Based on what I'm seeing you can probably just convert that data into an arrow record batch and then somehow stick that into an in memory scan
a
yea ig the place I am getting stuck is how to do that
in practice
k
So it looks like the StreamReader can iterate out chunks of arrays, where each chunk holds one array for each column. you probably want to take those arrays, and I think we have some function to convert arrow2 arrays into Series. You can then combine those series into a daft Table for each chunk, and then combine those tables into a MicroPartition
I'm not sure the exact functions to do all of these but I am fairly certain they exist
relevant arrow2 docs in case you did not see them: https://docs.rs/arrow2/latest/arrow2/io/ipc/read/struct.StreamReader.html
a
ok yea I got that part ... creating the micropartition. I have done that already. I think the part that is confusing me https://github.com/Eventual-Inc/Daft/blob/f6eb9934c8a935928e6995e88261d0b8325d1631/src/daft-logical-plan/src/builder.rs#L676-L694 is how we want to convert Micropartition into python land because afaik it requires it being a pyobj
k
Hm, I know about as much as you on that part
a
aw ok well thanks anyway
c
This code here creates a dataframe from micropartitions: https://github.com/Eventual-Inc/Daft/blob/5dce4fb549a2430580176a1d75df4059adb53c74/daft/dataframe/dataframe.py#L480 This means it will 1. create the cache entry with the micropartitions, 2. create the logical plan builder via the in_memory_scan function with the cache entry. So you just need to extract the logical plan builder somehow into rust. You can call the function from rust. See https://github.com/Eventual-Inc/Daft/pull/2992/files#diff-03b0d338dafc8e570fbe105e9beabad30198dd3739899b3e9e8effa34176eafaR94-R105 for an example of rust code calling a python function that acceps micropartitions