Hi, I am trying to read a subset of csv files stor...
# general
d
Hi, I am trying to read a subset of csv files stored under s3 prefix but can't seem to find a way to do it with daft. Can read a single object but not an array. Is it an expected behaviour?
k
Hey Denys, could you elaborate on what you are trying to accomplish?
d
Hi Kevin, trying to load a bunch of CSV files stored in S3 bucket for the further processing
I think I actually fixed the issue, I had a string interpolation in my call and it didn't work but after I removed it and left the list only it worked
source_df = daft.read_csv(f"{source_files}", has_headers=False)
replaced with
source_df = daft.read_csv(source_files, has_headers=False)
The interpolation was a leftover from the previous experimnet
k
Ok sounds good, let us know if you run into any issues then!
✅ 1
d
Actually, I do have a question. The underlying goal of my script is to do the comparison between 2 dataframes which represent the s3 inventory information in particular object name and the size. Basically confirm that s3 objects in one bucket are identical to the objects in the different bucket. Amount of objects can go into 10s of millions per folder
I was looking for a way to do comparison using daft itself without converting dataframe into pandas or alike but couldn't quite like anything like that
Can you suggest anything?
This is what I did in duckdb but it doesn't work in daft SQL
Copy code
WITH table_1(bucket_name, object_key, size) AS 
                (select  column00, column01, column05 from source_table
                ), table_2(bucket_name, object_key, size) AS
                (select  column00, column01, column05 from source_table
                )
                Select
                (SELECT sha256(list(table_1)::text) FROM table_1) = 
                (SELECT sha256(list(table_2)::text) FROM table_2) as is_identical;
k
So what you would like to accomplish is essentially checking if every row in each table is identical, right?
d
Correct - every row in one df should have corresponding identical row in the other
k
One way to do this would be to call
iter_rows
(docs here) on both tables and iterate through them together (using something like Python's zip function). This is not super efficient since it's executed in Python instead of via Daft but I don't think there's a good alternative way to compare two tables element-wise at the moment
Is there any information about the ordering or uniqueness of the data? We could perhaps use that to our advantage to do something more efficient
d
hm, yeah - I think it kind of defeats the purpose in this case
data is already ordered so just need to go line by line
k
Do you expect the (bucket_name, object_key, size) tuple to be unique for each row in a table?
d
yes
k
Oh in that case what you can do is a join on the two tables with those keys. If they all match you should have a new table that has the same length as the original two.
Copy code
len(table_1.join(table_2, on=["bucket_name", "object_key", "size"]).collect())
This can be done in a distributed manner by Daft
✅ 1
d
Good point, I will give it a go - thank you
👍 1