Denys Tyshetskyy
01/09/2025, 9:38 PMKevin Wang
01/09/2025, 9:40 PMDenys Tyshetskyy
01/09/2025, 9:41 PMDenys Tyshetskyy
01/09/2025, 9:42 PMDenys Tyshetskyy
01/09/2025, 9:43 PMsource_df = daft.read_csv(f"{source_files}", has_headers=False)
replaced with
source_df = daft.read_csv(source_files, has_headers=False)
The interpolation was a leftover from the previous experimnetKevin Wang
01/09/2025, 9:44 PMDenys Tyshetskyy
01/09/2025, 9:49 PMDenys Tyshetskyy
01/09/2025, 9:49 PMDenys Tyshetskyy
01/09/2025, 9:50 PMDenys Tyshetskyy
01/09/2025, 9:50 PMWITH table_1(bucket_name, object_key, size) AS
(select column00, column01, column05 from source_table
), table_2(bucket_name, object_key, size) AS
(select column00, column01, column05 from source_table
)
Select
(SELECT sha256(list(table_1)::text) FROM table_1) =
(SELECT sha256(list(table_2)::text) FROM table_2) as is_identical;Kevin Wang
01/09/2025, 9:54 PMDenys Tyshetskyy
01/09/2025, 9:54 PMKevin Wang
01/09/2025, 10:01 PMiter_rows (docs here) on both tables and iterate through them together (using something like Python's zip function). This is not super efficient since it's executed in Python instead of via Daft but I don't think there's a good alternative way to compare two tables element-wise at the momentKevin Wang
01/09/2025, 10:01 PMDenys Tyshetskyy
01/09/2025, 10:02 PMDenys Tyshetskyy
01/09/2025, 10:03 PMKevin Wang
01/09/2025, 10:04 PMDenys Tyshetskyy
01/09/2025, 10:06 PMKevin Wang
01/09/2025, 10:08 PMlen(table_1.join(table_2, on=["bucket_name", "object_key", "size"]).collect())
This can be done in a distributed manner by DaftDenys Tyshetskyy
01/09/2025, 10:12 PM