Ammar Bitar
11/25/2024, 1:43 PMdf_part = df_img.into_partitions(20)
for part in df_part.iter_partitions():
part.write_deltalake(table_path, mode="append", io_config=io_config)
# df_img.write_deltalake(table_path, mode="append", io_config=io_config)
does .write_deltalake() not work as expected on parts of the partitions?
The process is still running, and I'm fairly certain some partitions have been processed, but I don't see any table being written/uploaded.Kevin Wang
11/25/2024, 6:51 PMwrite_deltalake is exposed for the DataFrame type and not MicroPartition. In other words, you can call it on df_part or df_img directly instead of part. Would that work for you?Ammar Bitar
11/26/2024, 7:27 AM.into_partitions(n) are there any benchmarks ?
And thank you!Kevin Wang
11/26/2024, 9:07 PMinto_partitions and repartition both give a dataframe with n partitions, they just do it differently. into_partitions splits up/combines existing partitions whereas repartition actually globally shuffles all of the data into the new partition count. If you need your data to be partitioned by specific columns, use repartition and set the partition_by parameter, otherwise into_partitions is more performant!