Is this not doing what I think it's doing? ```df_p...
# general
a
Is this not doing what I think it's doing?
Copy code
df_part = df_img.into_partitions(20)
for part in df_part.iter_partitions():
    part.write_deltalake(table_path, mode="append", io_config=io_config)

# df_img.write_deltalake(table_path, mode="append", io_config=io_config)
does
.write_deltalake()
not work as expected on parts of the partitions? The process is still running, and I'm fairly certain some partitions have been processed, but I don't see any table being written/uploaded.
✅ 1
k
Hi Ammar,
write_deltalake
is exposed for the DataFrame type and not MicroPartition. In other words, you can call it on
df_part
or
df_img
directly instead of
part
. Would that work for you?
✅ 1
a
Yes, realized afterwards that my assumption was wrong, I thought the micropartitions were actual df's - but changed to `.repartition(n)`since then. I also saw that `.repartition(n)`is (much?) more expensive than
.into_partitions(n)
are there any benchmarks ? And thank you!
k
Hey Ammar, I don't have specific benchmarks and it depends highly on workload, but
into_partitions
and
repartition
both give a dataframe with
n
partitions, they just do it differently.
into_partitions
splits up/combines existing partitions whereas
repartition
actually globally shuffles all of the data into the new partition count. If you need your data to be partitioned by specific columns, use
repartition
and set the
partition_by
parameter, otherwise
into_partitions
is more performant!
👍 1