If you have partitions specified, what's the diffe...
# daft-dev
c
If you have partitions specified, what's the difference between
overwrite
and
overwrite-partitions
? ex:
Copy code
df.write_csv("data", write_mode="overwrite", partition_cols=["a"])
df.write_csv("data", write_mode="overwrite-partitions", partition_cols=["a"])
I'm trying to add
overwrite
to spark, but I broke a few tests & having a hard time understanding what the actual expected output is here.
c
overwrite will overwrite all files in the target dir. overwrite-partitions will only overwrite files in the partition dirs that were written to. Lets say you do a write with partition_col = 'a' and the values of 'a' in your data are [1,2,3]. With
overwrite
, all files in the target dir will be overwritten. With
overwrite-partitions
, only the files in directories
target_dir/partition_1/...
,
target_dirpartition_2/...
and
target_dir/partition_3/...
will be overwritten. If you had say
targert_dir/partition_4
,
target_dir/partition_5
, these would not be overwriten because the data did not contain partition values 4 and 5
💡 1