Yuri Gorokhov
10/25/2024, 7:25 PMwrite_parquet?
/arrow/cpp/src/arrow/filesystem/s3fs.cc:753: CompletedMultipartUpload got error embedded in a 200 OK response: SlowDown ("Please reduce your request rate."), retry = 1jay
10/25/2024, 7:28 PMYuri Gorokhov
10/25/2024, 7:30 PMYuri Gorokhov
10/25/2024, 7:30 PMYuri Gorokhov
10/25/2024, 7:31 PMmin_rows_per_group int, default 0
Minimum number of rows per group. When the value is greater than 0, the dataset writer will batch incoming data and only write the row groups to the disk when sufficient rows have accumulated.Yuri Gorokhov
10/25/2024, 7:42 PMYuri Gorokhov
10/25/2024, 7:43 PM.write_parquet(
tmp_doc_partition_path,
partition_cols=[partition_col, ])Yuri Gorokhov
10/25/2024, 7:43 PMYuri Gorokhov
10/25/2024, 7:43 PMjay
10/25/2024, 7:57 PMjay
10/25/2024, 7:57 PMKyle
10/26/2024, 12:30 AMjay
10/26/2024, 12:33 AM.read -> .write of the fragmented files, it should do pretty well in terms of reducing fragmentation!
Funny enough, in effect you’re using the filesystem/S3 here as effectively a shuffle service 😛Kyle
10/26/2024, 12:35 AMjay
10/26/2024, 12:40 AM_*parquet_target_filesize=*_512MBjay
10/26/2024, 12:41 AMscan_tasks_min_size_bytes/scan_tasks_max_size_bytes targeting to bring the data sizes to within 96/384MBKyle
10/26/2024, 12:50 AM