When writing (with write_parquet) to an S3 path, I...
# daft-dev
k
When writing (with write_parquet) to an S3 path, I get this error. How can I avoid this?
Copy code
OSError: When completing multiple part upload for key 'xxx/fb8794e8-d0e8-4324-a919-3a9761feacea-0.parquet' in bucket 'abc': AWS Error UNKNOWN (HTTP status 400) during CompleteMultipartUpload operation: Unable to parse ExceptionName: EntityTooSmall Message: Your proposed upload is smaller than the minimum allowed object size. Each part (except the last part) must meet the minimum part size requirement.
k
Hm is this a persistent issue? Does it have the same error if you run the query again?
j
Maybe a problem with the PyArrow writer… might be time to invest in our own
👍 1
k
Oh it is indeed not always happening. I also have a problem whereby the job fails with file not being found more often. There's one case it succeeded but took about 20 times as long as usual. It seems to be perhaps that the stability of my S3 service is not that good..
k
Huh. Are you using AWS or something else?
k
I think it's an s3 wrapper over an internal service 🫠
j
Ah… that would definitely explain it. We’ve found the stability of S3-based services to be extremely variable
Is this perhaps a workaround for Daft not having HDFS support?
k
Yeah it is
j
Ok @Sammy Sidhu looking at it
k
Nice thanks 😊😊😊