I’m seeing 3 other PRs that seem ready to merge/cl...
# daft-dev
j
I’m seeing 3 other PRs that seem ready to merge/close to the finish line for this release? • @Raunak Bhagat [BUG] Bump up max_header_size #3068 • @Desmond Cheong [FEAT] Support hive partitioned reads #3029 • @Andrew Gazelka [FEATURE] add min_hash alternate hashers #3052
r
I'm debugging the egress part of this. Have breakpoints throughout and looking through the writes on why I'm getting a
ArrowCapacityError
. The
pa.Table
's schema has
large_string
everywhere in it. No signs of regular
strings
. So I'm assuming everything is
i64
encoded already. Not sure why we're hitting the capacity error of 2gb.
Copy code
(Pdb) schema
repo: large_string
file: large_string
code: large_string
file_length: int64
avg_line_length: double
max_line_length: int64
extension_type: large_string
j
Could this be just a pyarrow datasets issue then?
r
Potentially. Going to talk to @Sammy Sidhu about this in a minute.
j
Ok, we can also merge a fix to the read parquet first, since writing parquet is a separate issue
r
The reason could be because of the translation back from Arrow to Parquet string encoding. Parquet encodes strings with a 4-byte prefix for the length. Arrow has a double-buffer mechanism in which the second buffer holds offsets. During ingress, we have to convert the lengths into offsets. This involves a summation operation. On the otherhand, during egress, we have to convert the summation back into offsets. My current hypothesis is that PyArrow probably has some check somewhere that if the sum is greater than (2^31 - 1), then it throws an error, even if the individual offsets themselves may be less than (2^31 - 1). I.e.:
Copy code
Parquet string encoding (first is length, second is the data):
[len_1, string_1]
...
[len_n, string_n]

E.g.:
[5, "Hello"]
[6, "world!"]
Copy code
Arrow string encoding:
[string_1...string_n]
[len_1, len_1 + len_2, ..., (len_1 + ... + len_n)]
So even if
len_i <= (2^31 - 1)
for
i := 1,...,n
, maybe PyArrow throws an error if
len_1 + ... + len_n >= (2^32 - 1)
?
But ya, let's merge this in since ingress is fixed.
d
Maybe we can cut tomorrow? Very possible to get both csv and hive in
@Sammy Sidhu and I banged out the last items for csv just now
a
🕺