Anyone know why the iterator for `arrow2::array::B...
# daft-dev
r
Anyone know why the iterator for
arrow2::array::BinaryArray<i64>
produces an iterator over
Option<&[u8]>
? I would have expected it to produce an iterator over `Option<i64>`s...
k
A binary array stores variable length binary values which can't always be small enough to store in an i64
❤️ 1
r
So I can assume that each
&[u8]
subslice inside of each element is of the same length, correct?
k
could you elaborate on what you mean?
r
So each element inside of the iterator is an
Option<&[u8]>
, right. Thus, we should observe that:
Copy code
let first: Option<&[u8]> = binary_array.next().unwrap();
let second: Option<&[u8]> = binary_array.next().unwrap();

first.unwrap().len() == second.unwrap().len()
(And thus similarly for all other elements)
k
So that is not true for a BinaryArray. The lengths of each slice is defined by the offset array stored in the binary array.
In a FixedSizeBinaryArray you can assume that each row has the same length
It's similar to a ListArray/FixedSizeListArray if you are familiar with that
r
Oh. So then in a BinaryArray, the minimum size required to represent a number is used? And then in a FixedSizeBinaryArray, a set size is used to represent a number.
k
A binary array usually isn't used to denote a number if that's what you're asking. It's usually used as a way to store arbitrary objects that are serialized into binary values. For example you could binary serialize rust structs and store them into a binary array. Since the serialized length of each struct instance may vary depending on its value, you would use a variable size binary array
c
For further explaination. the arrow spec has 2 different
BinaryArray
implementations • Binary • LargeBinary in
arrow2
, the generic parameter is the offset, so it's really mutually exclusive to either
i32 | i64
Binary
is equivalent to
BinaryArray<i32>
and
LargeBinary
is
BinaryArray<i64>
. The parameterized type really is just an optimization. I believe we mostly use the
LargeBinary
implementation, but in some scenario's where you know you wont exceed the limitations of
BinaryArray<i32>
, (approx 2gb per item), then it could be a good choice as it has a smaller footprint than the i64 offset TL;DR; It's very conceptually similar to a
Bytes
type/class, and the
i64
represents the offset, not the inner data type.