Hello, I'm getting an error when using `.write_del...
# general
a
Hello, I'm getting an error when using
.write_deltalake()
:
daft.exceptions.DaftCoreException: DaftError::ValueError All images in a column must have the same dtype, but got: UInt16 and UInt8
Can it be that there's a type mismatch from on_error=null ?
Copy code
def computed_dict(images):
    computed_dict = {
        'height': im_height(images),
        'width' : im_width(images),
        'channels': im_channels(images),
        'other': im_other(images),
        'name': str(im_uuid(images))
    }
    return computed_dict


df_img = df.where(df['path'].str.endswith('.png') | df['path'].str.endswith('.jpg')) #.limit(5000)
key_to_dtype = {
    "height": daft.DataType.int16(),
    "width": daft.DataType.int16(),
    "channels": daft.DataType.int8(),
    "other": daft.DataType.float32(),
    "name": daft.DataType.string()
}

df_img = df_img.with_column(
    'temp_results',
    df_img['path']
    .url.download(max_connections=64, io_config=io_config, on_error='null')
    .image.decode(on_error='null')
    .apply(computed_dict, return_dtype=daft.DataType.python())
)


for key in ["height", "width", "channels", "other", "name"]:
    df_img = df_img.with_column(
        key,
        df_img['temp_results']
        .apply(lambda d, k=key: d.get(k), return_dtype=key_to_dtype[key])
    )
Appreciate any help! 🙂
c
Looks like the error is coming from
image.decode
, https://github.com/Eventual-Inc/Daft/blob/48284a89e8d405223ec82d8893f3ea3e99ebf3a6/src/daft-image/src/series.rs#L44 . Off the top of my head, I don't think it's because of
on_error='null'
, and i think it could be because you are combining pngs and jpgs in the same column
df_img = df.where(df['path'].str.endswith('.png') | df['path'].str.endswith('.jpg'))
, which could be decoded into different image buffer formats.
Try separating them out to different columns and running the image decode separately?
a
I already run this, and get 1 parquet file successfully written (which includes RGB/Grayscale images in both PNG/Jpg formats, ~1mil images) There seem to be some subset of images in another unsupported colortype. Also, I'm not actively storing the binaries nor the decoded images. But shouldn't `.image.decode(on_error='null')`allow me to just skip them instead of breaking the whole execution?
c
Yeah, looks currently
on_error
is only applied for the decoding part, and not when validating across all the decoded images. @jay do you think it makes sense to skip images with different dtypes if
on_error
is null?
🙌 1
j
I think we actually have an option to coerce them all into the same mode…
decode(mode="RGB")
This should ensure that everything will get decoded into the same mode. Would that help here?
✅ 1
a
Yes that's what I ended up using - I wasn't sure how it would interact for grayscale images though! Thanks 🙂
j
It converts the grayscale images to RGB on the fly
@Colin Ho one thing we could do here is give a better error message indicating to use the
mode
kwarg?
c
pikachu shocked today i learned haha
Sounds good, will add message