Hey <@U042126MG49> something to check so like we w...
# general
a
Hey @jay something to check so like we were discussing on linkedin, Im reading my iceberg table from Polaris catalog. So far reading using pyicerberg is ok and then converting it to a daft dataframe is fairly straightforward.
from pyiceberg.catalog import load_catalog
import pyarrow.parquet as pq
catalog = load_catalog('demo_catalog', \
**{'uri': '<https://xxxxxx.snowflakecomputing.com/polaris/api/catalog>', \
'warehouse': 'demo_catalog', \
'credential': 'sasasasas', \
'scope':'PRINCIPAL_ROLE:ALL', \
'client.region':'us-west-2'})
table = catalog.load_table("taxi.taxi_dataset")
df = daft.read_iceberg(table)
However, Im trying to write a new row to the iceberg table through the catalog but cant seem to be able to do so
df_write = daft.sql("select * from df limit 1")
written_df = df_write.write_iceberg(df, mode="append")
Got an error stating OSError: When initiating multiple part upload for key 'taxi/taxi_dataset/data/ee044039-1820-4918-b58d-a7361e5ec333-0.parquet' in bucket 'adrian-iceberg-us-demo': AWS Error UNKNOWN (HTTP status 301) during CreateMultipartUpload operation: Unable to parse ExceptionName: PermanentRedirect Message: The bucket you are attempting to access must be addressed using the specified endpoint. Please send all future requests to this endpoint.
I tried doing as what we discussed on linkedin and tried doing
written_df = df_write.write_iceberg(table, mode="append", io_config=<http://daft.io|daft.io>.IOConfig(s3=<http://daft.io|daft.io>.S3Config(region_name="us-west-2")))
but got an error
TypeError: got an unexpected keyword argument 'io_config'
j
Oh indeed this seems to be missing. It seems we have this functionality for
df.write_deltalake
but not yet for
df.write_iceberg
Do you mind making an issue for the team? We’d love to knock it out quickly
a
sure let me create an issue on github
Is there a way instead I can convert the daft dataframe to pyiceberg and then save it back instead?
j
You can, and that would require exporting the data out of Daft. Something like:
<http://df.to|df.to>_arrow_iter()
would return you an iterator of PyArrow record batches, which you can then use pyiceberg to write directly
My guess is that pyiceberg is also using the pyarrow writers though, so you might need to configure it a little bit there
a
did it another way for now... works but if possible if we can call the method natively instead of doing the arrow conversions would b great!
df_write_pyarrow = <http://df_write.to|df_write.to>_arrow()
table.append(
df=df_write_pyarrow
)
j
Yup! We’ll get on it 🥳
k
Hi @adrian lee, for the PyIceberg table you fetched with
table = catalog.load_table("taxi.taxi_dataset")
, could you give me the keys of the dictionary
_table_.io.properties
? My suspicion is that the
client.region
is set but we get the region from
s3.region
instead. So a quick workaround, if I am correct, is to set
s3.region
when you call
load_catalog
to "us-west-2", and I will also get a fix in to pick the region info up from
client.region
as well
j
Oh yeah interesting and good catch — I wonder if this is failing because it’s a Polaris-specific config that we’re not catching!
k
Either way, I made a PR that implements both my suspected solution as well as a way to pass in a custom IO config, so this should solve the issue regardless. @jay could you take a look? https://github.com/Eventual-Inc/Daft/pull/3633
🔥 2
j
Gentle bump here @adrian lee :
Hi @adrian lee, for the PyIceberg table you fetched with
table = catalog.load_table("taxi.taxi_dataset")
, could you give me the keys of the dictionary
_table_.io.properties
? My suspicion is that the
client.region
is set but we get the region from
s3.region
instead. So a quick workaround, if I am correct, is to set
s3.region
when you call
load_catalog
to “us-west-2", and I will also get a fix in to pick the region info up from
client.region
as well (edited)
We think that Polaris might be behaving differently from other catalogs
a
Oh sorry for missing this.
Copy code
{'default-base-location': '<s3://adrian-iceberg-us-demo/>', 'uri': '<https://xxxxxxxx.snowflakecomputing.com/polaris/api/catalog>', 'warehouse': 'demo_catalog', 'credential': 'wxxxxxxxxx', 'scope': 'PRINCIPAL_ROLE:ALL', 'client.region': 'us-west-2', 'token': 'ver:1-hint:45969035153414-xxxxxxxx', 'prefix': 'demo_catalog', 'write.parquet.compression-codec': 'zstd', 's3.access-key-id': 'xxxxxxxx', 's3.secret-access-key': 'xxxxxx', 's3.session-token': 'xxxxxxx', 'expiration-time': '1734949724000'}
k
Yep that matches what I expect then. @jay could you give my PR a review? we’ll get a fix in our next release
👍 1
One compelling reason for switching to iceberg-rust I’m starting to realize is that we only need to make sure Daft works with one version of an iceberg client library 😄