adrian lee
12/20/2024, 6:08 AMfrom pyiceberg.catalog import load_catalog
import pyarrow.parquet as pq
catalog = load_catalog('demo_catalog', \
**{'uri': '<https://xxxxxx.snowflakecomputing.com/polaris/api/catalog>', \
'warehouse': 'demo_catalog', \
'credential': 'sasasasas', \
'scope':'PRINCIPAL_ROLE:ALL', \
'client.region':'us-west-2'})
table = catalog.load_table("taxi.taxi_dataset")
df = daft.read_iceberg(table)
However, Im trying to write a new row to the iceberg table through the catalog but cant seem to be able to do so
df_write = daft.sql("select * from df limit 1")
written_df = df_write.write_iceberg(df, mode="append")
Got an error stating OSError: When initiating multiple part upload for key 'taxi/taxi_dataset/data/ee044039-1820-4918-b58d-a7361e5ec333-0.parquet' in bucket 'adrian-iceberg-us-demo': AWS Error UNKNOWN (HTTP status 301) during CreateMultipartUpload operation: Unable to parse ExceptionName: PermanentRedirect Message: The bucket you are attempting to access must be addressed using the specified endpoint. Please send all future requests to this endpoint.
I tried doing as what we discussed on linkedin and tried doing
written_df = df_write.write_iceberg(table, mode="append", io_config=<http://daft.io|daft.io>.IOConfig(s3=<http://daft.io|daft.io>.S3Config(region_name="us-west-2")))
but got an error TypeError: got an unexpected keyword argument 'io_config'jay
12/20/2024, 6:39 AMdf.write_deltalake but not yet for df.write_iceberg
Do you mind making an issue for the team? We’d love to knock it out quicklyadrian lee
12/20/2024, 6:42 AMadrian lee
12/20/2024, 6:59 AMadrian lee
12/20/2024, 8:48 AMjay
12/20/2024, 8:49 AM<http://df.to|df.to>_arrow_iter() would return you an iterator of PyArrow record batches, which you can then use pyiceberg to write directlyjay
12/20/2024, 8:50 AMadrian lee
12/20/2024, 9:27 AMdf_write_pyarrow = <http://df_write.to|df_write.to>_arrow()
table.append(
df=df_write_pyarrow
)jay
12/20/2024, 9:37 AMKevin Wang
12/20/2024, 8:51 PMtable = catalog.load_table("taxi.taxi_dataset") , could you give me the keys of the dictionary _table_.io.properties? My suspicion is that the client.region is set but we get the region from s3.region instead. So a quick workaround, if I am correct, is to set s3.region when you call load_catalog to "us-west-2", and I will also get a fix in to pick the region info up from client.region as welljay
12/20/2024, 11:24 PMKevin Wang
12/21/2024, 7:25 AMjay
12/23/2024, 9:25 AMHi @adrian lee, for the PyIceberg table you fetched withWe think that Polaris might be behaving differently from other catalogs, could you give me the keys of the dictionarytable = catalog.load_table("taxi.taxi_dataset")? My suspicion is that the_table_.io.propertiesis set but we get the region fromclient.regioninstead. So a quick workaround, if I am correct, is to sets3.regionwhen you calls3.regionto “us-west-2", and I will also get a fix in to pick the region info up fromload_catalogas well (edited)client.region
adrian lee
12/23/2024, 9:29 AM{'default-base-location': '<s3://adrian-iceberg-us-demo/>', 'uri': '<https://xxxxxxxx.snowflakecomputing.com/polaris/api/catalog>', 'warehouse': 'demo_catalog', 'credential': 'wxxxxxxxxx', 'scope': 'PRINCIPAL_ROLE:ALL', 'client.region': 'us-west-2', 'token': 'ver:1-hint:45969035153414-xxxxxxxx', 'prefix': 'demo_catalog', 'write.parquet.compression-codec': 'zstd', 's3.access-key-id': 'xxxxxxxx', 's3.secret-access-key': 'xxxxxx', 's3.session-token': 'xxxxxxx', 'expiration-time': '1734949724000'}Kevin Wang
12/23/2024, 8:10 PMKevin Wang
12/23/2024, 8:12 PM