Ok my Catalog API stuff is finally ready for a rev...
# daft-dev
j
Ok my Catalog API stuff is finally ready for a review https://github.com/Eventual-Inc/Daft/pull/3036 (cc @Cory Grinstead here) For other folks, feel free to let me know what you think of the API!
Copy code
import daft

###
# Register a PyIceberg Catalog
# TODO: we should detect this from a YAML or something
###

from pyiceberg.catalog import load_catalog
catalog = load_catalog(...)

daft.catalog.register_python_catalog(catalog)

###
# Adding named tables
###

df = daft.register_table(df, "foo")

###
# Reading tables
###

df1 = daft.read_table("foo")  # first priority is named tables
df2 = daft.read_table("x.y.z")  # next priority is the registered default catalog
df3 = daft.read_table("x.y.z", catalog_name="my_other_catalog")  # Supports named catalogs other than default one
p
In the sample code: “df = daft.register_table(df, “foo”) “, is the first parameter to the call register_table supposed to be catalog variable defined earlier?
j
So I was thinking
daft.register_table
just registers “named tables” and doesn’t do any table creation in the catalog. These are just temporary tables which can then be referenced from SQL or from
daft.read_table
There will be other APIs we will need to build for table creation in catalogs, probably a
df.write_table(…)
which will actually persist the table into a catalog!
p
I see. That makes sense. Thanks Jay for the clarification.