Hi folks, I want to read only certain columns in m...
# general
h
Hi folks, I want to read only certain columns in my parquet dataset. Is there a way to do this efficiently with read_parquet()?
j
Hey @Henry T! You can actually do this with a
.select
after the fact.
Copy code
df = daft.read_parquet(...)
df = df.select("c1", "c2", "c3")
df.show()
Since Daft is lazy, we do automatic pushdowns into your read 🙂
h
nice thanks, just want to confirm that with this, the other columns won’t be read right?
since i have large columns that i don’t need to read, i want to be more efficient/optimized
j
Yes correct. You can verify this by checking the plan:
Copy code
df.explain(True)
Take a look at your optimized plan — it should show that the Scan now has the appropriate columns pushed into it!
h
awesome thank you for the quick response