Hi. Is there any way to create/get a unique row in...
# general
c
Hi. Is there any way to create/get a unique row index for a dataframe. I want to load static read only parquet files etc and add new columns by creating new parquet files which can be joined onto the original dataframe. I see the docs discuss not needing row indexes for operations (when compared to Pandas) but in this case I want to augment a read only set of files. Thx.
k
Hi @Clive Cox, if you need unique row indices, we do have a currently non-public API for that in the form of
DataFrame._add_monotonically_increasing_id()
, which will generate a column of unique IDs. Would that suit your use case?
c
Maybe, but if the parquet files I am loading are read-only is there a way to save this column separately and join back on to dataframe as needed?
k
Hm, may I ask what is your use case for persisting this column?
c
We'd want to create other columns from ML operations and other scoring operations to join back to the original dataset. Those scores might be for all the dataset rows or subsets. If the dataset has a unique index or a hashable set of columns we can use that. Just wondered if there was anyway to still do this when its more of a pure data frame with no primary key as it were. Ideally, maybe a unique reference such as parquet-file URL and row index which is consistent.
k
Ah so if I am understanding correctly, you are creating new columns that you would like to write out, and you want to maintain the row mapping without also copying the data from the original columns
c
yes that's right
k
@Colin Ho do you know if monotonically increasing ID yields the same ID every time, if it's just read parquet + add monotonically increasing id? If so, Clive could probably just call that every time after the read to insert it
c
yes it will yield the same id
Ideally, maybe a unique reference such as parquet-file URL and row index which is consistent.
Also, you can try the
file_path_column
parameter on read_parquet, e.g.
daft.read_parquet(url, file_path_column='path')
which will automatically include the source path(s) as a column
c
Thanks for the suggestions