Ammar Bitar
11/26/2024, 3:23 PMdf_img = df_img.with_column(
'temp_results',
df_img['path']
.url.download(max_connections=8, io_config=io_config)
.image.decode()
.apply(computed_dict, return_dtype=daft.DataType.python())
)
will the images / binaries be flushed after the computation, or after each partition, or when can I assume the memory to be freed?Colin Ho
11/27/2024, 2:00 AMurl.download will be flushed once image.decode is completed, and the images from image.decode will be flushed once .apply is completed.Ammar Bitar
11/28/2024, 3:36 PMColin Ho
11/28/2024, 11:36 PMNeil Wadhvana
11/29/2024, 4:41 PMdf_img = df_img.with_column(
'temp_results',
df_img['path']
.url.download(max_connections=8, io_config=io_config)
.image.decode())
df_img = df_img.with_column('path', df_img['path'].apply(computed_dict, return_dtype=daft.DataType.python())
)Colin Ho
11/29/2024, 8:25 PMdf.explain(True)
Daft consolidates all transformations into a plan, and you can see each step when doing .explain(True) , example below. In a project step for example, daft will execute some number of expressions, i.e. image_decode(download(col(path))) as temp_results per partition. You can think of the execution of these expressions as like doing a depth first search. Retrieve the column, 'path', do the download, then decode.Colin Ho
11/29/2024, 8:27 PMurl_download is not used or reference any more in any other expression, it should be dropped.Neil Wadhvana
11/29/2024, 8:28 PMColin Ho
11/29/2024, 8:39 PM