Neil Wadhvana
11/18/2024, 3:25 AMurl.download() expression. What would be the most efficient/daft-friendly way to perform an equivalent url.upload()? The current API only seems to support uploading to a directory/S3 prefix and doesn't allow setting up a column for desired output URLs.jay
11/18/2024, 6:47 PMimport daft
df = daft.from_pydict({"foo": [b"data1", b"data2", b"data3"]})
df = df.with_column("bar", df["foo"].url.upload("my_dir/"))
df.collect()Neil Wadhvana
11/18/2024, 6:49 PMjay
11/18/2024, 6:49 PMjay
11/18/2024, 6:50 PMjay
11/18/2024, 6:51 PMdf["foo"].url.upload(df["uploaded_urls"])Neil Wadhvana
11/18/2024, 6:51 PMjay
11/18/2024, 6:52 PMNeil Wadhvana
11/18/2024, 6:53 PMdownload() handle/report failures?jay
11/18/2024, 6:53 PMon_error="raise,null" flag which dictates behaviorjay
11/18/2024, 6:54 PMnull will set the result to null and log an errorNeil Wadhvana
11/18/2024, 6:55 PMjay
11/18/2024, 6:55 PMjay
11/18/2024, 6:56 PMdf = df.with_column("uploaded_url", df["foo"].url.upload(df["target_urls"]))jay
11/18/2024, 6:56 PMtarget_urls and the uploaded_urlsjay
11/18/2024, 6:56 PMNeil Wadhvana
11/18/2024, 6:56 PMnull seems like a meaningful output. That way, the user can still control the behavior of raise without it breaking API consistency with download()?jay
11/18/2024, 6:57 PMNeil Wadhvana
11/18/2024, 6:57 PMjay
11/18/2024, 6:57 PMjay
11/18/2024, 6:57 PMjay
11/18/2024, 6:57 PMNeil Wadhvana
11/18/2024, 7:02 PMNo simply because the rest of the year is busy for me, but I expect I may have some small contributions to offer next year to daft!jay
11/18/2024, 7:23 PM