Bumping this issue :grimacing: <https://github.com...
# daft-dev
k
Bumping this issue 😬 https://github.com/Eventual-Inc/Daft/issues/3008 Need a better way to fail gracefully when there are malformed inputs 🫠 Any ideas? Currently running this.. but it takes forever
Copy code
@daft.udf(return_dtype = daft.DataType.bool())
def check_path(path_string_series):
    path_strings = path_string_series.to_pylist()
    results = []
    for path in path_strings:
        if ".parquet" in path:
            try:
                df = daft.read_parquet(path).collect()
                del df
                results.append(True)
            except Exception as e:
                results.append(False)
                print_with_context(f"Failed to read {path} due to {e}")
        else:
            results.append(False)
    return results
j
Yeah sorry we’re coming out of a really busy week. We’ll take a look at this this week
k
Thank you!! Yes I noticed that the pypi releases were also flying off the press this week ☺️
❤️ 1
j
I wonder if we should add a new File type… Then you can do like
df[“path”].file_from_url().valid_parquet()
or something
k
I guess the important thing would be how to efficiently define and run valid_parquet.. If I UDF it and collect it, it seems a lot less efficient than how a single read_parquet would read them and spit the error out
👍 1
j
It would probably always be more efficient to ignore on-the-fly as we’re reading parquet
But perhaps more unexpected user behavior
k
Perhaps setting an execution config or read parquet param? Like the url.download config which has on_error
To make it even less likely to be abused it can be that there needs to be a log path for a choice of it not raising an error so that the actual inputs used/discarded can be logged somewhere
k
Hi @Kyle, the UDF you have is likely slow because it requires daft to fully materialize every parquet file. Here's an alternative using pyarrow that should only read the metadata, could you give it a try?
Copy code
import pyarrow
import pyarrow.parquet as pq

@daft.udf(return_dtype = daft.DataType.bool())
def check_path(path_string_series):
    path_strings = path_string_series.to_pylist()
    results = []
    for path in path_strings:
        if ".parquet" in path:
            try:
                parquet_file = pq.ParquetFile(path)
                results.append(True)
                del parquet_file
            except pyarrow.lib.ArrowIOError as e:
                results.append(False)
                print_with_context(f"Failed to read {path} due to {e}")
        else:
            results.append(False)
    return results
k
Works much faster now! 😄 Also found that I should split using into_partitions after the glob so that it doesn't just run the UDF with one cpu. Thanks @Kevin Wang!!
🙌 2
Used to take hours with no end haha 🫠 Now it finishes in seconds 😄 The job with an error has parquets throwing ArrowInvalid so I added that as well!
❤️ 3