Hi again! I'm looking to compute the hashes of my ...
# general
a
Hi again! I'm looking to compute the hashes of my objects, currently doing as such:
Copy code
@daft.udf(return_dtype=daft.DataType.string())
class UDF_S3FileHasher:
    def __init__(self, endpoint, access, secret, hash_algo: str = "sha256", chunk_size: int = 8192) -> None:
        self.hash_algo = hash_algo
        self.chunk_size = chunk_size
        self.s3_client = boto3.client(
            's3',
            endpoint_url=endpoint,
            aws_access_key_id=access,
            aws_secret_access_key=secret,
            verify=False,
        )    

    def __call__(self, paths) -> list:
        def compute_hash(s3_path: str) -> str:
            bucket, key = s3_path.replace("s3://", "").split("/", 1)
            hasher = hashlib.new(self.hash_algo)
            response = self.s3_client.get_object(Bucket=bucket, Key=key)
            for chunk in response["Body"].iter_chunks(chunk_size=self.chunk_size):
                hasher.update(chunk)
            return hasher.hexdigest()

        return [compute_hash(path) for path in paths.to_pylist()]
Now, I was wondering, instead of going directly through boto3, could I leverage daft (for instance maybe
.url.download()
)? (specifically, I'd be interested if we can also control the chunk-size)
k
Hi @Ammar Bitar, we do not currently have a chunk size configuration in
Expression.url.download
, but maybe what you could do is use it to get your data into a binary column, and then use a UDF to split it into chunks and hash it.
That way, you can take advantage of Daft's built-in downloading capability and only use the UDF to compute hashes
a
I'll look into it, but I explicitly want to avoid storing the binaries in the table.