Ammar Bitar
12/06/2024, 7:36 AM@daft.udf(return_dtype=daft.DataType.string())
class UDF_S3FileHasher:
def __init__(self, endpoint, access, secret, hash_algo: str = "sha256", chunk_size: int = 8192) -> None:
self.hash_algo = hash_algo
self.chunk_size = chunk_size
self.s3_client = boto3.client(
's3',
endpoint_url=endpoint,
aws_access_key_id=access,
aws_secret_access_key=secret,
verify=False,
)
def __call__(self, paths) -> list:
def compute_hash(s3_path: str) -> str:
bucket, key = s3_path.replace("s3://", "").split("/", 1)
hasher = hashlib.new(self.hash_algo)
response = self.s3_client.get_object(Bucket=bucket, Key=key)
for chunk in response["Body"].iter_chunks(chunk_size=self.chunk_size):
hasher.update(chunk)
return hasher.hexdigest()
return [compute_hash(path) for path in paths.to_pylist()]
Now, I was wondering, instead of going directly through boto3, could I leverage daft (for instance maybe .url.download() )? (specifically, I'd be interested if we can also control the chunk-size)Kevin Wang
12/06/2024, 7:08 PMExpression.url.download, but maybe what you could do is use it to get your data into a binary column, and then use a UDF to split it into chunks and hash it.Kevin Wang
12/06/2024, 7:09 PMAmmar Bitar
12/06/2024, 11:30 PM