I have a question about partitioning/clustering sp...
# daft-dev
k
I have a question about partitioning/clustering specs that I'm wondering if anyone had any thoughts about. So in a cross join (which I am implementing), the result should technically have the clustering spec of both the left and right sides of the join. However, if they are clustered differently, there's no way of currently specifying a clustering spec that is both, say, hash and range clustered. This leads me to two questions: 1. is there a good choice of what side's clustering spec to use in this case? 2. should we modify our clustering specs to allow multiple clustering methods?
Ok after offline discussion with Sammy: • hash clustering assumes that for each hash there is only one partition that has it, so that won't be preserved in a cross join • range clustering is technically still valid in a cross join if you do the range clustered side as the outer loop of the cross product. however we think that some of our code actually assumes that range clustered partitions are sorted which they will no longer be • the only way then to preserve clustering with those assumptions is if one side has only one partition, in which case we can take the other side's partitioning scheme. I will do that.
🔥 1