<https://github.com/Eventual-Inc/Daft/issues/2958>...
# daft-dev
a
https://github.com/Eventual-Inc/Daft/issues/2958 @jay can you explain this more to me on Slack / when I come into office?
👍 1
k
When i looked at the minhash rs code again i noticed that it also truncates to i32 so my hypothesis that the difference was coming from the truncation seems to not be the case.. the other thing i was thinking was, would the arrays default to max i32 values upon initialization or on error?
j
Could we also clarify — what is your goal here @Kyle? Are you trying to get to exact parity in terms of semantics with the bigcode codebase? That might be difficult as you’ve seen there are many subtleties in terms of implementation detail. Otherwise, are you trying to just generally increase the collision rate of deduplication? Perhaps there are better strategies for doing so.
Or one more question — are you perhaps noticing unexpected behavior in the minhash dedupe? E.g. matching hashes on documents that shouldn’t be matches. That could be indicative of something like:
would the arrays default to max i32 values upon initialization or on error?
k
The ideal I had in mind was to decompose minhash such that it could be run as a set of expressions instead of just minhash - then I could adjust the bytes I want to keep, or select the hashing algorithm The "parity" part is also a goal because I have yet to figure out why the deduplication is less effective now and that's something I need to justify if we were to switch systems. It's not so much parity as it is proving that it is better, but parity + faster would just be an easy explanation So far it's matching less than expected on daft but I'm not sure why. It could also be my understanding of how to use the resulting pairs that's not quite right..