The ideal I had in mind was to decompose minhash such that it could be run as a set of expressions instead of just minhash - then I could adjust the bytes I want to keep, or select the hashing algorithm
The "parity" part is also a goal because I have yet to figure out why the deduplication is less effective now and that's something I need to justify if we were to switch systems. It's not so much parity as it is proving that it is better, but parity + faster would just be an easy explanation
So far it's matching less than expected on daft but I'm not sure why. It could also be my understanding of how to use the resulting pairs that's not quite right..