Yeah I’m guessing that’s some kind of really big data skew then. We need to figure out a way to maybe make smaller physical partitions/do some kind of streaming aggregation here potentially to avoid a naive materialization of all that data at once.
Let me try to reproduce something to see if I can get a big skew to give us a big resource request like this
I’ll have to chat with @Sammy Sidhu tomorrow too