I'm working in <bedmap> to rewrite <PixPlot> using...
# general
f
I'm working in bedmap to rewrite PixPlot using Daft. This is a tool to create an embeddings visualization: • Validate a folder of images • Generate embeddings for each image • Generate a 2D layout from the embeddings • Build a nice web visualization The Daft API feels really nice for this. However, I hit OOM errors frequently. This notebook captures some of the struggles. The goal for this tool is that it runs on various hardware, with various pretrained
timm
models, without much thought from the user. Are there any tips for getting things to run memory efficiently?
❤️ 2
j
Super cool. I see you're running into some memory issues:
@Colin Ho might be able to help here. I believe maybe morsel sizing could be used here to help you do your workload in a more streaming fashion
I'm guessing we're downloading all 2000 images and decoding them at the same time
We're still working on making this more "automagic" -- i.e. figure out morsel sizes more effectively at each step so as to not overwhelm your system, but still get good parallelism.
👍 1
f
Thanks for the quick answer @jay ! You're correct, I finished writing my original question -- I'm messaging here about OOM issues. In this case, I've copied the images locally, but I think I could be loading them in memory in too big of chunks. I've experimented with morsel sizes as small as 50 -- I will try again with something very small like 5.
Update: I still get OOM with
morsel_size = 1
. It's strange to me that the memory issue is a function of number of rows (images) if
morsel_size
and/or
@udf(..., batch_size)
are doing what I would expect and doing this in a streaming fashion.
j
Hmm interesting. I see also that 2000 images was successful in the case of
mobilenetv3_large_100
It could have something to do with the
vit
models taking a large amount of memory to execute, and maybe some combination of that coupled with the fact that they are much slower so Daft ends up performing a full materialization of your dataset. I'll let @Colin Ho provide more guidance here
👍 1
f
Yes, it must be some combination like this of model size and images count. Thanks, sounds good!
c
One more thing you can try is setting a `resource_request`: https://www.getdaft.io/projects/docs/en/stable/api_docs/udf.html#resource-requests either
num_cpus
or
memory_bytes
, which should give more manual control of the parallelism of your udf
👍 1
f
Update on this- I set up a plain ol
torch
dataloader and I eventually get OOM errors there too. I can’t point as
daft
as the culprit then. I’ll post again here if’n’when I get this fixed and get this
daft
-based tool in working shape.
❤️ 2
j
Awesome good to know!