If I have a cpu based text classifier that I'd lik...
# daft-dev
k
If I have a cpu based text classifier that I'd like to use for inference across many rows, what's the best way to init once and reuse across my ray cluster? I tried to use the actor pool projection as mentioned in a previous issue but it seems that with concurrency 1, the work being done simultaneously is reduced, and with higher concurrency the tasks get stuck at the projection.
đź‘€ 1
j
@Kevin Wang can advise!
Also could you elaborate on what “tasks getting stuck” means?
k
Meaning i tried matching the with_concurrency to the number of cores I had but they just init without doing work as fast as without the actor pool projection enabled so I'm not sure if they are "stuck with accepting no other task other than initializing the model" or are just slower in doing the inferencing
k
Ah so that will not work because if your actor pool uses up all of the cores, no other work can be scheduled. I'm working on adding better logging and errors around that
This is on Ray right?
k
So it's more of a serving model such that there's cores dedicated to the serving once the model is initialized? Yes it's on ray
If I wanted to just initialize the model per core/machine once but have 100 models inferencing simultaneously is that possible?
k
It's a limitation of Ray, since I believe ray assigns resources to actors instead of actor method invocations. That means for any work to actually be scheduled, you need to have some amount of CPUs+memory left over from the models. So if you do like concurrency=(# cpus - 1) the pipeline should flow properly. A lower value might also be good depending on how much of a bottleneck the stateful UDF actually is versus other parts of the execution
k
I see.. so probably it'll just be better to keep it in the udf to be initialised per method call instead if the overheads are not particularly high? Does the concurrent model initialization avoid going to a single actor and instead distributes it across the actors as much as possible?
k
If your UDF initialization isn't costly, then yes using a stateful UDF probably won't bring significant benefits. I'm not entirely sure how to answer your second question, are you asking if you set concurrency=n, if Daft distributes work across all of the n actors? The answer to that is yes
k
Yes that's what I was asking about, thanks!!
k
np! hope you are able to get it working well
j
Keep us posted! Also if you send us a sample of what kind of UDF you’re running that will be super helpful for our knowledge as well
k
It's a small fasttext model inferencing. I simply have to initialize the model and call it on a text column and I found it weird that pyspark runs it faster than daft so I'm trying to figure out the optimizations i need to do haha
‼️ 1
j
Are you just doing read -> inference -> write?
k
Yep!