Any idea how I should interpret this task failure ...
# general
g
Any idea how I should interpret this task failure based on Ray driver node logs. It doesn't look like any of the worker nodes ever died, my only guess is a brief networking issue on our side. The job gets stuck at this point and does not retry (sitting with one task waiting for resources). Note that I have run this same Ray job multiple times successfully.
Copy code
[2025-03-07 12:50:48,676 I 1537 1586] normal_task_submitter.cc:465: Retrying attempt to schedule task (id: 624ca5a5314469a2c8af12bf4124c94820da58ba05000000 name: Project-Filter-WriteFile [Stage:4]) at remote node (id: 6UH"�L�뻉�"ds���{��T!^ ip: 10.216.202.76). Try again on a local node. Error: RpcError: RPC Error message: Socket closed; RPC Error details:
[2025-03-07 12:50:48,700 I 1537 1586] raylet_client.cc:403: Error getting task result: RpcError: RPC Error message: failed to connect to all addresses; last error: UNAVAILABLE: ipv4:10.216.202.76:37383: recvmsg:Connection reset by peer; RPC Error details:
[2025-03-07 12:50:48,700 W 1537 1586] normal_task_submitter.cc:672: Failed to fetch task result with status RpcError: RPC Error message: failed to connect to all addresses; last error: UNAVAILABLE: ipv4:10.216.202.76:37383: recvmsg:Connection reset by peer; RPC Error details: node id: 36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e ip: 10.216.202.76
[2025-03-07 12:50:48,700 W 1537 1586] task_manager.cc:1103: Task attempt c1ab73323ba089daffffffffffffffffffffffff05000000 failed with error NODE_DIED Fail immediately? 0, status RpcError: RPC Error message: Socket closed; RPC Error details: , error info error_message: "Task failed due to the node (where this task was running) was dead or unavailable.\n\nThe node IP: 10.216.202.76, node ID: 36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e\n\nThis can happen if the instance where the node was running failed, the node was preempted, or raylet crashed unexpectedly (e.g., due to OOM) etc.\n\nTo see node death information, use `ray list nodes --filter \"node_id=36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e\"`, or check Ray dashboard cluster page, or search the node ID in GCS log, or use `ray logs raylet.out -ip 10.216.202.76`"
error_type: NODE_DIED

[2025-03-07 12:50:48,700 I 1537 1586] task_manager.cc:1000: task c1ab73323ba089daffffffffffffffffffffffff05000000 retries left: 3, oom retries left: -1, task failed due to oom: 0
[2025-03-07 12:50:48,700 I 1537 1586] task_manager.cc:1004: Attempting to resubmit task c1ab73323ba089daffffffffffffffffffffffff05000000 for attempt number: 0
[2025-03-07 12:50:48,700 I 1537 1586] core_worker.cc:440: Will resubmit task after a 0ms delay: Type=NORMAL_TASK, Language=PYTHON, Resources: {memory: 6.09078e+08, CPU: 1, }, function_descriptor={type=PythonFunctionDescriptor, module_name=daft.runners.ray_runner, class_name=, function_name=single_partition_pipeline, function_hash=ea0eb0a72e264ab8992bcb09a86dc7da}, task_id=c1ab73323ba089daffffffffffffffffffffffff05000000, task_name=Project-Filter-WriteFile [Stage:4], job_id=05000000, num_args=10, num_returns=2, max_retries=3, depth=1, attempt_number=1, runtime_env_hash=1762274293, eager_install=1, setup_timeout_seconds=600
[2025-03-07 12:50:48,701 I 1537 1586] raylet_client.cc:389: Error returning worker: RpcError: RPC Error message: failed to connect to all addresses; last error: UNAVAILABLE: ipv4:10.216.202.76:37383: recvmsg:Connection reset by peer; RPC Error details:
[2025-03-07 12:50:54,635 I 1537 1586] core_worker.cc:641: Event stats:
j
Does the suggested
ray list nodes --filter \"node_id=36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e\
give us any info?
g
I see this in dashboard_agent.log of the node:
Copy code
2025-03-07 12:50:49,146 ERROR reporter_agent.py:1234 -- Error publishing node physical stats.
Traceback (most recent call last):
 File "/home/ray/anaconda3/lib/python3.12/site-packages/ray/dashboard/modules/reporter/reporter_agent.py", line 1231, in _run_loop
 await publisher.publish_resource_usage(self._key, json_payload)
 File "/home/ray/anaconda3/lib/python3.12/site-packages/ray/_private/gcs_pubsub.py", line 141, in publish_resource_usage
 await self._stub.GcsPublish(req)
 File "/home/ray/anaconda3/lib/python3.12/site-packages/grpc/aio/_call.py", line 318, in __await__
 raise _create_rpc_error(
grpc.aio._call.AioRpcError: <AioRpcError of RPC that terminated with:
 status = StatusCode.UNAVAILABLE
 details = "failed to connect to all addresses; last error: UNKNOWN: ipv4:10.220.100.248:6379: Failed to connect to remote host: Connection refused"
 debug_error_string = "UNKNOWN:Error received from peer {grpc_message:"failed to connect to all addresses; last error: UNKNOWN: ipv4:10.220.100.248:6379: Failed to connect to remote host: Connection refused", grpc_status:14, created_time:"2025-03-07T12:50:49.145826375-08:00"}"
>
potentially related to this problem, I have seen more frequent cases of jobs just "hanging" with "waiting for scheduling" where I eventually just have to kill the job. in the latest one today I am not finding any errors anywhere....
what are the core logs I should look at to debug, there are a lot available in ray 😅
j
Yeah this does seem like a Ray issue hmm
I’m guessing it’s some kind of ephemeral node networking issue from the looks of it
g
ephemeral node networking issue
this is my thinking as well, checking with teams on my side 🤞
still debugging, another error:
Copy code
ray.exceptions.RayTaskError(ObjectFetchTimedOutError): ray::HashJoin [Stage:7]() (pid=25150, ip=10.216.13.29)
  At least one of the input arguments for this task could not be computed:
ray.exceptions.ObjectFetchTimedOutError: Failed to retrieve object a8b93cde8e35f7a1ffffffffffffffffffffffff1000000002000000. To see information about where this ObjectRef was created in Python, set the environment variable RAY_record_ref_creation_sites=1 during `ray start` and `ray.init()`.

Fetch for object a8b93cde8e35f7a1ffffffffffffffffffffffff1000000002000000 timed out because no locations were found for the object. This may indicate a system-level bug.
j
What version of Daft are you using btw?
g
latest 0.4.7
is this anything to worry about?
Copy code
ScanWithTask-Project-Aggregate-FanoutHash [Stage:2]:   0%|          | 0/1 [00:00<?, ?it/s](raylet) WARNING: 20 PYTHON worker processes have been started on node: 9bc176467c0d5c671e8fce26129006dabd6a67a5142f1074c04341f8 with address: 10.216.139.94. This could be a result of using a large number of actors, or due to tasks blocked in ray.get() calls (see <https://github.com/ray-project/ray/issues/3644> for some discussion of workarounds).
j
I have seen more frequent cases of jobs just "hanging" with "waiting for scheduling"
Is this the job waiting to be scheduled, or a task inside of a job?
g
the job was already running, that at some point tasks stop getting scheduled
👍 1
all these errors feel like they are stemming from the same root cause
j
Yeah it does, feels like potentially some kind of ephemeral networking issue that Ray isn't able to recover from
What version of Ray is this running?
g
2.35.0
I can provide the data/query to try to debug
but may be hard if it is some k8s network issue on our side
oof:
Copy code
Normal   Created                 56m                kubelet                  Created container fluentbit
  Normal   Started                 56m                kubelet                  Started container fluentbit
  Normal   Pulled                  10m (x2 over 56m)  kubelet                  Container image "<http://cssacrprod.azurecr.io/istio/proxyv2:1.21.2|cssacrprod.azurecr.io/istio/proxyv2:1.21.2>" already present on machine
  Normal   Created                 10m (x2 over 56m)  kubelet                  Created container istio-proxy
  Normal   Started                 10m (x2 over 56m)  kubelet                  Started container istio-proxy
  Warning  OOMKilled               9m12s              oomkiller                Container not found was killed by oomkiller, app: , limit:unknown, total-vm:2651400kB, anon-rss:1971892kB, file-rss:43088kB, shmem-rss:0kB
I didn't see any indication of OOM in ray logs though
j
Ah… if Kubernetes kills the pod, Ray itself will get killed. Thus Ray won’t give any information here.
g
ok, I think this is fixed with smarter repartitioning prior to a gnarly join
argh, fails again, I don't quite understand why it is not retrying when a node is lost, I swear it did
j
How large of a workload/join is this in terms of approx GB and num partitions? Having the query plan might be helpful here too so we can maybe try to repro on our end.
g
this is a self join of ~5M rows --> 330M rows. there is skew in the keys for which I do a simple "salted join" to account for. will share the query plan soon.