Garrett Weaver
03/07/2025, 8:58 PM[2025-03-07 12:50:48,676 I 1537 1586] normal_task_submitter.cc:465: Retrying attempt to schedule task (id: 624ca5a5314469a2c8af12bf4124c94820da58ba05000000 name: Project-Filter-WriteFile [Stage:4]) at remote node (id: 6UH"�L�뻉�"ds���{��T!^ ip: 10.216.202.76). Try again on a local node. Error: RpcError: RPC Error message: Socket closed; RPC Error details:
[2025-03-07 12:50:48,700 I 1537 1586] raylet_client.cc:403: Error getting task result: RpcError: RPC Error message: failed to connect to all addresses; last error: UNAVAILABLE: ipv4:10.216.202.76:37383: recvmsg:Connection reset by peer; RPC Error details:
[2025-03-07 12:50:48,700 W 1537 1586] normal_task_submitter.cc:672: Failed to fetch task result with status RpcError: RPC Error message: failed to connect to all addresses; last error: UNAVAILABLE: ipv4:10.216.202.76:37383: recvmsg:Connection reset by peer; RPC Error details: node id: 36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e ip: 10.216.202.76
[2025-03-07 12:50:48,700 W 1537 1586] task_manager.cc:1103: Task attempt c1ab73323ba089daffffffffffffffffffffffff05000000 failed with error NODE_DIED Fail immediately? 0, status RpcError: RPC Error message: Socket closed; RPC Error details: , error info error_message: "Task failed due to the node (where this task was running) was dead or unavailable.\n\nThe node IP: 10.216.202.76, node ID: 36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e\n\nThis can happen if the instance where the node was running failed, the node was preempted, or raylet crashed unexpectedly (e.g., due to OOM) etc.\n\nTo see node death information, use `ray list nodes --filter \"node_id=36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e\"`, or check Ray dashboard cluster page, or search the node ID in GCS log, or use `ray logs raylet.out -ip 10.216.202.76`"
error_type: NODE_DIED
[2025-03-07 12:50:48,700 I 1537 1586] task_manager.cc:1000: task c1ab73323ba089daffffffffffffffffffffffff05000000 retries left: 3, oom retries left: -1, task failed due to oom: 0
[2025-03-07 12:50:48,700 I 1537 1586] task_manager.cc:1004: Attempting to resubmit task c1ab73323ba089daffffffffffffffffffffffff05000000 for attempt number: 0
[2025-03-07 12:50:48,700 I 1537 1586] core_worker.cc:440: Will resubmit task after a 0ms delay: Type=NORMAL_TASK, Language=PYTHON, Resources: {memory: 6.09078e+08, CPU: 1, }, function_descriptor={type=PythonFunctionDescriptor, module_name=daft.runners.ray_runner, class_name=, function_name=single_partition_pipeline, function_hash=ea0eb0a72e264ab8992bcb09a86dc7da}, task_id=c1ab73323ba089daffffffffffffffffffffffff05000000, task_name=Project-Filter-WriteFile [Stage:4], job_id=05000000, num_args=10, num_returns=2, max_retries=3, depth=1, attempt_number=1, runtime_env_hash=1762274293, eager_install=1, setup_timeout_seconds=600
[2025-03-07 12:50:48,701 I 1537 1586] raylet_client.cc:389: Error returning worker: RpcError: RPC Error message: failed to connect to all addresses; last error: UNAVAILABLE: ipv4:10.216.202.76:37383: recvmsg:Connection reset by peer; RPC Error details:
[2025-03-07 12:50:54,635 I 1537 1586] core_worker.cc:641: Event stats:jay
03/07/2025, 9:02 PMray list nodes --filter \"node_id=36554822e14cb0ebbb891afe22647383020797957b14e3a0d154215e\ give us any info?Garrett Weaver
03/07/2025, 9:02 PM2025-03-07 12:50:49,146 ERROR reporter_agent.py:1234 -- Error publishing node physical stats.
Traceback (most recent call last):
File "/home/ray/anaconda3/lib/python3.12/site-packages/ray/dashboard/modules/reporter/reporter_agent.py", line 1231, in _run_loop
await publisher.publish_resource_usage(self._key, json_payload)
File "/home/ray/anaconda3/lib/python3.12/site-packages/ray/_private/gcs_pubsub.py", line 141, in publish_resource_usage
await self._stub.GcsPublish(req)
File "/home/ray/anaconda3/lib/python3.12/site-packages/grpc/aio/_call.py", line 318, in __await__
raise _create_rpc_error(
grpc.aio._call.AioRpcError: <AioRpcError of RPC that terminated with:
status = StatusCode.UNAVAILABLE
details = "failed to connect to all addresses; last error: UNKNOWN: ipv4:10.220.100.248:6379: Failed to connect to remote host: Connection refused"
debug_error_string = "UNKNOWN:Error received from peer {grpc_message:"failed to connect to all addresses; last error: UNKNOWN: ipv4:10.220.100.248:6379: Failed to connect to remote host: Connection refused", grpc_status:14, created_time:"2025-03-07T12:50:49.145826375-08:00"}"
>Garrett Weaver
03/10/2025, 3:59 PMGarrett Weaver
03/10/2025, 4:04 PMjay
03/10/2025, 4:11 PMjay
03/10/2025, 4:12 PMGarrett Weaver
03/10/2025, 4:12 PMephemeral node networking issuethis is my thinking as well, checking with teams on my side 🤞
Garrett Weaver
03/10/2025, 6:30 PMray.exceptions.RayTaskError(ObjectFetchTimedOutError): ray::HashJoin [Stage:7]() (pid=25150, ip=10.216.13.29)
At least one of the input arguments for this task could not be computed:
ray.exceptions.ObjectFetchTimedOutError: Failed to retrieve object a8b93cde8e35f7a1ffffffffffffffffffffffff1000000002000000. To see information about where this ObjectRef was created in Python, set the environment variable RAY_record_ref_creation_sites=1 during `ray start` and `ray.init()`.
Fetch for object a8b93cde8e35f7a1ffffffffffffffffffffffff1000000002000000 timed out because no locations were found for the object. This may indicate a system-level bug.jay
03/10/2025, 6:31 PMGarrett Weaver
03/10/2025, 6:32 PMGarrett Weaver
03/10/2025, 6:32 PMScanWithTask-Project-Aggregate-FanoutHash [Stage:2]: 0%| | 0/1 [00:00<?, ?it/s](raylet) WARNING: 20 PYTHON worker processes have been started on node: 9bc176467c0d5c671e8fce26129006dabd6a67a5142f1074c04341f8 with address: 10.216.139.94. This could be a result of using a large number of actors, or due to tasks blocked in ray.get() calls (see <https://github.com/ray-project/ray/issues/3644> for some discussion of workarounds).jay
03/10/2025, 6:37 PMI have seen more frequent cases of jobs just "hanging" with "waiting for scheduling"Is this the job waiting to be scheduled, or a task inside of a job?
Garrett Weaver
03/10/2025, 6:38 PMGarrett Weaver
03/10/2025, 6:38 PMjay
03/10/2025, 6:39 PMjay
03/10/2025, 6:46 PMGarrett Weaver
03/10/2025, 6:53 PMGarrett Weaver
03/10/2025, 7:03 PMGarrett Weaver
03/10/2025, 7:04 PMGarrett Weaver
03/10/2025, 8:03 PMNormal Created 56m kubelet Created container fluentbit
Normal Started 56m kubelet Started container fluentbit
Normal Pulled 10m (x2 over 56m) kubelet Container image "<http://cssacrprod.azurecr.io/istio/proxyv2:1.21.2|cssacrprod.azurecr.io/istio/proxyv2:1.21.2>" already present on machine
Normal Created 10m (x2 over 56m) kubelet Created container istio-proxy
Normal Started 10m (x2 over 56m) kubelet Started container istio-proxy
Warning OOMKilled 9m12s oomkiller Container not found was killed by oomkiller, app: , limit:unknown, total-vm:2651400kB, anon-rss:1971892kB, file-rss:43088kB, shmem-rss:0kBGarrett Weaver
03/10/2025, 8:05 PMjay
03/10/2025, 9:50 PMGarrett Weaver
03/10/2025, 10:46 PMGarrett Weaver
03/10/2025, 11:37 PMjay
03/11/2025, 12:27 AMGarrett Weaver
03/11/2025, 8:01 PM