Signal connection times out on the "v0 path" at agent join, forcing a fallback that adds 0.5–5s of call-setup latency

Setup: voice agents on livekit-agents 1.4.6 / livekit 1.1.2 / livekit-api 1.0.7 / livekit-protocol 1.1.1 (Python 3.11), outbound SIP telephony, LiveKit Cloud (project p_3tqm7ro6kbs). Each call = one agent worker process; the worker connects to the room via ctx.connect() right after the job starts.

What we saw: within a ~19-second window, 5 near-simultaneous outbound calls (across 4 different worker processes) all logged, at room-join time:

livekit_api::signal_client:287 - signal connection failed on v0 path: Timeout("signal connection timed out")

In every case the SDK then recovered — Connected to LiveKit followed ~0.3–0.5s later for 4 of the 5 calls, and those calls proceeded normally (voicemail detection / conversation, clean shutdown). So the warning was non-fatal.

But one call was worse: room RM_qxjzLwVsFw2F logged the v0 timeout twice (13:28:56 and 13:29:01 UTC) plus:

The room connection was not established within 10 seconds after calling job_entry.
This might mean that job_ctx.connect() was never invoked, or that no AgentSession was started ...

and only connected ~5 seconds after the first attempt. Our audio egress (recording) consequently started ~5s late on that call.

Affected rooms (project p_3tqm7ro6kbs), all 2026-06-08:

Room ID Job ID First v0 timeout (UTC) Outcome
RM_kBfppXUVVE7Q AJ_PSNKRFeqwpwx 13:28:42.804 recovered ~0.5s → completed normally (~94s)
RM_w8sdJUdCogke AJ_iEDtHdH5ioc8 13:28:48.895 recovered ~0.5s → voicemail → clean shutdown
RM_DiFYT4TAudNJ AJ_y7VquhMGuZqp 13:28:49.361 agent connected; remote SIP participant never joined (see note)
RM_CNnwLf7tPqR3 AJ_eTmKwpmqWu85 13:28:55.834 recovered ~0.5s → voicemail → clean shutdown
RM_qxjzLwVsFw2F AJ_VyLSqagwr8e6 13:28:56.081 two v0 timeouts + 10s room-connection warning → connected ~5s late

Please help investigate this

Hi Zaheer,

Could this be a capacity issue with your agents? I see a lot of ‘no worker available’ for your project around this time and, looking at other projects besides your own for the timeline covered by your 5 failing rooms, I don’t see any other projects who had calls that were slow to connect, so it seems to be isolated to your project.

Hello Darryn,

The worker got connected for these calls within 200ms and it started the process - the above log appeared when connecting to LiveKit.

livekit_api::signal_client:287 - signal connection failed on v0 path: Timeout("signal connection timed out")

I see a lot of ‘no worker available’ for your project around this time and, looking at other projects besides your own for the timeline covered by your 5 failing rooms

We have autoscaling enabled and also ondemand instances running - so there is very small chance we get this. Are you seeing this across multiple days? Is there a way for me to check which Rooms OR Calls the no worker available appeared?

I don’t want to mislead you, that ‘no worker available’ is part of the standard process to find an agent to handle the job - if no workers are available then the job scheduler will back off and retry a few seconds later - in the cases you highlighted the jobs were accepted, so it’s not the root cause.

I did double-check the server logs again around RM_qxjzLwVsFw2F and there’s nothing in there which would explain why the again was slow to join.

I ran your error message through Claude, and it spat out the following (along with a load of other explanation):

The agent process was too CPU-saturated (a single worker running far too many concurrent jobs during a call burst) to complete the signal WebSocket handshake within connect_timeout, so the Rust SDK aborted the v0 connect attempt with Timeout and retried — succeeding once CPU freed up.

Obviously take that with a grain of salt, but it does track with there being no clues in the server logs.

Thanks for rechecking this Darryn.

I am sure that we did not have a high CPU spike - we have autoscaling setup that scales with a good amount of resources and doesn’t allow many concurrent calls on a single pod

See above image - the time is in IST here (18:58 IST - 13:28 UTC), which shows that we had enough CPU resources to handle those calls.

Will see if I can check more details on our side too.

@darryncampbell - seeing this again today for the below rooms. And I don’t see any high resource usage too. Would really appreciate if you can help us understand this better

Error Message

livekit_api::signal_client:287:livekit_api::signal_client - signal connection failed on v0 path: Timeout(\"signal connection timed out\")"

Affected Rooms

  ┌──────────────────┬─────────────────┬─────────────────────────────────────────────────────────────────────┐
  │ First v0 timeout │     Room ID     │                   Setup delay (entry → connected)                   │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:07:24         │ RM_wEey2AcGhxfH │ ~5s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:12:56         │ RM_8sxYdkKV9fef │ ~5s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:13:23         │ RM_cHCojGmsaoQw │ ~13s (2 v0 timeouts + "room connection not established within 10s") │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:13:27         │ RM_JCvmB5kHNJNR │ ~5s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:15:31         │ RM_wgjMbdjGtbqE │ ~8s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:15:40         │ RM_67ZrHjxkuin5 │ ~5s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:15:42         │ RM_k9MMKqqx7oiy │ ~9s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:15:48         │ RM_BEivJYLhHyLZ │ ~9s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:15:55         │ RM_tz9ifzekMdsf │ ~5s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:16:16         │ RM_SPUYFMaqLb6M │ ~6s                                                                 │
  ├──────────────────┼─────────────────┼─────────────────────────────────────────────────────────────────────┤
  │ 15:16:24         │ RM_r5z2darvm8Rb │ ~8s                                                                 │
  └──────────────────┴─────────────────┴─────────────────────────────────────────────────────────────────────┘

This smells of the Agent’s main thread getting blocked somehow. Do you have raw logs from the agent? Another possibility is that the agent establishing the WebSocket is being delayed. Do you have any network logs that may show a block of some kind?

Thanks @CWilson.

Raw log line (Rust signal client, one per failed attempt):

livekit_api::signal_client:287 - signal connection failed on v0 path: Timeout("signal connection timed out")
  1. Agent main thread blocked? No, we don’t have a blocking thread and we call ctx.connect() as soon as call start. Below is the logging pattern we see for all these cases
{"message": "received job request", "level": "INFO", "name": "livekit.agents", "job_id": "AJ_G2qB2BM7WJ3c", "dispatch_id": "AD_qrixUCtWz3h7", "room": "tcn_fsa_use2_000_outbound_prod_deployment_+13213810097_t2rSLUzhZA7N", "room_id": "RM_wEey2AcGhxfH", "agent_name": "prod_inbound", "resuming": false, "enable_recording": false, "timestamp": "2026-06-09T15:07:19.149428+00:00"}

{"message": "livekit_api::signal_client:287:livekit_api::signal_client - signal connection failed on v0 path: Timeout(\"signal connection timed out\")", "level": "WARNING", "name": "livekit", "room_name": "tcn_fsa_use2_000_outbound_prod_deployment_+13213810097_t2rSLUzhZA7N", "worker_id": "AW_tZhvxupaFce7", "call_id": "call_AJ_G2qB2BM7WJ3c", "pid": 199833, "job_id": "AJ_G2qB2BM7WJ3c", "room_id": "RM_wEey2AcGhxfH", "timestamp": "2026-06-09T15:07:24.249696+00:00"}

{"levelname": "INFO", "name": "agent_orchestrator", "process": 199833, "event": "Connected to LiveKit", "message": "", "pathname": "/home/app/agent-orchestrator/agent_orchestrator/service/call_service.py", "lineno": 2088, "call_id": "call_AJ_G2qB2BM7WJ3c", "time_taken": 5556, "room_name": "tcn_fsa_use2_000_outbound_prod_deployment_+13213810097_t2rSLUzhZA7N", "tenant_id": "fsa_use2_000", "timestamp": "2026-06-09T15:07:24.721242+00:00"}
  1. Agent slow to establish the WebSocket? No - we call ctx.connect() immediately at job start

  2. Network block? No — we don’t notice any packet drops from the pod’s NIC logs

Happy to share full raw agent logs for any of these rooms (e.g. RM_cHCojGmsaoQw, RM_wgjMbdjGtbqE, RM_67ZrHjxkuin5).

Thanks again for helping out with response. Apologies for the delay, had to pull out network logs to confirm this once

Looking through the data I see AW_tZhvxupaFce7 and AW_x8txT3Ev6nYW tends to take longer than all the rest.

Is it possible to get agent logs from that worker between 2026-06-09 15:00:00 - 2026-06-09 15:30:00 UTC? I would like to try and correlate server logs with agent logs.

If it is possible free to DM to me. But let me know here if you do share so I know to check my DM.

If that timeframe is too broad then a few minutes areound 2026-06-09 15:13:18 would be helpful. and 2026-06-08 14:29:23

I spent quite a bit of time digging through the logs on this and the other thread. I can see that jobs are being offered to the workers but they are not taking the jobs. This is what is caussing the assignment delay.

Here is a graph of no worker and max attempts reached:

Here are example of the logs messages on our side:

To go any deeper than that we would need to do a detailed review of your agent logs particularly the ones mentioned above as those have the mose.

I have to head out but the server literally can’t find a free worker during bursts. Only 5 workers exist, 2 carry 78% of traffic, and bursts exhaust the pool. Crucially, this explains your “CPU was fine” point: the binding limit is max concurrent jobs per worker, not CPU — so a CPU-based autoscaler never reacts while the dispatcher is already throwing “max attempts reached.” The fix is worker capacity/headroom and the per-worker job-concurrency limit (and better distribution off the two hot workers), not CPU scaling.

Have a look at this doc

Thanks a lot for digging deeper @CWilson :raising_hands:

I have shared the log files over Slack DM. I am also investigating on our scaling logic and if we can modify that to use more metrics than just CPU and Memory

I checked the logs. That pretty much re-confirms the issue:

2026-06-09T15:22:39.999468+00:00 INFO livekit.agents worker is at full capacity, marking as unavailable [load=0.7106701000000005 threshold=0.7]

So it won’t take new jobs when offered until the load falls below the threshold.

See Self-hosted deployments | LiveKit Documentation

I checked that too - it was very transient. The way cpu is being measured is every 2.5 secs per 1.4.6 and we are on CGroupV2. If you check the logs around it - it switches back to being available after the same second. Unfortunately there is no workerId logged here to identify which worker logged what.

@CWilson - continuing our conversation of signal connection timeout in this thread.

We now have good headroom for our agent pods along with autoscaling to serve the current traffic without delays - you should notice lesser worker is at full capacity and zero failures of maximum attempts reached. This is just FYI.

Coming to the signaling connection issue - as per your analysis is a networking issue and NOT a capacity issue. I digged through all our vpcflowlogs from our pod - no packets had dropped and don’t see port exhaustion errors on our NAT gateway.

The below is the log we are still seeing for the below rooms:
livekit_api::signal_client:287 - signal connection failed on v0 path: Timeout("signal connection timed out")

All 13 occurrences (project p_3tqm7ro6kbs):

  ┌────────────┬─────────────────┬─────────────────┐
  │ Time (UTC) │     Room ID     │     Worker      │
  ├────────────┼─────────────────┼─────────────────┤
  │ 01:59:44   │ RM_h6BiM7yDYd2B │ AW_8vLrg6NX7uMv │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:08:11   │ RM_YoTqjYV6L2W7 │ AW_yLX7Bv8Bd432 │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:08:14   │ RM_pyfSxbSMUNtr │ AW_dVcPdmxwMBnJ │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:11:43   │ RM_viqZWyBwprUv │ AW_tU9QX4DdqtWU │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:12:10   │ RM_hYLfqi3N2cLz │ AW_U8JhQadVJuHB │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:13:31   │ RM_fmWoHHuddG3Y │ AW_dVcPdmxwMBnJ │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:14:39   │ RM_Hf4is9dcsqfZ │ AW_yLX7Bv8Bd432 │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:15:33   │ RM_5w8mMZmsThaV │ AW_yLX7Bv8Bd432 │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:18:00   │ RM_RZy6ZSd35Ttt │ AW_U8JhQadVJuHB │
  ├────────────┼─────────────────┼─────────────────┤
  │ 16:18:02   │ RM_RqxqgzSMGZxX │ AW_tU9QX4DdqtWU │
  ├────────────┼─────────────────┼─────────────────┤
  │ 17:28:56   │ RM_tMvrdWvsELkC │ AW_tU9QX4DdqtWU │
  ├────────────┼─────────────────┼─────────────────┤
  │ 17:29:10   │ RM_CUcFx3ZTWDK7 │ AW_tU9QX4DdqtWU │
  ├────────────┼─────────────────┼─────────────────┤
  │ 17:29:15   │ RM_AXFFcirDv3NB │ AW_yLX7Bv8Bd432 │
  └────────────┴─────────────────┴─────────────────┘

Can you please check what happened here :folded_hands:

Around these signal failure cases. I’m also seeing the room connection stall even though ctx.connect() is the first thing our entrypoint does — the SDK logs The room connection was not established within 10 seconds after calling job_entry.

In two such cases, the moment the connection finally established we also got:

livekit_ffi::server::room:149 - audio filter cannot be enabled: LiveKit Cloud is required

Why would this above log appear ^?

  ┌───────────────────┬─────────────────┬─────────────────┬───────────┐
  │ Connected (UTC,   │      Room       │       Job       │ Connect   │
  │    2026-06-11)    │                 │                 │   time    │
  ├───────────────────┼─────────────────┼─────────────────┼───────────┤
  │ 16:10:05          │ RM_GzBSh8Hn7XXa │ AJ_UCK4eHXf3eFY │ ~137s     │
  ├───────────────────┼─────────────────┼─────────────────┼───────────┤
  │ 16:15:41          │ RM_qmx8g9BacSTt │ AJ_qLTXcjWfwdyZ │ ~136s     │
  └───────────────────┴─────────────────┴─────────────────┴───────────┘

I think this may be related to this issue I just created:

I will need to hear from the experts.

But I think that is just a bad side effect of some network-related issue that I am not sure how we can pinpoint.

I looked through recent customer issues over the past week, and I am not seeing any widespread signal connection timed out reports.

I will need your agent logs to dig any deeper for this. Maybe for AW_tU9QX4DdqtWU.

Shared the logs over DM

@Zaheer_Abbas, how often are you seeing signal connection timed out? Is it on going or is it only that group?