LiveKit Inference timeouts happening since past 2 weeks

Thanks, I have deployed an agent with that configuration and see similar observability data, so that mystery is solved. Digging a little more in your traces, the failed llm_request lists the model as openai/gemma-4-31b-it. Do you know where that comes from? That is not a model I would expect to see based on the above, or any internal fallback logic we have.

Thanks, i see fallback model is using openai/gemma-4-31b-it, really strange, so i think its an issue, i will dig more and check the confguration for fallbacks again.

But again we get to part 1 bug, about timeouts, so when timeout would happen, then only it will go to fallback model, so still things about timeout is an open issue?

i have fixed fallback model issue,

I’m still seeing a lot of session timeouts.

https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_xJMEoEU3RPuu/observability?mode=transcript
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_LXTJTPcXgP7G/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_tWYovDJHZ7nE/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_BKHJvp3VJd9q/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_CgdjnxxuQHJN/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_s7f94Ysj5Wbz/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_saCbm9F2RpgP/observability

But timeout still happens as below example

Thanks,

Engineering have informed me they have made a fix which should improve the timeout situation for models with a regional endpoint (so, your original gpt-5.4-mini I would now expect to not see the level of timeouts you reported at the start of this thread)

This will not improve the Gemma situation, since this is currently in the US only, which is why you continue to see timeouts with that model when called from the EU.

I see evidence in the session you shared of the fallback adapter working to the gpt-5.4-mini, but for now I recommend not invoking Gemma from an EU session.

Thanks so much, i hope they support gemma soon in eu region and then finally the timeouts would be fixeed, myself and many other out there would be really keen to have it, i think its fine i can keep using gemma, will reduce timeout setting to half and use fallback as openai, as 70% of the time it works and its really faster than 5.4 mini around ~300 ttft for me.

@darryncampbell
So this is not region specific issue, i changed the region to us as you can check for this session, still its happening multiple times in single session like 20+ times. So definitely not regional issue or not specific to gemma model, its livekit inference wide issue for all models.

https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_F9LoqhQ3m2e6/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_Hm58jsoziEsZ/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_zrkcjtiV8tTr/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_4rbjLgDmnoZ6/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_MCQeWZfjeu4D/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_7sMdrmR7DdGd/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_LSTAt5zfZF3N/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_LSCFjsW57WPa/observability

Hi Kaushal, I was going to message you this morning saying we had an incident yesterday:

There is some overlap with the sessions you shared, but it does not account for the sessions that happened after 1500 UTC. I’m investigating whether the incident may have had a longer duration than what is reported on that page.

Hey, thanks but multiple sessions have timeouts which are outside of incident time like
few sessions below which i can find, there are others too, i will also check how it goes for few upcoming days

https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_ediYk2frXMfC/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_NXeDa8yw4kvc/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_Bwp89dKHKTWa/observability

Does likevit team records this timeouts globablly somewhere for livekit inference for all projects? as it might be very easy for team to then investigate on why it started suddently and when? as its reallly something after 14th july that the issue has started and some days its okay and some days its worst.

I am recording almost everthing on db so that is why was able to notice, otherwise mostly people wont notice this timeouts as 70-80% of sessions it works. I hope fix is added soon for livekit inference

I took a quick glance at the sessions you shared, and they look different than what I saw before. We have passed those off to the inference team for deeper review.

Do you have agent logs for those sessions that I can review? To me this looks like an issue with Agent logic but can’t be sure without some extra details.

It would also be helpful to see how you configured your LLM, STT, and TTS. If you don’t mind sharing a snippet of that, please include any timeouts you set and any fallbacks you’re using.

So above all session i shared has just agent observability tracing enabled, so you can still check tracing tab for full details, there is no audio, transcript or logs recorded for those sessions.

By the way issue still happens today too, here are some sessions

https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_yL8an4dcEjj6/observability
https://cloud.livekit.io/projects/p_3scf9tt1e2i/sessions/RM_fCiDx9Xhj3J6/observability

Observability is helpful, but agent logs would help a lot too.

Here’s the sample config configuration snippet, Sorry i cant enable full observability as i its not free for us in plan. i have enabled for some days to share some data with you.

ai-pipeline-config.md (4.1 KB)

Thank you. That is helpful. We are looking into it.

Some error captured in sentry if that might be helpful

InternalServerError
provider: vercel model: google/gemma-4-31b-it, message: POST "https://ai-gateway.vercel.sh/v1/chat/completions": 503 Service Unavailable {"message":"Service temporarily unavailable. Please try again shortly.","type":"service_unavailable_error","param":{"error":"Service temporarily unavailable. Please try again shortly.","type":"service_unavailable_error","statusCode":503}}.  {"error":{"message":"Service temporarily unavailable. Please try again shortly.","type":"service_unavailable_error","param":{"error":"Service temporarily unavailable. Please try again shortly.","type":"service_unavailable_error","statusCode":503}},"providerMetadata":{"gateway":{"routing":{"originalModelId":"google/gemma-4-31b-it","resolvedProvider":"novita","fallbacksAvailable":["parasail","cerebras"],"canonicalSlug":"google/gemma-4-31b-it","modelAttemptCount":1,"modelAttempts":[{"canonicalSlug":"google/gemma-4-31b-it","success":false,"providerAttemptCount":2,"providerAttempts":[{"provider":"parasail","credentialType":"system","success":false,"error":"Service temporarily unavailable","startTime":1785940674867,"endTime":1785940675057,"statusCode":503},{"provider":"cerebras","credentialType":"system","success":false,"error":"Service temporarily unavailable","startTime":1785940675057,"endTime":1785940675171,"statusCode":503}]}],"totalProviderAttemptCount":2,"skippedProviderAttempts":[{"credentialType":"system","provider":"novita","reason":"zdr_not_supported"}]},"generationId":"gen_01KZ95R28A00KPMFCXTTB3J9VQ"}}}: POST "https://ai-gateway.vercel.sh/v1/chat/completions": 503 Service Unavailable {"message":"Service temporarily unavailable. Please try again shortly.","type":"service_unavailable_error","param":{"error":"Service temporarily unavailable. Please try again shortly.","type":"service_unavailable_error","statusCode":503}}
APITimeoutError
Request timed out.

chat google/gemma-4-31b-it


Hey Kaushal,

Thanks for you patience on this issue. Our Gemma offering has become quite popular and we’re seeing scaling issues as we try to meet demand, which is the cause of the periodic timeouts. I’m from the Inference engineering team and we’ve identified some bottlenecks with Gemma and we’re actively working on addressing them. Once of the fixes went out last night, the data indicates an improvement in Gemma availability. Would you be able to share your stats? I’m curious if you’re seeing an improvement.

Also, here’s a quick summary of what we’re observing. We get a daily ramp of requests for Gemma, part of this ramp is a storm of requests with large contexts, this storm has slowly increased over time. We’ve recently crossed a threshold such that our Gemma deployment can’t keep up with tokenizing these large contexts…especially as these conversations carry on. Specifically, we’ve identified a potential CPU bottleneck tokenizing these large contexts. We’re tuning the deployment to handle this volume of large contexts more efficiently.

Hey, thanks. I hope it will be fixed soon, i will be checking today as well as next week, as for weekends we ususally saw lesser timeouts, so next week would be ideal to check if issue is fixed or not.

Below are the full timeout statistics starting from 11 July 2026 in table, broken down into four 6-hour time periods per day. I haven’t included any dates before 11 July, as there were minimal to no timeouts at all before then. I’ve also added a chart to help visualise when the issue first started, how it progressed over time, and which days had relatively low timeout counts versus the days when the problem became significantly worse. This should give you a much clearer picture of how the issue evolved, including how it escalated from 14 July onwards.

Also to add, we were initially using GPT-5.4 Mini in the EU-Central region, before switching to Gemma 4 on 22 July. Both models were running in the EU-Central region. You then suggested that the timeouts might be related to the region, as Gemma 4 is better suited for the US region, so I also switched our deployment to the US region on 4 August while continuing to use Gemma 4.

Date Model Region Timeout count
Before 22 Jul GPT-5.4 Mini EU-Central 214
22 Jul – 3 Aug Gemma 4 EU-Central 247
From 4 Aug Gemma 4 US 350

I have experienced the highest timeouts in the US region.

Table stats data with date/time-range
timeout_stats.md (7.2 KB)

@Adrian_Cowham

To add on, this was not issue just for gemma model, i was experiencing it also with other models, so initially with openai gpt 5.4 mini too. So the fix should be global? for all models that we can use via livekit inference.

So if its not fixed next week, we may need to think of moving away from inference as its been degrading the quality of sessions, i hope we get same quality/latency with zeor to minimal timeouts as it was working before past few weeks

@Adrian_Cowham @darryncampbell

So its still happening, 30 times total till now.

Timeouts on Gemma? I’m not seeing a high TTFT right now. Do you still have your fallbackadapter timeout set to 2.5s?