Gemini-3.8-live rejected with vertexai=True in livekit-plugins-google, although Vertex AI serves it (EU multi-region)

Hi LiveKit community,

We run a realtime voice agent on Gemini Live (google.realtime.RealtimeModel), currently gemini-3.1-flash-live-preview through the Gemini Developer API. We want to move it to Vertex AI in the EU multi-region for data residency and enterprise terms. gemini-3.8-live is the natural target, but the plugin blocks it before any connection is made.

Versions

  • livekit-agents 1.8.3
  • livekit-plugins-google 1.8.3 (we also checked main at the time of writing: same code)
  • google-genai 2.25.0

Repro

from livekit.plugins import google

llm = google.realtime.RealtimeModel(
model=“gemini-3.8-live”,
vertexai=True,
project=“”,
location=“eu”,
)

Result

ValueError: Model ‘gemini-3.8-live’ is a Gemini API model, but vertexai=True.
Use a VertexAI model (e.g., ‘gemini-live-2.5-flash-native-audio’) or set vertexai=False.

This comes from _validate_model_api_match() in realtime/realtime_api.py: gemini-3.8-live and gemini-3.8-live-extended-thinking are listed only in KNOWN_GEMINI_API_MODELS, while KNOWN_VERTEXAI_MODELS only contains gemini-live-2.5-flash-native-audio.

Vertex AI does serve this model. We called google-genai directly (client.aio.live.connect, vertexai=True, service account with roles/aiplatform.user, no allowlist request) on 2026-10-01:

gemini-3.8-live:

  • eu: works
  • us: works
  • europe-west4: 1008 “Publisher model … was not found”
  • global: 1008 “Publisher model … was not found”

gemini-live-2.5-flash-native-audio:

  • europe-west4: works
  • eu: 1008 “Publisher model … was not found”

On eu, gemini-3.8-live streamed audio replies normally over 8 test sessions, and responded faster than 3.1 on the Developer API in the same conditions (median 1.07 s vs 1.29 s from end of user speech to first audio chunk). The plugin itself also lists gemini-3.8-live in api_proto.py and applies 3.8-specific handling (MODELS_WITHOUT_REPLY_PLACEHOLDER, MODELS_DEFAULT_NON_BLOCKING), so the rest of the code path looks ready.

Suggested fix

gemini-3.8-live (and probably gemini-3.8-live-extended-thinking) is available on both APIs. The validation should only raise for models that exist on one side only, for example:

KNOWN_VERTEXAI_ONLY_MODELS = frozenset({“gemini-live-2.5-flash-native-audio”})
KNOWN_GEMINI_API_ONLY_MODELS = frozenset({
“gemini-3.1-flash-live-preview”,
“gemini-2.5-flash-native-audio-preview-12-2025”,
})
KNOWN_DUAL_API_MODELS = frozenset({
“gemini-3.8-live”,
“gemini-3.8-live-extended-thinking”,
})

_validate_model_api_match() would then skip KNOWN_DUAL_API_MODELS. It might also help to mention in the docs that the Live API on Vertex is not available on global, and that 3.8 Live is served on the eu / us multi-regions rather than single regions like europe-west4.

A related difference you may want to handle

On the Gemini Developer API, gemini-3.8-live rejects thinking_config=ThinkingConfig(thinking_level=…) with 1007 Thinking level is not supported for this model. On Vertex eu, the same config was accepted. If the plugin forwards thinking_config as-is, this difference between the two APIs is worth documenting.

Questions

  1. Is a fix planned to accept gemini-3.8-live with vertexai=True? We’re happy to open a PR along the lines above if that helps.
  2. Until it ships, what is the recommended way to use it on Vertex without patching the plugin?

Thanks!

Hi @Akiniesta, I have been able to use gemini-3.8-livein Livekit using the Gemini API by setting the GOOGLE_API_KEY env variable. And it has been working fine.

I’m trying to understand what your issue is. I would like to ask, have you tried using the model gemini-3.8-live in Livekit using Gemini Live API before moving it to Vertex. If yes, has that worked for you?

When using Vertex to authenticate, you will have to provide GOOGLE_APPLICATION_CREDENTIALS with the proper roles. I believe you’re doing that.

From the error messages you’ve provided, it shows that gemini-3.8-live is accepted by Livekit when using the Gemini API, but Livekit is rejecting it when using Vertex. Is that correct?

But when you used gemini-3.8-live with google genai, both Gemini API & Vertex. Is that correct?

Yes, this is right. I agree that this should be highlighted in the Livekit Docs.

Hi, thanks for looking into this!

To answer your questions:

  1. Gemini API (GOOGLE_API_KEY): yes, gemini-3.8-live works for us too, both directly with google-genai and through the LiveKit plugin (google.realtime.RealtimeModel, livekit-plugins-google 1.8.3). In the plugin, generate_reply() works, and so do replies to user audio streamed in real time.
  2. LiveKit rejects it with Vertex: correct. The error is raised by the plugin itself, before any network call or credential check:

ValueError: Model ‘gemini-3.8-live’ is a Gemini API model, but vertexai=True.
Use a VertexAI model (e.g., ‘gemini-live-2.5-flash-native-audio’) or set vertexai=False.

It comes from _validate_model_api_match() in livekit/plugins/google/realtime/realtime_api.py: gemini-3.8-live and gemini-3.8-live-extended-thinking are listed only in KNOWN_GEMINI_API_MODELS, while KNOWN_VERTEXAI_MODELS only contains gemini-live-2.5-flash-native-audio. We checked main as well and the lists are the same.

It isn’t an authentication issue: with the same service account (roles/aiplatform.user, via GOOGLE_APPLICATION_CREDENTIALS), gemini-live-2.5-flash-native-audio works fine on Vertex through the plugin.
3. google-genai directly, on both APIs: correct. With the same service account and vertexai=True, gemini-3.8-live streams audio replies normally with location=“eu” (8 out of 8 test sessions) and location=“us”. It returns 1008 Publisher model … was not found on global and europe-west4, so it seems to be served on the multi-regions only.

So the model works on both APIs at the Google level. Only the plugin’s model/API validation prevents using it with vertexai=True.

Minimal repro (livekit-agents / livekit-plugins-google 1.8.3):

from livekit.plugins import google

google.realtime.RealtimeModel(
model=“gemini-3.8-live”,
vertexai=True,
project=“”,
location=“eu”,
)

→ ValueError before any connection

Suggested fix: treat models served by both APIs as valid for either, and only raise for models that exist on one side. For example:

KNOWN_VERTEXAI_ONLY_MODELS = frozenset({“gemini-live-2.5-flash-native-audio”})
KNOWN_GEMINI_API_ONLY_MODELS = frozenset({
“gemini-3.1-flash-live-preview”,
“gemini-2.5-flash-native-audio-preview-12-2025”,
})
KNOWN_DUAL_API_MODELS = frozenset({
“gemini-3.8-live”,
“gemini-3.8-live-extended-thinking”,
})

_validate_model_api_match() would then skip KNOWN_DUAL_API_MODELS. It might also be worth mentioning in the docs that the Live API on Vertex isn’t available on global, and that 3.8 Live is served on the eu / us multi-regions.

One related difference: on the Gemini API, gemini-3.8-live rejects thinking_config=ThinkingConfig(thinking_level=…) with 1007 Thinking level is not supported for this model, while Vertex eu accepted the same config. Worth knowing if the plugin forwards thinking_config as-is.

We’re happy to open a PR along these lines if that helps. Until then, is there a recommended way to use gemini-3.8-live on Vertex without patching the plugin?

Thanks again!

Thanks, for reporting. I can reproduce this and I’ve raised this with our agents team.

I’m not sure if a separate table for KNOWN_DUAL_API_MODELS would be the right approach - I’m not a maintainer of the code, but that feels like one more thing to keep up to date manually.

Claude did suggest a workaround which seems to work in my testing, to provide what is (as far as I can tell) an undocumented model name:

model="publishers/google/models/gemini-3.8-live",
vertexai=True,

Good catch on the difference on thinking_config behaviour, I’ve also raised that with the agents team as it will surely affect how we properly handle the fix for this.

Thanks a lot @Darryn ! I confirm the workaround works on our side: publishers/google/models/gemini-3.8-live with vertexai=True and location=“eu”, through google.realtime.RealtimeModel (livekit-plugins-google 1.8.3).

Both generate_reply() and replies to streamed user audio work over several runs. We’ll use it until the fix lands and switch back to the plain model name then. Agreed on the dual-API table, a model registry that doesn’t need manual upkeep would be better.

Thanks for raising the thinking_config difference too.