Hi LiveKit community,
We run a realtime voice agent on Gemini Live (google.realtime.RealtimeModel), currently gemini-3.1-flash-live-preview through the Gemini Developer API. We want to move it to Vertex AI in the EU multi-region for data residency and enterprise terms. gemini-3.8-live is the natural target, but the plugin blocks it before any connection is made.
Versions
- livekit-agents 1.8.3
- livekit-plugins-google 1.8.3 (we also checked main at the time of writing: same code)
- google-genai 2.25.0
Repro
from livekit.plugins import google
llm = google.realtime.RealtimeModel(
model=“gemini-3.8-live”,
vertexai=True,
project=“”,
location=“eu”,
)
Result
ValueError: Model ‘gemini-3.8-live’ is a Gemini API model, but vertexai=True.
Use a VertexAI model (e.g., ‘gemini-live-2.5-flash-native-audio’) or set vertexai=False.
This comes from _validate_model_api_match() in realtime/realtime_api.py: gemini-3.8-live and gemini-3.8-live-extended-thinking are listed only in KNOWN_GEMINI_API_MODELS, while KNOWN_VERTEXAI_MODELS only contains gemini-live-2.5-flash-native-audio.
Vertex AI does serve this model. We called google-genai directly (client.aio.live.connect, vertexai=True, service account with roles/aiplatform.user, no allowlist request) on 2026-10-01:
gemini-3.8-live:
- eu: works
- us: works
- europe-west4: 1008 “Publisher model … was not found”
- global: 1008 “Publisher model … was not found”
gemini-live-2.5-flash-native-audio:
- europe-west4: works
- eu: 1008 “Publisher model … was not found”
On eu, gemini-3.8-live streamed audio replies normally over 8 test sessions, and responded faster than 3.1 on the Developer API in the same conditions (median 1.07 s vs 1.29 s from end of user speech to first audio chunk). The plugin itself also lists gemini-3.8-live in api_proto.py and applies 3.8-specific handling (MODELS_WITHOUT_REPLY_PLACEHOLDER, MODELS_DEFAULT_NON_BLOCKING), so the rest of the code path looks ready.
Suggested fix
gemini-3.8-live (and probably gemini-3.8-live-extended-thinking) is available on both APIs. The validation should only raise for models that exist on one side only, for example:
KNOWN_VERTEXAI_ONLY_MODELS = frozenset({“gemini-live-2.5-flash-native-audio”})
KNOWN_GEMINI_API_ONLY_MODELS = frozenset({
“gemini-3.1-flash-live-preview”,
“gemini-2.5-flash-native-audio-preview-12-2025”,
})
KNOWN_DUAL_API_MODELS = frozenset({
“gemini-3.8-live”,
“gemini-3.8-live-extended-thinking”,
})
_validate_model_api_match() would then skip KNOWN_DUAL_API_MODELS. It might also help to mention in the docs that the Live API on Vertex is not available on global, and that 3.8 Live is served on the eu / us multi-regions rather than single regions like europe-west4.
A related difference you may want to handle
On the Gemini Developer API, gemini-3.8-live rejects thinking_config=ThinkingConfig(thinking_level=…) with 1007 Thinking level is not supported for this model. On Vertex eu, the same config was accepted. If the plugin forwards thinking_config as-is, this difference between the two APIs is worth documenting.
Questions
- Is a fix planned to accept gemini-3.8-live with vertexai=True? We’re happy to open a PR along the lines above if that helps.
- Until it ships, what is the recommended way to use it on Vertex without patching the plugin?
Thanks!