Загрузка видео...
Не удалось загрузить видео
🔈 today we're introducing two new live dialogue models Gemini 3.8 Live: built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking: built for high-complexity tasks, with increased intelligence and multi-step reasoning. these audio models are available to build... show more
65,207 просмотров • 1 день назад •via X (Twitter)
Комментарии: 30

We're so excited to bring these models to our customers! Please try and share your feedback

Why the release data is 20 August 2026 in both models ??

Sharing screen freezes the voice conversation.

Could you make Gemini Live's voice sound more natural? It sounds very robotic compared to Chatgpt's voices.

Trying the live dialogue models in Studio.

Google will release Gemini Live (just to pass Open AI on something) Gemini Flash Lite (for uses no one asked for) Gemini Cyber (because Dario is hyping cyberfears) before releasing Gemini Pro. which is a same because Gemini pro models used to be frontier tier

中文支持如何

Live dialogue plus visual grounding is a powerful combination.

Great. I wonder how does this model perform in foreign language accuracy

Would be handy to compare both on the same conversation, including cost and response time.

No frontier model from google yet... really looking for a model thats more efficient in 3d and computer use than astra... as a team gemini user, really hoping this wait is worth it...

NOTICE : do not waste your time fucking with this. It doesn't even work. @GoogleAIStudio ; [Notice: gemini-3.8-flash is at high demand (503). Auto-falling back to gemini-3.5-flash...]

This sounds terrible.

I think the development or app creation section should be updated with (🟢🟡🔴)

Gemini的视觉理解力依然傲视群雄

很赞

is it a full duplex model?

Gemini 3.8 Live just showed up in AI Studio live dialogue. going to burn a call on it tonight and see if latency feels different from the old live stack.

The async function-calling detail is the real interface shift. A voice model that can acknowledge, work, and return with grounded context feels less like a turn-based chatbot and more like a task runner. Latency communication becomes part of trust.

Gr8

The Live vs Extended Thinking split is clearer here than in the consumer posts. One for cheap dialogue, one for calls that actually need to think mid-conversation.

Is Gemini 3.8 live extended thinking also more expensive at API cost then normal Gemini 3.8 live ?

I can finally build a voice assistant that understands my screen without draining my API credits in a single afternoon.

the launch pitch says fluid dialogue, but the thread is already reporting freezes during screen sharing. reliability in the awkward edge cases will decide this product.

two at once is huge

Google’s new voice models don’t pause to “think.” Gemini 3.8 Live reasons while it talks and runs tools in the background. Extended Thinking just took #1 on speech-to-speech. If voice agents finally work, which product dies first - phone trees or meeting notes?

would love for high quality, cheap TTS. pleaaaassseeee

Please make it work in Faroese. We really need a good transcription model to help handicapped people in our country.

what's the interruption latency, that's the real live-audio metric

The split is the interesting bit. In a live audio model, extended thinking means the model goes quiet before it speaks — you spend latency to buy deliberation, and nothing gives you both. So the two SKUs are really one dial: how long will a user wait before a voice comes back?



