Video wird geladen...
Video konnte nicht geladen werden
Introducing our most advanced Gemini Audio models yet 🗣 Gemini 3.8 Live and 3.8 Live Extended Thinking let you speak, collaborate, and execute tasks seamlessly, meaning conversing with AI just got a lot more natural. So, what’s the difference between these two models? Let’s break it down: — Gemini... show more
224,698 Aufrufe • vor 2 Tagen •via X (Twitter)
27 Kommentare

Gemini 3.8 Live is rolling out to: — Consumers: in Search Live — Developers: in public preview in the Gemini API via @googleaistudio — Enterprises: in private preview via Gemini Enterprise (coming soon to Gemini Enterprise Customer Experience) Gemini 3.8 Live Extended Thinking is rolling out to: — Consumers: in Gemini Live in the @GeminiApp, plus Google AI Pro and Ultra subscribers in @GoogleWorkspace in @GoogleDocs, and all Google AI subscribers in @gmail and Keep — Developers: in public preview in the Gemini API via @GoogleAIStudio — Enterprises: in private preview via Gemini Enterprise (coming soon to Gemini Enterprise Customer Experience and @GoogleWorkspace business customers)

No one wants this. We want a model with intelligence and agentic ability that can rival Fable or Sol/Astra. You're falling behind.

You've everything but why you're losing ai race?

Excited to use it

we are going to have 3-4 ChatGPT moments before xmas

Fix 3.8 Flash speed....

@GoogleAI intrigued by the new models! Curious about how they handle real-time tasks without lag.

Will try this

would please, for the love of FSM, bring them to Google Chat and Google Meet?

Real-time DIY guidance just by pointing your camera at the problem is 🔥 This is where voice AI gets really practical.

🚢🚢🚢

the actual product bet is "point your camera at a broken bike chain and get audio instructions," that's a much harder real-world use case than text chat, it requires real-time visual grounding plus natural interruption handling plus enough domain knowledge to give correct repair steps, not just describe what it sees. if this works reliably, it's a genuinely useful category (hands-busy, eyes-busy tasks), if it doesn't, it's a demo that falls apart the moment the pipe leak isn't the exact one in the training data

Regarding Story Understanding Gemini 3.8 Flash is actually better then current Opus.

Your timing is so ... Google. Everyone is watching @bot - is the model so bad you had to bury it?

Useful split. I’d evaluate these on the messy middle, not the demo: interruption recovery, instruction fidelity after a context switch, and time-to-completed action. For marketing ops, “sounds natural” matters less than finishing the workflow without a human reset.

Repair Gemini 3.8 Flash speed instead!!!

Switching between 97 languages mid-conversation sounds impressive. Curious how much latency Extended Thinking adds when it's reasoning and talking at the same time.

Google if you need help blink twice

Voice is still the interface most people actually use.

“Reasoning and speaking in parallel” is impressive. But an agent narrating its progress is still the agent telling us what it did. For consequential tasks, narration isn’t evidence. Autonomous execution needs an independent verification layer.

Real-time voice + vision + reasoning is a powerful combination. This feels much closer to having a truly useful AI assistant.

Try to remove ai slop(in designing) and ai hallucinations in upcoming gemini 4 / 4 pro models

@CodeeNCX - Some additional help with that project 😀

The shift from latent inference to parallelized, narrated reasoning signals a transition toward agentic workflows. We are moving from mere chat interfaces to persistent, multimodal system executors that redefine human-computer...

Voice + video matters when the agent distinguishes what it actually saw from what it inferred. In a repair, an interruption can change the plan mid-step. Can developers inspect that observation trail in the Live API, or is it hidden inside the conversation?

Gemini 3.8 Live’s mid-sentence barge-in across 97 languages is the hard part. Extended Thinking narrating progress turns latency into UX. Audio finally feels like an interface, not a prompt.

How different are the latency and interruption behaviors between Live and Live Extended Thinking in real conversations?


