Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Build voice agents that can respond quickly or think longer without changing your stack. Gemini 3.8 Live is available today in LiveKit Agents. Choose gemini-3.8-live for low-latency audio or gemini-3.8-live-extended-thinking for async reasoning. Check the docs for more info:

18,003 Aufrufe • vor 3 Tagen •via X (Twitter)

11 Kommentare

Profilbild von Thor 雷神 ⚡️
Thor 雷神 ⚡️vor 3 Tagen

Awesome way of visualising the difference! @codeSTACKr's demos are always fire 🔥 Thanks so much for shipping on Day0 🚀

Profilbild von Patrick Loeber
Patrick Loebervor 3 Tagen

🔥🔥

Profilbild von Ronit Jain
Ronit Jainvor 3 Tagen

Fast and deep reasoning in one stack is the right split. latency budgets finally get a product knob.

Profilbild von A.W.E.S.O.M.-O 4000
A.W.E.S.O.M.-O 4000vor 3 Tagen

Having both low latency and extended thinking options without changing the stack is a great flexibility win

Profilbild von Mira Synth Tech
Mira Synth Techvor 3 Tagen

Great flexibility fast or deep thinking, same stack.

Profilbild von Matt @ Stepwright
Matt @ Stepwrightvor 3 Tagen

This is awesome! Thanks for the demo :)

Profilbild von Fajar M Reza
Fajar M Rezavor 3 Tagen

Voice agents should separate speaking state from action state during tool calls.

Profilbild von Mitchell Agoma
Mitchell Agomavor 3 Tagen

Keeping the stack fixed should make this a useful regression comparison. I would separate the silence timeout from the tool-execution deadline, then test both models with a slow tool and a caller correction. A quiet audio stream must not trigger a duplicate task or premature success. My detailed QA approach:

Profilbild von AI Mastery Guide
AI Mastery Guidevor 3 Tagen

extended thinking for voice agents??

Profilbild von saietta
saiettavor 3 Tagen

the interesting question is whether you can switch modes mid-call or it's locked in per session. real conversations don't sort cleanly into fast vs slow upfront, the hard part is deciding on the fly whether this specific question is worth eating the extended-thinking latency for.

Profilbild von k0n00
k0n00vor 3 Tagen

On a phone line turnComplete is not idle. Extended thinking keeps speaking fillers while interaction_status stays in progress, so the next generation still needs to know the turn is not over.

Ähnliche Videos