Sensitive content

This media may contain sensitive content.

正在加载视频...

视频加载失败

Hello Gemini 2.5 Flash-Lite! So fast, it codes *each screen* on the fly (Neural OS concept 👇). The frontier isn't always about large models and beating benchmarks. In this case, a super fast & good model can unlock drastic use cases. Read more:

405,746 次观看 • 1 年前 •via X (Twitter)

10 条评论

Serçiya^ سەرچیا 的头像
Serçiya^ سەرچیا1 年前

What are some useful cases for this?

AJ - e/acc 🚀🇨🇴🇵🇷 的头像
AJ - e/acc 🚀🇨🇴🇵🇷1 年前

It's FAST!

HarrisonAIX 的头像
HarrisonAIX1 年前

what's the Gemini VScode extension using nowadays? Gemini 2.5 Flash-Lite? From what I gather, users can't justify ultra and pro is a necessity. The extension being updated would be next level.

Chris 的头像
Chris1 年前

Apps could just be APIs, where the users can just ask for whatever UI they want and the LLMs will build it

Suman 的头像
Suman1 年前

While aiming for better responsiveness with Neural OS ideas, we need to think carefully about how well the model works in different situations, keep user data safe in complex systems, and ensure clear rules for user experiences that change.

Chanchal 💬 的头像
Chanchal 💬1 年前

mind blowing 😲

BlockRadar 的头像
BlockRadar1 年前

Alpha tech, WAGMI

Jesahn O 的头像
Jesahn O1 年前

Best use cases?

Djasnive Rajaona 的头像
Djasnive Rajaona1 年前

@GoogleDeepMind What the heck

Sic Semper Salustri🏴‍☠️ 的头像
Sic Semper Salustri🏴‍☠️1 年前

@GoogleDeepMind 👀

相关视频

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 次观看 • 2 个月前