正在加载视频...

视频加载失败

There is a subtle architecture shift happening in voice AI. The voice stack is becoming part of the agent's execution loop. Cartesia is combining the listening and speaking paths around that loop. Sonic-3.6 turns text into speech (90ms latency) and Ink-2 turns speech into text (100ms transcript latency), faster...

308,266 次观看 • 4 天前 •via X (Twitter)

19 条评论

Chahat Sharma 的头像
Chahat Sharma4 天前

@cartesia tbh i haven't tried cartesia yet. is the 90ms consistent or does it spike under load

Gill 的头像
Gill4 天前

@cartesia Sub 100ms turnaround finally makes interruptions feel natural

Shinka - AI 的头像
Shinka - AI4 天前

@cartesia Voice makes latency a benchmark you can feel, and sub-100ms streaming just became the price of admission.

AI Mastery Guide 的头像
AI Mastery Guide4 天前

@cartesia 90ms latency is insane

0xKachm 的头像
0xKachm4 天前

@cartesia Sub-100ms latency is brutal, but the hidden trade-off is that streaming speed often comes at the cost of prosody and emotional range. the real edge in voice agents isn't just speed - it's making the machine sound less like a machine.

ParkRider 的头像
ParkRider4 天前

@cartesia Sp we get free upgrades to business and flight changes if we use their platform? Sold!

sake 的头像
sake4 天前

@cartesia In voice agents, latency stops being a benchmark number once it starts shaping the conversation.

Prisma Voices 的头像
Prisma Voices3 天前

@cartesia Latency is the whole game on live calls. People forgive a slightly synthetic voice much faster than they forgive a 700 ms pause. The pause is what makes callers say 'hello?' and talk over the agent. Shaving 100 ms off each side makes the interruption problem much smaller.

Kemal Ege Aktemur 的头像
Kemal Ege Aktemur2 天前

@cartesia the 100ms threshold is such a useful framing. once listening and speaking share the execution loop, turn-taking becomes a systems problem, not just a better tts demo.

Fajar M Reza 的头像
Fajar M Reza4 天前

@cartesia Shared listening-speaking loops could reduce latency beyond isolated voice models.

Jason傑森 🇭🇰 | 🛠️ 的头像
Jason傑森 🇭🇰 | 🛠️3 天前

@cartesia 音箱刚更新俩新技术想想都激动。

Ajay Yadav 的头像
Ajay Yadav4 天前

@cartesia Voice AI is becoming part of the agent loop, not just an output layer. Huge shift for real-time agents.

Muhammad usman 的头像
Muhammad usman4 天前

@cartesia Sonic-3.6 and Ink-2 really are #1 on both arenas right now. The latency numbers are vendor-stated model latency though, not full round-trip, still impressive, just a nuance worth knowing. love it man

Dipanshu Kushwaha 的头像
Dipanshu Kushwaha4 天前

@cartesia This is pretty cool! It's exciting to see how voice AI is evolving and becoming more integrated.

Colbert 的头像
Colbert4 天前

@cartesia It is wild how much the latency gap is closing. Integrating the voice stack directly into the agent loop really feels like the missing piece for natural interaction.

Agnes Broad 的头像
Agnes Broad4 天前

@cartesia #1 on their own leaderboard. I don't think so.

nomad.carpenter 的头像
nomad.carpenter4 天前

@cartesia 100ms is the difference between a voice agent feeling turn-based and feeling present; putting both speech directions inside the execution loop matters more than another benchmark point.

Manish | Skygnosis 的头像
Manish | Skygnosis4 天前

90ms is fast enough that the bottleneck stops being the model and becomes everything around it — sip routing, barge-in detection, tool latency. saw a great postmortem this week: only 15% of build time on a phone agent went to conversation quality, the rest was telephony and failure handling. speed just moves where the real cost sits

AI Quanting 的头像
AI Quanting4 天前

@cartesia 90ms on the tts end isnt where a conversation feels slow. most of the gap is endpointing, waiting long enough to be sure the person actually stopped, and thats a few hundred ms you cant just cut without talking over people

相关视频

Learn to build conversational AI voice agents in "Building AI Voice Agents for Production", created in collaboration with LiveKit and RealAvatar, and taught by dsa (Co-founder & CEO of LiveKit), Shayne (Developer Advocate, LiveKit), and Nedelina Teneva (Head of AI at RealAvatar, an AI Fund portfolio company). Voice agents combine speech and reasoning capabilities to enable real-time conversations. They're already being used to support customer service, to improve accessibility in healthcare, for entertainment applications, and for talk therapy. In this course, you’ll learn to build voice agents that listen, reason, and respond naturally. You’ll follow the architecture used to create the "AI Andrew" Avatar, a collaborative project between and RealAvatar that responds to users in what sounds like my voice. You’ll build a voice agent from scratch and deploy it to the cloud, enabling support for many simultaneous users. What you’ll learn: - Understand the fundamentals of voice agents, including key components like speech-to-text (STT), text-to-speech (TTS), and LLMs, and how latency is introduced at each layer. - Explore voice agent architectures and the trade-offs between modular pipelines and speech-to-speech APIs. - Explore how platforms like LiveKit mitigate latency issues with optimized networking infrastructure and low-latency communication protocols. - Learn how to connect client devices to voice agents using WebRTC—and why it outperforms HTTP and WebSocket for low-latency audio streaming. - Incorporate voice activity detection (VAD), end-of-turn detection, and context management to detect turns, handle interruptions, and manage conversational flow. - Understand the trade-offs between latency, quality, and cost in an example in which you build a voice agent and change its voice. - Equip your agent with metrics to measure latency at each stage of the voice pipeline and learn the key levers you can pull to make your agent faster and more responsive. The voice agents built in this course also incorporate voice technology from , a supporting contributor to the project. By the end of this course, you'll have learned the components of an AI voice agent pipeline, combined them into a system with low-latency communication, and deployed them on cloud infrastructure so it scales to many users. I’m looking forward to seeing what voice agents you build from this course! Please sign up here:

Andrew Ng

87,810 次观看 • 1 年前

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

328,306 次观看 • 15 天前

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 次观看 • 3 个月前