Loading video...
Video Failed to Load
There is a subtle architecture shift happening in voice AI. The voice stack is becoming part of the agent's execution loop. Cartesia is combining the listening and speaking paths around that loop. Sonic-3.6 turns text into speech (90ms latency) and Ink-2 turns speech into text (100ms transcript latency), faster... show more
308,266 views • 4 days ago •via X (Twitter)
19 Comments

@cartesia tbh i haven't tried cartesia yet. is the 90ms consistent or does it spike under load

@cartesia Sub 100ms turnaround finally makes interruptions feel natural

@cartesia Voice makes latency a benchmark you can feel, and sub-100ms streaming just became the price of admission.

@cartesia 90ms latency is insane

@cartesia Sub-100ms latency is brutal, but the hidden trade-off is that streaming speed often comes at the cost of prosody and emotional range. the real edge in voice agents isn't just speed - it's making the machine sound less like a machine.

@cartesia Sp we get free upgrades to business and flight changes if we use their platform? Sold!

@cartesia In voice agents, latency stops being a benchmark number once it starts shaping the conversation.

@cartesia Latency is the whole game on live calls. People forgive a slightly synthetic voice much faster than they forgive a 700 ms pause. The pause is what makes callers say 'hello?' and talk over the agent. Shaving 100 ms off each side makes the interruption problem much smaller.

@cartesia the 100ms threshold is such a useful framing. once listening and speaking share the execution loop, turn-taking becomes a systems problem, not just a better tts demo.

@cartesia Shared listening-speaking loops could reduce latency beyond isolated voice models.

@cartesia 音箱刚更新俩新技术想想都激动。

@cartesia Voice AI is becoming part of the agent loop, not just an output layer. Huge shift for real-time agents.

@cartesia Sonic-3.6 and Ink-2 really are #1 on both arenas right now. The latency numbers are vendor-stated model latency though, not full round-trip, still impressive, just a nuance worth knowing. love it man

@cartesia This is pretty cool! It's exciting to see how voice AI is evolving and becoming more integrated.

@cartesia It is wild how much the latency gap is closing. Integrating the voice stack directly into the agent loop really feels like the missing piece for natural interaction.

@cartesia #1 on their own leaderboard. I don't think so.

@cartesia 100ms is the difference between a voice agent feeling turn-based and feeling present; putting both speech directions inside the execution loop matters more than another benchmark point.

90ms is fast enough that the bottleneck stops being the model and becomes everything around it — sip routing, barge-in detection, tool latency. saw a great postmortem this week: only 15% of build time on a phone agent went to conversation quality, the rest was telephony and failure handling. speed just moves where the real cost sits

@cartesia 90ms on the tts end isnt where a conversation feels slow. most of the gap is endpointing, waiting long enough to be sure the person actually stopped, and thats a few hundred ms you cant just cut without talking over people


