Loading video...

Video Failed to Load

Go Home

This new 2B open-source speech model transcribes in real time with 200ms latency 🤯 It’s called Confucius R2T2 by NetEase Youdao and commits each word the moment it's recognized and never rewrites it. 100% open source.

31,677 views • 3 days ago •via X (Twitter)

18 Comments

Alvaro Cintas's profile picture
Alvaro Cintas3 days ago

Almost every streaming ASR system is built to emit something at every step. When the audio is ambiguous, it falls back on language priors and produces its best guess, because staying silent isn't really an option in the architecture. R2T2 introduces an explicit wait state instead. If there isn't enough acoustic evidence yet, it holds. When evidence arrives, it commits.

Alvaro Cintas's profile picture
Alvaro Cintas3 days ago

What I find interesting is how this reframes the latency conversation. The field keeps chasing "predict faster." This design argues the harder skill is knowing when a prediction is premature and would cost you more downstream than a short pause. Link: Open source, vLLM backend. Built for voice agents.

Alvaro Cintas's profile picture
Alvaro Cintas2 days ago

Here are the links to the repo: and Hugging Face:

FluidVoice's profile picture
FluidVoice2 days ago

Awesome find! Benchmarking it to see if it's a good fit for Fluidvoice and my users ;)

Swipecraft's profile picture
Swipecraft2 days ago

What is the inference cost ? we have already NLLB model there right ? so what is the special about this model ?

Qamar Rizwani's profile picture
Qamar Rizwani3 days ago

Nice real time transcription, @Build_Vertical.

triet.bui's profile picture
triet.bui2 days ago

Wow, thanks for the information 🤯

Gregor's profile picture
Gregor3 days ago

the 'never rewrites' claim is the real tradeoff here. most streaming models correct earlier words as context rolls in. genuinely asking how WER compares to Whisper on phonetically ambiguous audio

Ali Mirza | AI Agents & Automation's profile picture
Ali Mirza | AI Agents & Automation3 days ago

200ms immutable token emission cuts voice agent turn-taking latency in half.

Manu_TechAndGames's profile picture
Manu_TechAndGames2 days ago

Not completely open source. Companies with 1B revenues need a specific license.

Aaron's profile picture
Aaron3 days ago

damn that's awesome, trying it right away

Muhammad Ayan's profile picture
Muhammad Ayan2 days ago

the wait state is the clever bit guessing less can beat speaking sooner

💧🌞John Troughton's profile picture
💧🌞John Troughton3 days ago

How can we get it?

Alvaro Cintas's profile picture
Alvaro Cintas3 days ago

Here:

Thach Nguyen's profile picture
Thach Nguyen3 days ago

big if true, esp for ai voice agents. how is it compared to deepgram in terms of accuracy?

James Camarota's profile picture
James Camarota2 days ago

Committing each word without rewriting is a bold design. How does it handle names or phrases that only become clear from later context?

Abat's profile picture
Abat2 days ago

That's great!

Max Bevza's profile picture
Max Bevza2 days ago

finally an open source model that actually stays bullish

Related Videos