Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

This new 2B open-source speech model transcribes in real time with 200ms latency 🤯 It’s called Confucius R2T2 by NetEase Youdao and commits each word the moment it's recognized and never rewrites it. 100% open source.

31,677 Aufrufe • vor 3 Tagen •via X (Twitter)

18 Kommentare

Profilbild von Alvaro Cintas
Alvaro Cintasvor 3 Tagen

Almost every streaming ASR system is built to emit something at every step. When the audio is ambiguous, it falls back on language priors and produces its best guess, because staying silent isn't really an option in the architecture. R2T2 introduces an explicit wait state instead. If there isn't enough acoustic evidence yet, it holds. When evidence arrives, it commits.

Profilbild von Alvaro Cintas
Alvaro Cintasvor 3 Tagen

What I find interesting is how this reframes the latency conversation. The field keeps chasing "predict faster." This design argues the harder skill is knowing when a prediction is premature and would cost you more downstream than a short pause. Link: Open source, vLLM backend. Built for voice agents.

Profilbild von Alvaro Cintas
Alvaro Cintasvor 2 Tagen

Here are the links to the repo: and Hugging Face:

Profilbild von FluidVoice
FluidVoicevor 2 Tagen

Awesome find! Benchmarking it to see if it's a good fit for Fluidvoice and my users ;)

Profilbild von Swipecraft
Swipecraftvor 3 Tagen

What is the inference cost ? we have already NLLB model there right ? so what is the special about this model ?

Profilbild von Qamar Rizwani
Qamar Rizwanivor 3 Tagen

Nice real time transcription, @Build_Vertical.

Profilbild von triet.bui
triet.buivor 2 Tagen

Wow, thanks for the information 🤯

Profilbild von Gregor
Gregorvor 3 Tagen

the 'never rewrites' claim is the real tradeoff here. most streaming models correct earlier words as context rolls in. genuinely asking how WER compares to Whisper on phonetically ambiguous audio

Profilbild von Ali Mirza | AI Agents & Automation
Ali Mirza | AI Agents & Automationvor 3 Tagen

200ms immutable token emission cuts voice agent turn-taking latency in half.

Profilbild von Manu_TechAndGames
Manu_TechAndGamesvor 2 Tagen

Not completely open source. Companies with 1B revenues need a specific license.

Profilbild von Aaron
Aaronvor 3 Tagen

damn that's awesome, trying it right away

Profilbild von Muhammad Ayan
Muhammad Ayanvor 3 Tagen

the wait state is the clever bit guessing less can beat speaking sooner

Profilbild von 💧🌞John Troughton
💧🌞John Troughtonvor 3 Tagen

How can we get it?

Profilbild von Alvaro Cintas
Alvaro Cintasvor 3 Tagen

Here:

Profilbild von Thach Nguyen
Thach Nguyenvor 3 Tagen

big if true, esp for ai voice agents. how is it compared to deepgram in terms of accuracy?

Profilbild von James Camarota
James Camarotavor 2 Tagen

Committing each word without rewriting is a bold design. How does it handle names or phrases that only become clear from later context?

Profilbild von Abat
Abatvor 2 Tagen

That's great!

Profilbild von Max Bevza
Max Bevzavor 3 Tagen

finally an open source model that actually stays bullish

Ähnliche Videos