Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We're releasing Voice of Reason, a speech-native model that does math out loud. Give it a spoken problem without transcription nor text LLM in the loop, and it reasons and answers in speech. GSM8K goes from 27.3% for GLM-4-Voice to 77.1%. Link in 🧵

19,297 Aufrufe • vor 16 Tagen •via X (Twitter)

10 Kommentare

Profilbild von kyutai
kyutaivor 16 Tagen

Math is a good benchmark for models’ intelligence but speech models typically lag behind text ones. Speech models have to answer out loud immediately, but cascading around a text model (transcribe, think, speak) adds too much latency. How do we bring them up to par ?

Profilbild von kyutai
kyutaivor 16 Tagen

We start from GLM-4-Voice, finetune it on synthesized spoken question answering, then post-train it with RL against a verifiable reward. The model answers straight-out, with no extra reasoning tokens at all, and reaches 70.3% accuracy, above some past speech models that needed explicit reasoning traces.

Profilbild von kyutai
kyutaivor 16 Tagen

We can further enable thinking without speaking using the streaming reasoning scheme of STITCH (Chiang et al., The model reasons in silent chunks while its previous sentence is still playing: 77.1% on GSM8K, without added thinking latency.

Profilbild von kyutai
kyutaivor 16 Tagen

Paper: (COLM 2026) Weights: directly answers. for extra reasoning tokens while speaking.

Profilbild von Deep
Deepvor 16 Tagen

i build voice agents and the transcription step always loses something. doing it natively is the right call.

Profilbild von sentient being
sentient beingvor 16 Tagen

kyutai always working on impressive stuff!!!

Profilbild von unfair.so
unfair.sovor 15 Tagen

reasoning out loud

Profilbild von Mathias Heinke
Mathias Heinkevor 16 Tagen

Gewichte offen als glm-4-voice-of-reason-stitch-9b, Lizenz glm-4-voice-license, Basis GLM-4-Voice-9B. Laut rechnen ohne Textmodell dazwischen: ich hätte nicht gedacht, dass das schon geht 🫡

Profilbild von Audemy
Audemyvor 16 Tagen

gsm8k answers are plain numbers, so each one has exactly one reading. the hard part for audio-first math is the other direction: "one over x plus y" is two different fractions, and speech has to commit to something like explicit end-fraction markers.

Profilbild von Stephen
Stephenvor 15 Tagen

STITCH turns speech's serial nature into free compute — reasoning in silent chunks while the previous sentence plays. The open question: non-verifiable tasks, where speaking forces commitment before thinking finishes.

Ähnliche Videos