Video wird geladen...
Video konnte nicht geladen werden
We're releasing Voice of Reason, a speech-native model that does math out loud. Give it a spoken problem without transcription nor text LLM in the loop, and it reasons and answers in speech. GSM8K goes from 27.3% for GLM-4-Voice to 77.1%. Link in 🧵
19,297 Aufrufe • vor 16 Tagen •via X (Twitter)
10 Kommentare

Math is a good benchmark for models’ intelligence but speech models typically lag behind text ones. Speech models have to answer out loud immediately, but cascading around a text model (transcribe, think, speak) adds too much latency. How do we bring them up to par ?

We start from GLM-4-Voice, finetune it on synthesized spoken question answering, then post-train it with RL against a verifiable reward. The model answers straight-out, with no extra reasoning tokens at all, and reaches 70.3% accuracy, above some past speech models that needed explicit reasoning traces.

We can further enable thinking without speaking using the streaming reasoning scheme of STITCH (Chiang et al., The model reasons in silent chunks while its previous sentence is still playing: 77.1% on GSM8K, without added thinking latency.

Paper: (COLM 2026) Weights: directly answers. for extra reasoning tokens while speaking.

i build voice agents and the transcription step always loses something. doing it natively is the right call.

kyutai always working on impressive stuff!!!

reasoning out loud

Gewichte offen als glm-4-voice-of-reason-stitch-9b, Lizenz glm-4-voice-license, Basis GLM-4-Voice-9B. Laut rechnen ohne Textmodell dazwischen: ich hätte nicht gedacht, dass das schon geht 🫡

gsm8k answers are plain numbers, so each one has exactly one reading. the hard part for audio-first math is the other direction: "one over x plus y" is two different fractions, and speech has to commit to something like explicit end-fraction markers.

STITCH turns speech's serial nature into free compute — reasoning in silent chunks while the previous sentence plays. The open question: non-verifiable tasks, where speaking forces commitment before thinking finishes.

