正在加载视频...
视频加载失败
Introducing Monsoon ASR ⚡️ Voice Arena Speech recognition does not have a model problem anymore. It has a data problem. The best ASR systems are approaching human-level performance in English. But move into the long tail of the world's languages, especially real, conversational speech and error rates can still... show more
12 条评论

For the last decade, much of ASR progress came from building better models. For the next billion voice users, a huge part of the remaining problem is different: We need better data, in far more languages, representing how people actually speak. That is what we're building with Monsoon. 50 languages today. 100 next by February 2027. And a 1000 languages soon with single-digit semantic WER in every language. Live on VoiceArena. Here's a quick preview of the utterances that make up the Monsoon dataset.

The gap is enormous. On FLEURS, ElevenLabs Scribe v2 gets: US English → 4.2% WER Hausa → 19.5% Amharic → 28.2% Somali → 40.0% Yoruba → 41.4% That is 4.7X–9.9X more errors than English. And FLEURS is read speech. Real conversations - accents, interruptions, noise, informal vocabulary, different devices are harder still

So we built Monsoon. A large-scale training corpus for real-world ASR across 50 languages and six regions spanning South Asia, Southeast Asia, Africa, Europe, MENA and the Americas, from Bengali, Gujarati and Hausa to Javanese, Arabic dialects and regional Englishes.

Each corpus consists of spontaneous, single-speaker audio from two-party conversations, recorded one speaker per channel and built through a rigorous gating and verification protocol: - classifier gating for language, speaker gender and non-speech events - five levels of human transcription, admitted only at ≥ 98% audited accuracy - selection for acoustic and lexical coverage Below: Whisper medium fine-tuned on Monsoon corpus vs commercial ASR.

Do the gains stop once you have enough hours? So we trained the same Whisper Medium recipe on: 1,000 h → 2,000 h → ~3,500 h across Hindi, Tamil and Telugu. WER kept falling on every curve. At the full corpus size, none had flattened. For long-tail speech recognition, we are nowhere near done collecting data.

How do we build lexical and acoustic diversity? Elicitation. ~5,000 deduplicated topics per language span 35 everyday domains and systematically elicit numbers, procedures, named entities, places and products. Selection. Segments are ranked for lexical rarity and stratified by DNSMOS and duration, preserving noisy and short-form speech as controlled tails rather than filtering them out. Coverage. In Hindi, Monsoon spans 13,731 speakers, 781 districts, 1,123 device models and 120,674 unique word types.

We also deliberately do not clean the real world out of the dataset. Short clips, noisy environments and difficult acoustic conditions are retained as a controlled part of the training distribution. On the noisiest subset of our Bengali STT benchmark, the Monsoon fine-tune ranks at 3 of the systems we evaluated, ahead of Gemini 3 Pro, MAI-Transcribe-2 and ElevenLabs Scribe v2. (Voice Arena does not build models for a commercial usecase but for proving the quality of our datasets. If you would like to check the results of this model on your internal benchmarks you can access the API available on the Monsoon webpage on Voice Arena)

Monsoon is available in three sizes Small - 500 hours / language Experiment with the data . Medium - 1000 to 3,500 hours / language Promises single digit word error rates Large - Custom - Your model is evaluated, failure modes are identified and a custom dataset is collected to achieve sub 5% WER.

A small number of datasets have defined progress in speech recognition: Switchboard and Fisher for conversational English, LibriSpeech and MLS for read speech at scale, Common Voice and FLEURS for language coverage, IndicVoices for Indic languages. Monsoon builds on that line of legacy datasets that have defined speech recognition. We are proud to bring you 50 languages today and are excited about completing the 1000 languages target by the end of next year. Live now at

@voicearena_ai Models are no longer the bottleneck. Good data is all you need .

@voicearena_ai improving real conversational speech across that many languages is a serious data challenge

@voicearena_ai single digit WER across 50 languages is an ambitious target the long tail language problem is where ASR still has a lot of room to improve
