Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Monsoon ASR ⚡️ Voice Arena Speech recognition does not have a model problem anymore. It has a data problem. The best ASR systems are approaching human-level performance in English. But move into the long tail of the world's languages, especially real, conversational speech and error rates can still...

564,495 görüntüleme • 17 saat önce •via X (Twitter)

12 Yorum

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

For the last decade, much of ASR progress came from building better models. For the next billion voice users, a huge part of the remaining problem is different: We need better data, in far more languages, representing how people actually speak. That is what we're building with Monsoon. 50 languages today. 100 next by February 2027. And a 1000 languages soon with single-digit semantic WER in every language. Live on VoiceArena. Here's a quick preview of the utterances that make up the Monsoon dataset.

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

The gap is enormous. On FLEURS, ElevenLabs Scribe v2 gets: US English → 4.2% WER Hausa → 19.5% Amharic → 28.2% Somali → 40.0% Yoruba → 41.4% That is 4.7X–9.9X more errors than English. And FLEURS is read speech. Real conversations - accents, interruptions, noise, informal vocabulary, different devices are harder still

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

So we built Monsoon. A large-scale training corpus for real-world ASR across 50 languages and six regions spanning South Asia, Southeast Asia, Africa, Europe, MENA and the Americas, from Bengali, Gujarati and Hausa to Javanese, Arabic dialects and regional Englishes.

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

Each corpus consists of spontaneous, single-speaker audio from two-party conversations, recorded one speaker per channel and built through a rigorous gating and verification protocol: - classifier gating for language, speaker gender and non-speech events - five levels of human transcription, admitted only at ≥ 98% audited accuracy - selection for acoustic and lexical coverage Below: Whisper medium fine-tuned on Monsoon corpus vs commercial ASR.

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

Do the gains stop once you have enough hours? So we trained the same Whisper Medium recipe on: 1,000 h → 2,000 h → ~3,500 h across Hindi, Tamil and Telugu. WER kept falling on every curve. At the full corpus size, none had flattened. For long-tail speech recognition, we are nowhere near done collecting data.

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

How do we build lexical and acoustic diversity? Elicitation. ~5,000 deduplicated topics per language span 35 everyday domains and systematically elicit numbers, procedures, named entities, places and products. Selection. Segments are ranked for lexical rarity and stratified by DNSMOS and duration, preserving noisy and short-form speech as controlled tails rather than filtering them out. Coverage. In Hindi, Monsoon spans 13,731 speakers, 781 districts, 1,123 device models and 120,674 unique word types.

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

We also deliberately do not clean the real world out of the dataset. Short clips, noisy environments and difficult acoustic conditions are retained as a controlled part of the training distribution. On the noisiest subset of our Bengali STT benchmark, the Monsoon fine-tune ranks at 3 of the systems we evaluated, ahead of Gemini 3 Pro, MAI-Transcribe-2 and ElevenLabs Scribe v2. (Voice Arena does not build models for a commercial usecase but for proving the quality of our datasets. If you would like to check the results of this model on your internal benchmarks you can access the API available on the Monsoon webpage on Voice Arena)

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

Monsoon is available in three sizes Small - 500 hours / language Experiment with the data . Medium - 1000 to 3,500 hours / language Promises single digit word error rates Large - Custom - Your model is evaluated, failure modes are identified and a custom dataset is collected to achieve sub 5% WER.

Shobhit Banga profil fotoğrafı
Shobhit Banga17 saat önce

A small number of datasets have defined progress in speech recognition: Switchboard and Fisher for conversational English, LibriSpeech and MLS for read speech at scale, Common Voice and FLEURS for language coverage, IndicVoices for Indic languages. Monsoon builds on that line of legacy datasets that have defined speech recognition. We are proud to bring you 50 languages today and are excited about completing the 1000 languages target by the end of next year. Live now at

AshutoshShrivastava profil fotoğrafı
AshutoshShrivastava17 saat önce

@voicearena_ai Models are no longer the bottleneck. Good data is all you need .

Aryan SMM profil fotoğrafı
Aryan SMM16 saat önce

@voicearena_ai improving real conversational speech across that many languages is a serious data challenge

Artists Voyage 🔶 profil fotoğrafı
Artists Voyage 🔶16 saat önce

@voicearena_ai single digit WER across 50 languages is an ambitious target the long tail language problem is where ASR still has a lot of room to improve

Benzer Videolar

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,715 görüntüleme • 5 ay önce

Putin's famous historic Munich speech in 2007 This is a PHENOMINAL speech. EVERYONE needs to hear it! What prophecies have come true? NATO’s eastward expansion foments tension Putin’s Munich speech: "NATO expansion does not have any relation with the modernization of the Alliance itself, or with ensuring security in Europe. On the contrary, it represents a serious provocation that reduces the level of mutual trust." Democracy by diktat won’t work Putin stressed in his Munch speech: "[The observance of human rights] is an important task. We support this. But this does not mean interfering in the internal affairs of other countries, and especially not imposing a regime that determines how these states should live and develop. It is obvious that such interference does not promote the development of democratic states at all. On the contrary, it makes them dependent and, as a consequence, politically and economically unstable." An arms race will follow Putin stated in his historic Munich speech: "No one can feel that international law is like a stone wall that will protect them. Of course, such a policy stimulates an arms race… The potential danger of the destabilization of international relations is connected with obvious stagnation in the disarmament issue." Unresolved Iranian nuclear problem Putin noted in his historic Munich speech: "If the international community does not find a reasonable solution for resolving this conflict of interests, the world will continue to suffer similar, destabilizing crises… We are going to constantly fight against the threat of the proliferation of weapons of mass destruction." Long-term contracts and energy security Putin stated in his Munich speech: "And now about whether our government cabinet is able to operate responsibly in resolving issues linked to energy deliveries and ensuring energy security. Of course, it can! Moreover, all that we have done and are doing is designed to achieve only one goal, namely to transfer our relations with consumers and countries that transport our energy to market-based, transparent principles and long-term contracts." Unipolar world’s fall Putin said in his Munich speech: "I consider that the unipolar model is not only unacceptable but also impossible in today’s world. The model itself is flawed because at its basis there is and can be no moral foundations for modern civilization. There is no reason to doubt that the economic potential of the new centers of global economic growth will inevitably be converted into political influence and will strengthen multipolarity." From friends to foes Putin said in his Munich speech: "He (U.S. President George W. Bush) says, 'I proceed from the fact that Russia and the USA will never be opponents and enemies again. I agree with him." Another world Putin said in his Munich speech: "Russia is a country with a history that spans more than a thousand years and has practically always used the privilege to carry out an independent foreign policy." NATO’s expansion, a unipolar world, disarmament problems, the erosion of the OSCE as an institution, the Iranian nuclear problem and Europe’s energy security - TASS has summarized Putin’s warnings and prophecies from his Munich speech that have come true simply because nobody turned an attentive ear to them.

Ignorance, the root and stem of all evil

220,401 görüntüleme • 2 yıl önce

QVAC SDK 0.14.0 is live. This release makes the on-device stack faster on mobile, ships the developer-agent path, and takes local text-to-speech to 31 languages. Main highlights: - OpenCode and OpenClaw. The first official OpenCode plugin, plus a maintained OpenClaw compatibility path, both built on managed mode and qvac serve. Point a coding agent at a local model with far less setup and far fewer surprises. - Brain-computer interface transcription, on the SDK. Take recorded neural signal data and decode it into text, fully on-device, no cloud. Stream it in chunks through a simple API. In 0.14 it runs GPU-accelerated on iOS. - Text to Speech in 31 languages with our Supertonic3 upgrade. VOICE AND SPEECH - Supertonic3 multilingual TTS, 5 languages to 31. - Chatterbox and Supertonic now run on the Android GPU, with lower memory use (especially on iOS), quantized s3gen Chatterbox support, and a fix for Chatterbox occasionally emitting random speech. - Whisper transcription now runs on the iOS GPU. Parakeet runs on the Android GPU, with steadier real-time streaming. VISION AND OCR - VLM multi-tile batching: high-resolution Pan and Scan images are encoded in one pass instead of tile by tile, for faster vision throughput. - OCR on ggml (EasyOCR and DocTR) reaches full speed parity with the onnx path, across Metal, OpenCL, and Vulkan. PLATFORM AND RELIABILITY - Dynamic compute backends on Linux: one build picks the right backend at runtime, and opens the door to ROCm and CUDA support without per-backend builds. - Thinking tokens are kept out of the model context, so reasoning no longer fills the KV cache. SDK 0.14.0 is now leaner and faster to start. Let’s build.

QVAC

23,995,874 görüntüleme • 3 ay önce