Loading video...

Video Failed to Load

Go Home

MICROSOFT OPEN SOURCED A 7B PARAMETER MODEL THAT TRANSCRIBES 60 MINUTES OF AUDIO IN A SINGLE PASS and it's completely free VIBEVOICE ASR no chunking, no context loss, full speaker diarization baked in not just speech to text..not a basic wrapper who spoke, when they spoke, exactly what they...

1,375,300 views • 5 months ago •via X (Twitter)

32 Comments

Seth's profile picture
Seth5 months ago

MIT means free code, not free compute. Self-hosting a 7B ASR means a GPU instance running 24/7. That's almost certainly more $/mo than a Deepgram bill. There's no Microsoft-hosted free API tier here.

Stef's profile picture
Stef5 months ago

“Free” Who owns the data now? Who can parse and use that content now? Who can leverage that content now? Yeah… “free” is always a Trojan horse.

timdrchslr's profile picture
timdrchslr5 months ago

Use groq with whisper large v3 turbo for only 0,04$ per hour, hours of audio transcribed in a couple of seconds!!

Nomad's profile picture
Nomad5 months ago

Diarization baked in is cool, but 7B parameters kinda defeats the point. You can’t roll this out on device. Parakeet at 0.6b or Granite 1b are a better fit for os local

EventsLooped's profile picture
EventsLooped5 months ago

Whisper models are also free to run on your machine

VC Intern's profile picture
VC Intern5 months ago

we’re entering the phase where: speech becomes a solved layer the value won’t be transcription it’ll be what you do with it after

Alpha Batcher's profile picture
Alpha Batcher5 months ago

i love free open sourced models >3 applause to Microsoft

Julian Goldie SEO's profile picture
Julian Goldie SEO5 months ago

that’s actually a big drop

Sanchit's profile picture
Sanchit5 months ago

oh wow i had been using elevenlabs API inside of kilo for some work but it was slightly costly. I am going to try this out. 7B is not a lot to run locally

tang | AI Product Maker's profile picture
tang | AI Product Maker5 months ago

the 'no chunking' part is the real unlock. tested 3 ASR pipelines for a video workflow last quarter and chunked diarization always lost ~6-8% speaker accuracy when voices swapped at the chunk boundary. how does VibeVoice deal with overlapping speakers in the same window?

Quadcode AI's profile picture
Quadcode AI5 months ago

funny how every week there's a revolutonary asr model that supposedly fixes everything at once if it actually nails long form audio and diarization cleanly then thats genuinely impressive but claims like that usually age weirdly

juaniam's profile picture
juaniam5 months ago

Or you can use typrapp with whisper large v3 turbo running locally for free and a nice ui on top

flamz's profile picture
flamz5 months ago

Nah, they are data farming peoples voices and you all are willingly giving it to them.

Layton Gott's profile picture
Layton Gott5 months ago

It's not completely free...

Sebastian Buzdugan's profile picture
Sebastian Buzdugan5 months ago

single pass is cool but diarization quality will decide if this actually ships

Max Bourke's profile picture
Max Bourke5 months ago

Hmm single pass sounds good but 7b is pretty big for local transcription… be curious to see if quantized versions perform well.

Victor Tavernari 's profile picture
Victor Tavernari 5 months ago

@perssua @lucas_montano

Eric's profile picture
Eric4 months ago

ignoring the caps lock hype, the actual signal here isn't the 60 minute context window. it's the native diarization.

CODIFY's profile picture
CODIFY5 months ago

That’s ground-breaking!

Ben Dickey's profile picture
Ben Dickey5 months ago

Wow,

Cyril Gupta's profile picture
Cyril Gupta5 months ago

Microsoft just collapsed half the ASR stack. One-pass transcription + built-in diarization = no chunking, no glue code. Paired with Ollama → fully local, private voice pipelines. Free isn’t the headline. Owning the stack is.

Goldengmgm🐸 Яebel 🪖's profile picture
Goldengmgm🐸 Яebel 🪖5 months ago

wen @AskVenice

Maxwell de Melo's profile picture
Maxwell de Melo5 months ago

Nice!

Facundo Campo · Dev's profile picture
Facundo Campo · Dev5 months ago

Es un cambio de rumbo!

DiffusionTales's profile picture
DiffusionTales5 months ago

I use SubtitleEdit and use the audio to text (whisper).

Adam Jaworski's profile picture
Adam Jaworski5 months ago

is it any good?

AIGoldenfinger's profile picture
AIGoldenfinger5 months ago

17GB ASR Speech-to-Text model? Are you kidding me? I will run a Film Production Pipelines on 12GB GPU, that would generate the Images, The videos and also the Conversations And also the Music (the score). Microsoft Speech-to-text... ROFL - Get Out Of Here. Go fix your Windows MS

⭕💡's profile picture
⭕💡5 months ago

When you say 🆓 what does it mean?

Terry Nunn's profile picture
Terry Nunn5 months ago

Don’t just open source the weights, host it for us for free! No one normal person can afford the GPUs needed for this.

Tabish Gillani's profile picture
Tabish Gillani5 months ago

Is it available on

Rey Neill's profile picture
Rey Neill5 months ago

Hm

AI's Nest's profile picture
AI's Nest5 months ago

We are entering an era where the most successful researchers won't be the best coders, but the best "Product Managers" of their own ideas

Related Videos

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,715 views • 5 months ago

NVIDIA JUST DROPPED A FREE AI MODEL THAT READS PDFS, WATCHES VIDEOS, LISTENS TO AUDIO, AND UNDERSTANDS YOUR SCREEN SIMULTANEOUSLY. Not one at a time. ALL AT ONCE. In a single pass. It is called Nemotron 3 Nano Omni and it runs 9 times faster than every other multimodal model currently available. Think about what that actually means for how you work. Right now you are switching between tools constantly. One tool for transcribing your call recordings. A different tool for analyzing your client PDFs. Another tool for processing your training videos. A separate workflow for understanding what is happening on your screen. Four tools. Four contexts. Four different outputs you have to manually synthesize into one decision. Nemotron 3 Nano Omni does all of it in one model. One pass. One output. The use cases that just got dramatically simpler: Meeting recordings where you need the transcript, the visual context, and the document references all analyzed together. Training videos where the audio, the slides, and the on-screen demonstrations all feed into one coherent summary. Client PDFs where you need the document content cross-referenced against your screen data and your call notes simultaneously. Sales call transcripts analyzed alongside the proposals and the CRM data in one unified pass. This is not a marginal improvement on existing multimodal models. It is a 9x speed increase on a capability that was already changing how people work. Free. From NVIDIA. Available right now. Bookmark this before everyone catches on. Follow CyrilXBT for every AI capability shift the moment it drops.

CyrilXBT

37,847 views • 5 months ago