Video yükleniyor...
Video Yüklenemedi
MICROSOFT OPEN SOURCED A 7B PARAMETER MODEL THAT TRANSCRIBES 60 MINUTES OF AUDIO IN A SINGLE PASS and it's completely free VIBEVOICE ASR no chunking, no context loss, full speaker diarization baked in not just speech to text..not a basic wrapper who spoke, when they spoke, exactly what they... show more
1,375,300 görüntüleme • 5 ay önce •via X (Twitter)
32 Yorum

MIT means free code, not free compute. Self-hosting a 7B ASR means a GPU instance running 24/7. That's almost certainly more $/mo than a Deepgram bill. There's no Microsoft-hosted free API tier here.

“Free” Who owns the data now? Who can parse and use that content now? Who can leverage that content now? Yeah… “free” is always a Trojan horse.

Use groq with whisper large v3 turbo for only 0,04$ per hour, hours of audio transcribed in a couple of seconds!!

Diarization baked in is cool, but 7B parameters kinda defeats the point. You can’t roll this out on device. Parakeet at 0.6b or Granite 1b are a better fit for os local

Whisper models are also free to run on your machine

we’re entering the phase where: speech becomes a solved layer the value won’t be transcription it’ll be what you do with it after

i love free open sourced models >3 applause to Microsoft

that’s actually a big drop

oh wow i had been using elevenlabs API inside of kilo for some work but it was slightly costly. I am going to try this out. 7B is not a lot to run locally

the 'no chunking' part is the real unlock. tested 3 ASR pipelines for a video workflow last quarter and chunked diarization always lost ~6-8% speaker accuracy when voices swapped at the chunk boundary. how does VibeVoice deal with overlapping speakers in the same window?

funny how every week there's a revolutonary asr model that supposedly fixes everything at once if it actually nails long form audio and diarization cleanly then thats genuinely impressive but claims like that usually age weirdly

Or you can use typrapp with whisper large v3 turbo running locally for free and a nice ui on top

Nah, they are data farming peoples voices and you all are willingly giving it to them.

It's not completely free...

single pass is cool but diarization quality will decide if this actually ships

Hmm single pass sounds good but 7b is pretty big for local transcription… be curious to see if quantized versions perform well.

@perssua @lucas_montano

ignoring the caps lock hype, the actual signal here isn't the 60 minute context window. it's the native diarization.

That’s ground-breaking!

Wow,

Microsoft just collapsed half the ASR stack. One-pass transcription + built-in diarization = no chunking, no glue code. Paired with Ollama → fully local, private voice pipelines. Free isn’t the headline. Owning the stack is.

wen @AskVenice

Nice!

Es un cambio de rumbo!

I use SubtitleEdit and use the audio to text (whisper).

is it any good?

17GB ASR Speech-to-Text model? Are you kidding me? I will run a Film Production Pipelines on 12GB GPU, that would generate the Images, The videos and also the Conversations And also the Music (the score). Microsoft Speech-to-text... ROFL - Get Out Of Here. Go fix your Windows MS

When you say 🆓 what does it mean?

Don’t just open source the weights, host it for us for free! No one normal person can afford the GPUs needed for this.

Is it available on

Hm

We are entering an era where the most successful researchers won't be the best coders, but the best "Product Managers" of their own ideas
