正在加载视频...

视频加载失败

MICROSOFT OPEN SOURCED A 7B PARAMETER MODEL THAT TRANSCRIBES 60 MINUTES OF AUDIO IN A SINGLE PASS and it's completely free VIBEVOICE ASR no chunking, no context loss, full speaker diarization baked in not just speech to text..not a basic wrapper who spoke, when they spoke, exactly what they...

1,375,300 次观看 • 5 个月前 •via X (Twitter)

32 条评论

Seth 的头像
Seth5 个月前

MIT means free code, not free compute. Self-hosting a 7B ASR means a GPU instance running 24/7. That's almost certainly more $/mo than a Deepgram bill. There's no Microsoft-hosted free API tier here.

Stef 的头像
Stef5 个月前

“Free” Who owns the data now? Who can parse and use that content now? Who can leverage that content now? Yeah… “free” is always a Trojan horse.

timdrchslr 的头像
timdrchslr5 个月前

Use groq with whisper large v3 turbo for only 0,04$ per hour, hours of audio transcribed in a couple of seconds!!

Nomad 的头像
Nomad5 个月前

Diarization baked in is cool, but 7B parameters kinda defeats the point. You can’t roll this out on device. Parakeet at 0.6b or Granite 1b are a better fit for os local

EventsLooped 的头像
EventsLooped5 个月前

Whisper models are also free to run on your machine

VC Intern 的头像
VC Intern5 个月前

we’re entering the phase where: speech becomes a solved layer the value won’t be transcription it’ll be what you do with it after

Alpha Batcher 的头像
Alpha Batcher5 个月前

i love free open sourced models >3 applause to Microsoft

Julian Goldie SEO 的头像
Julian Goldie SEO5 个月前

that’s actually a big drop

Sanchit 的头像
Sanchit5 个月前

oh wow i had been using elevenlabs API inside of kilo for some work but it was slightly costly. I am going to try this out. 7B is not a lot to run locally

tang | AI Product Maker 的头像
tang | AI Product Maker5 个月前

the 'no chunking' part is the real unlock. tested 3 ASR pipelines for a video workflow last quarter and chunked diarization always lost ~6-8% speaker accuracy when voices swapped at the chunk boundary. how does VibeVoice deal with overlapping speakers in the same window?

Quadcode AI 的头像
Quadcode AI5 个月前

funny how every week there's a revolutonary asr model that supposedly fixes everything at once if it actually nails long form audio and diarization cleanly then thats genuinely impressive but claims like that usually age weirdly

juaniam 的头像
juaniam5 个月前

Or you can use typrapp with whisper large v3 turbo running locally for free and a nice ui on top

flamz 的头像
flamz5 个月前

Nah, they are data farming peoples voices and you all are willingly giving it to them.

Layton Gott 的头像
Layton Gott5 个月前

It's not completely free...

Sebastian Buzdugan 的头像
Sebastian Buzdugan5 个月前

single pass is cool but diarization quality will decide if this actually ships

Max Bourke 的头像
Max Bourke5 个月前

Hmm single pass sounds good but 7b is pretty big for local transcription… be curious to see if quantized versions perform well.

Victor Tavernari  的头像
Victor Tavernari 5 个月前

@perssua @lucas_montano

Eric 的头像
Eric4 个月前

ignoring the caps lock hype, the actual signal here isn't the 60 minute context window. it's the native diarization.

CODIFY 的头像
CODIFY5 个月前

That’s ground-breaking!

Ben Dickey 的头像
Ben Dickey5 个月前

Wow,

Cyril Gupta 的头像
Cyril Gupta5 个月前

Microsoft just collapsed half the ASR stack. One-pass transcription + built-in diarization = no chunking, no glue code. Paired with Ollama → fully local, private voice pipelines. Free isn’t the headline. Owning the stack is.

Goldengmgm🐸 Яebel 🪖 的头像
Goldengmgm🐸 Яebel 🪖5 个月前

wen @AskVenice

Maxwell de Melo 的头像
Maxwell de Melo5 个月前

Nice!

Facundo Campo · Dev 的头像
Facundo Campo · Dev5 个月前

Es un cambio de rumbo!

DiffusionTales 的头像
DiffusionTales5 个月前

I use SubtitleEdit and use the audio to text (whisper).

Adam Jaworski 的头像
Adam Jaworski5 个月前

is it any good?

AIGoldenfinger 的头像
AIGoldenfinger5 个月前

17GB ASR Speech-to-Text model? Are you kidding me? I will run a Film Production Pipelines on 12GB GPU, that would generate the Images, The videos and also the Conversations And also the Music (the score). Microsoft Speech-to-text... ROFL - Get Out Of Here. Go fix your Windows MS

⭕💡 的头像
⭕💡5 个月前

When you say 🆓 what does it mean?

Terry Nunn 的头像
Terry Nunn5 个月前

Don’t just open source the weights, host it for us for free! No one normal person can afford the GPUs needed for this.

Tabish Gillani 的头像
Tabish Gillani5 个月前

Is it available on

Rey Neill 的头像
Rey Neill5 个月前

Hm

AI's Nest 的头像
AI's Nest5 个月前

We are entering an era where the most successful researchers won't be the best coders, but the best "Product Managers" of their own ideas

相关视频

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,715 次观看 • 5 个月前

NVIDIA JUST DROPPED A FREE AI MODEL THAT READS PDFS, WATCHES VIDEOS, LISTENS TO AUDIO, AND UNDERSTANDS YOUR SCREEN SIMULTANEOUSLY. Not one at a time. ALL AT ONCE. In a single pass. It is called Nemotron 3 Nano Omni and it runs 9 times faster than every other multimodal model currently available. Think about what that actually means for how you work. Right now you are switching between tools constantly. One tool for transcribing your call recordings. A different tool for analyzing your client PDFs. Another tool for processing your training videos. A separate workflow for understanding what is happening on your screen. Four tools. Four contexts. Four different outputs you have to manually synthesize into one decision. Nemotron 3 Nano Omni does all of it in one model. One pass. One output. The use cases that just got dramatically simpler: Meeting recordings where you need the transcript, the visual context, and the document references all analyzed together. Training videos where the audio, the slides, and the on-screen demonstrations all feed into one coherent summary. Client PDFs where you need the document content cross-referenced against your screen data and your call notes simultaneously. Sales call transcripts analyzed alongside the proposals and the CRM data in one unified pass. This is not a marginal improvement on existing multimodal models. It is a 9x speed increase on a capability that was already changing how people work. Free. From NVIDIA. Available right now. Bookmark this before everyone catches on. Follow CyrilXBT for every AI capability shift the moment it drops.

CyrilXBT

37,847 次观看 • 5 个月前