Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

MICROSOFT OPEN SOURCED A 7B PARAMETER MODEL THAT TRANSCRIBES 60 MINUTES OF AUDIO IN A SINGLE PASS and it's completely free VIBEVOICE ASR no chunking, no context loss, full speaker diarization baked in not just speech to text..not a basic wrapper who spoke, when they spoke, exactly what they...

1,375,300 görüntüleme • 5 ay önce •via X (Twitter)

32 Yorum

Seth profil fotoğrafı
Seth5 ay önce

MIT means free code, not free compute. Self-hosting a 7B ASR means a GPU instance running 24/7. That's almost certainly more $/mo than a Deepgram bill. There's no Microsoft-hosted free API tier here.

Stef profil fotoğrafı
Stef5 ay önce

“Free” Who owns the data now? Who can parse and use that content now? Who can leverage that content now? Yeah… “free” is always a Trojan horse.

timdrchslr profil fotoğrafı
timdrchslr5 ay önce

Use groq with whisper large v3 turbo for only 0,04$ per hour, hours of audio transcribed in a couple of seconds!!

Nomad profil fotoğrafı
Nomad5 ay önce

Diarization baked in is cool, but 7B parameters kinda defeats the point. You can’t roll this out on device. Parakeet at 0.6b or Granite 1b are a better fit for os local

EventsLooped profil fotoğrafı
EventsLooped5 ay önce

Whisper models are also free to run on your machine

VC Intern profil fotoğrafı
VC Intern5 ay önce

we’re entering the phase where: speech becomes a solved layer the value won’t be transcription it’ll be what you do with it after

Alpha Batcher profil fotoğrafı
Alpha Batcher5 ay önce

i love free open sourced models >3 applause to Microsoft

Julian Goldie SEO profil fotoğrafı
Julian Goldie SEO5 ay önce

that’s actually a big drop

Sanchit profil fotoğrafı
Sanchit5 ay önce

oh wow i had been using elevenlabs API inside of kilo for some work but it was slightly costly. I am going to try this out. 7B is not a lot to run locally

tang | AI Product Maker profil fotoğrafı
tang | AI Product Maker5 ay önce

the 'no chunking' part is the real unlock. tested 3 ASR pipelines for a video workflow last quarter and chunked diarization always lost ~6-8% speaker accuracy when voices swapped at the chunk boundary. how does VibeVoice deal with overlapping speakers in the same window?

Quadcode AI profil fotoğrafı
Quadcode AI5 ay önce

funny how every week there's a revolutonary asr model that supposedly fixes everything at once if it actually nails long form audio and diarization cleanly then thats genuinely impressive but claims like that usually age weirdly

juaniam profil fotoğrafı
juaniam5 ay önce

Or you can use typrapp with whisper large v3 turbo running locally for free and a nice ui on top

flamz profil fotoğrafı
flamz5 ay önce

Nah, they are data farming peoples voices and you all are willingly giving it to them.

Layton Gott profil fotoğrafı
Layton Gott5 ay önce

It's not completely free...

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan5 ay önce

single pass is cool but diarization quality will decide if this actually ships

Max Bourke profil fotoğrafı
Max Bourke5 ay önce

Hmm single pass sounds good but 7b is pretty big for local transcription… be curious to see if quantized versions perform well.

Victor Tavernari  profil fotoğrafı
Victor Tavernari 5 ay önce

@perssua @lucas_montano

Eric profil fotoğrafı
Eric4 ay önce

ignoring the caps lock hype, the actual signal here isn't the 60 minute context window. it's the native diarization.

CODIFY profil fotoğrafı
CODIFY5 ay önce

That’s ground-breaking!

Ben Dickey profil fotoğrafı
Ben Dickey5 ay önce

Wow,

Cyril Gupta profil fotoğrafı
Cyril Gupta5 ay önce

Microsoft just collapsed half the ASR stack. One-pass transcription + built-in diarization = no chunking, no glue code. Paired with Ollama → fully local, private voice pipelines. Free isn’t the headline. Owning the stack is.

Goldengmgm🐸 Яebel 🪖 profil fotoğrafı
Goldengmgm🐸 Яebel 🪖5 ay önce

wen @AskVenice

Maxwell de Melo profil fotoğrafı
Maxwell de Melo5 ay önce

Nice!

Facundo Campo · Dev profil fotoğrafı
Facundo Campo · Dev5 ay önce

Es un cambio de rumbo!

DiffusionTales profil fotoğrafı
DiffusionTales5 ay önce

I use SubtitleEdit and use the audio to text (whisper).

Adam Jaworski profil fotoğrafı
Adam Jaworski5 ay önce

is it any good?

AIGoldenfinger profil fotoğrafı
AIGoldenfinger5 ay önce

17GB ASR Speech-to-Text model? Are you kidding me? I will run a Film Production Pipelines on 12GB GPU, that would generate the Images, The videos and also the Conversations And also the Music (the score). Microsoft Speech-to-text... ROFL - Get Out Of Here. Go fix your Windows MS

⭕💡 profil fotoğrafı
⭕💡5 ay önce

When you say 🆓 what does it mean?

Terry Nunn profil fotoğrafı
Terry Nunn5 ay önce

Don’t just open source the weights, host it for us for free! No one normal person can afford the GPUs needed for this.

Tabish Gillani profil fotoğrafı
Tabish Gillani5 ay önce

Is it available on

Rey Neill profil fotoğrafı
Rey Neill5 ay önce

Hm

AI's Nest profil fotoğrafı
AI's Nest5 ay önce

We are entering an era where the most successful researchers won't be the best coders, but the best "Product Managers" of their own ideas

Benzer Videolar

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,715 görüntüleme • 5 ay önce

NVIDIA JUST DROPPED A FREE AI MODEL THAT READS PDFS, WATCHES VIDEOS, LISTENS TO AUDIO, AND UNDERSTANDS YOUR SCREEN SIMULTANEOUSLY. Not one at a time. ALL AT ONCE. In a single pass. It is called Nemotron 3 Nano Omni and it runs 9 times faster than every other multimodal model currently available. Think about what that actually means for how you work. Right now you are switching between tools constantly. One tool for transcribing your call recordings. A different tool for analyzing your client PDFs. Another tool for processing your training videos. A separate workflow for understanding what is happening on your screen. Four tools. Four contexts. Four different outputs you have to manually synthesize into one decision. Nemotron 3 Nano Omni does all of it in one model. One pass. One output. The use cases that just got dramatically simpler: Meeting recordings where you need the transcript, the visual context, and the document references all analyzed together. Training videos where the audio, the slides, and the on-screen demonstrations all feed into one coherent summary. Client PDFs where you need the document content cross-referenced against your screen data and your call notes simultaneously. Sales call transcripts analyzed alongside the proposals and the CRM data in one unified pass. This is not a marginal improvement on existing multimodal models. It is a 9x speed increase on a capability that was already changing how people work. Free. From NVIDIA. Available right now. Bookmark this before everyone catches on. Follow CyrilXBT for every AI capability shift the moment it drops.

CyrilXBT

37,847 görüntüleme • 5 ay önce