Загрузка видео...

Не удалось загрузить видео

На главную

MICROSOFT OPEN SOURCED A 7B PARAMETER MODEL THAT TRANSCRIBES 60 MINUTES OF AUDIO IN A SINGLE PASS and it's completely free VIBEVOICE ASR no chunking, no context loss, full speaker diarization baked in not just speech to text..not a basic wrapper who spoke, when they spoke, exactly what they...

1,375,300 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 32

Фото профиля Seth
Seth5 месяцев назад

MIT means free code, not free compute. Self-hosting a 7B ASR means a GPU instance running 24/7. That's almost certainly more $/mo than a Deepgram bill. There's no Microsoft-hosted free API tier here.

Фото профиля Stef
Stef5 месяцев назад

“Free” Who owns the data now? Who can parse and use that content now? Who can leverage that content now? Yeah… “free” is always a Trojan horse.

Фото профиля timdrchslr
timdrchslr5 месяцев назад

Use groq with whisper large v3 turbo for only 0,04$ per hour, hours of audio transcribed in a couple of seconds!!

Фото профиля Nomad
Nomad5 месяцев назад

Diarization baked in is cool, but 7B parameters kinda defeats the point. You can’t roll this out on device. Parakeet at 0.6b or Granite 1b are a better fit for os local

Фото профиля EventsLooped
EventsLooped5 месяцев назад

Whisper models are also free to run on your machine

Фото профиля VC Intern
VC Intern5 месяцев назад

we’re entering the phase where: speech becomes a solved layer the value won’t be transcription it’ll be what you do with it after

Фото профиля Alpha Batcher
Alpha Batcher5 месяцев назад

i love free open sourced models >3 applause to Microsoft

Фото профиля Julian Goldie SEO
Julian Goldie SEO5 месяцев назад

that’s actually a big drop

Фото профиля Sanchit
Sanchit5 месяцев назад

oh wow i had been using elevenlabs API inside of kilo for some work but it was slightly costly. I am going to try this out. 7B is not a lot to run locally

Фото профиля tang | AI Product Maker
tang | AI Product Maker5 месяцев назад

the 'no chunking' part is the real unlock. tested 3 ASR pipelines for a video workflow last quarter and chunked diarization always lost ~6-8% speaker accuracy when voices swapped at the chunk boundary. how does VibeVoice deal with overlapping speakers in the same window?

Фото профиля Quadcode AI
Quadcode AI5 месяцев назад

funny how every week there's a revolutonary asr model that supposedly fixes everything at once if it actually nails long form audio and diarization cleanly then thats genuinely impressive but claims like that usually age weirdly

Фото профиля juaniam
juaniam5 месяцев назад

Or you can use typrapp with whisper large v3 turbo running locally for free and a nice ui on top

Фото профиля flamz
flamz5 месяцев назад

Nah, they are data farming peoples voices and you all are willingly giving it to them.

Фото профиля Layton Gott
Layton Gott5 месяцев назад

It's not completely free...

Фото профиля Sebastian Buzdugan
Sebastian Buzdugan5 месяцев назад

single pass is cool but diarization quality will decide if this actually ships

Фото профиля Max Bourke
Max Bourke5 месяцев назад

Hmm single pass sounds good but 7b is pretty big for local transcription… be curious to see if quantized versions perform well.

Фото профиля Victor Tavernari 
Victor Tavernari 5 месяцев назад

@perssua @lucas_montano

Фото профиля Eric
Eric4 месяцев назад

ignoring the caps lock hype, the actual signal here isn't the 60 minute context window. it's the native diarization.

Фото профиля CODIFY
CODIFY5 месяцев назад

That’s ground-breaking!

Фото профиля Ben Dickey
Ben Dickey5 месяцев назад

Wow,

Фото профиля Cyril Gupta
Cyril Gupta5 месяцев назад

Microsoft just collapsed half the ASR stack. One-pass transcription + built-in diarization = no chunking, no glue code. Paired with Ollama → fully local, private voice pipelines. Free isn’t the headline. Owning the stack is.

Фото профиля Goldengmgm🐸 Яebel 🪖
Goldengmgm🐸 Яebel 🪖5 месяцев назад

wen @AskVenice

Фото профиля Maxwell de Melo
Maxwell de Melo5 месяцев назад

Nice!

Фото профиля Facundo Campo · Dev
Facundo Campo · Dev5 месяцев назад

Es un cambio de rumbo!

Фото профиля DiffusionTales
DiffusionTales5 месяцев назад

I use SubtitleEdit and use the audio to text (whisper).

Фото профиля Adam Jaworski
Adam Jaworski5 месяцев назад

is it any good?

Фото профиля AIGoldenfinger
AIGoldenfinger5 месяцев назад

17GB ASR Speech-to-Text model? Are you kidding me? I will run a Film Production Pipelines on 12GB GPU, that would generate the Images, The videos and also the Conversations And also the Music (the score). Microsoft Speech-to-text... ROFL - Get Out Of Here. Go fix your Windows MS

Фото профиля ⭕💡
⭕💡5 месяцев назад

When you say 🆓 what does it mean?

Фото профиля Terry Nunn
Terry Nunn5 месяцев назад

Don’t just open source the weights, host it for us for free! No one normal person can afford the GPUs needed for this.

Фото профиля Tabish Gillani
Tabish Gillani5 месяцев назад

Is it available on

Фото профиля Rey Neill
Rey Neill5 месяцев назад

Hm

Фото профиля AI's Nest
AI's Nest5 месяцев назад

We are entering an era where the most successful researchers won't be the best coders, but the best "Product Managers" of their own ideas

Похожие видео

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,715 просмотров • 5 месяцев назад

NVIDIA JUST DROPPED A FREE AI MODEL THAT READS PDFS, WATCHES VIDEOS, LISTENS TO AUDIO, AND UNDERSTANDS YOUR SCREEN SIMULTANEOUSLY. Not one at a time. ALL AT ONCE. In a single pass. It is called Nemotron 3 Nano Omni and it runs 9 times faster than every other multimodal model currently available. Think about what that actually means for how you work. Right now you are switching between tools constantly. One tool for transcribing your call recordings. A different tool for analyzing your client PDFs. Another tool for processing your training videos. A separate workflow for understanding what is happening on your screen. Four tools. Four contexts. Four different outputs you have to manually synthesize into one decision. Nemotron 3 Nano Omni does all of it in one model. One pass. One output. The use cases that just got dramatically simpler: Meeting recordings where you need the transcript, the visual context, and the document references all analyzed together. Training videos where the audio, the slides, and the on-screen demonstrations all feed into one coherent summary. Client PDFs where you need the document content cross-referenced against your screen data and your call notes simultaneously. Sales call transcripts analyzed alongside the proposals and the CRM data in one unified pass. This is not a marginal improvement on existing multimodal models. It is a 9x speed increase on a capability that was already changing how people work. Free. From NVIDIA. Available right now. Bookmark this before everyone catches on. Follow CyrilXBT for every AI capability shift the moment it drops.

CyrilXBT

37,847 просмотров • 5 месяцев назад