Loading video...

Video Failed to Load

Go Home

Just released on Hugging Face: Vui, a 100M open-source NotebookLM! 3 models: > Vui.BASE is the base checkpoint trained on 40k hours of audio conversations > Vui.ABRAHAM is a single speaker model that can reply with context awareness. > Vui.COHOST is checkpoint with two speakers that can talk to...

43,645 views • 1 year ago •via X (Twitter)

6 Comments

steven's profile picture
steven1 year ago

model here: code: here demo here:

HUDI's profile picture
HUDI1 year ago

🚀 Breaking News! HUDI Datamask Now Integrated with BNB Chain's Greenfield! 🚀 Securely back up your data wallet on decentralized storage—say goodbye to centralized servers! 🔐💾 For the first time in Web3 history, HUDI Datamask empowers you to own your data with the unmatched security of @BNBCHAIN's Greenfield. Experience true data sovereignty and privacy like never before! 🌐✨ Join the revolution in data ownership. Take control today! Learn more 👉 #HUDI #BNBChain #Greenfield #OwnYourData #Web3 #DataSovereignty

Harry Coultas Blum's profile picture
Harry Coultas Blum1 year ago

Creator here! You can also try the models at

Yossi Dahan's profile picture
Yossi Dahan1 year ago

The current emphasis seems to be on TTS, but I believe noise cancellation for real-world calls and semantic VAD are far more critical and needed at this stage than another TTS solution.

kfant's profile picture
kfant1 year ago

sounds good as fast as kokoro?

GTechne's profile picture
GTechne1 year ago

Vui.COHOST just redefined uncanny valley. Those non-speech tics are genius. How’s the latency on Abraham’s context replies?

Related Videos

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

220,977 views • 3 months ago