Loading video...

Video Failed to Load

Go Home

BOOM! Microsoft just released an upgraded VibeVoice Large ~10B Text to Speech model - MIT licensed ๐Ÿ”ฅ > Generate multi-speaker podcasts in minutes โšก > Works blazingly fast on ZeroGPU with H200 (FREE) Try it out today!

89,581 views โ€ข 11 months ago โ€ขvia X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

๐Ÿšจ JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does ๐Ÿ‘‡ VibeVoice ASR - Speech Recognition: โ†’ Processes 60 minutes of continuous audio in a single pass โ†’ Never slices audio into chunks so global context is never lost โ†’ Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously โ†’ Supports customized hotwords for domain specific accuracy โ†’ Works in 50+ languages natively โ†’ Already adopted by Hugging Face Transformers library โ†’ Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: โ†’ Generates up to 90 minutes of speech in a single pass โ†’ Supports up to 4 distinct speakers in one conversation โ†’ Natural turn taking and speaker consistency throughout โ†’ Expressive speech that captures emotional nuances โ†’ Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: โ†’ Only 300 millisecond first audible latency โ†’ Streams text input in real time โ†’ 0.5B parameters so it actually deploys anywhere โ†’ Robust long form generation up to 10 minutes โ†’ Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: โ†’ VibeVoice ASR 7B - available now on Hugging Face โ†’ VibeVoice Realtime 0.5B - try it on Colab right now โ†’ 50+ supported languages โ†’ 11 distinct English voice styles โ†’ 9 multilingual speaker voices โ†’ Already integrated into Hugging Face Transformers โ†’ Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. ๐Ÿ”– Save this before your competitors find it first. ๐Ÿ‘‡

Kanika

221,264 views โ€ข 4 months ago

You can now generate real-time speech that sounds conversational. Microsoft just open-sourced VibeVoice, a real-time text-to-speech system with ~300 ms first audio latency and streaming input. It handles long conversations without falling apart. ๐—ง๐—ต๐—ถ๐˜€ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐˜๐—ฒ๐˜€ ๐—น๐—ผ๐—ป๐—ด, ๐—บ๐˜‚๐—น๐˜๐—ถ-๐˜€๐—ฝ๐—ฒ๐—ฎ๐—ธ๐—ฒ๐—ฟ ๐˜€๐—ฝ๐—ฒ๐—ฒ๐—ฐ๐—ต. It produces up to 90 minutes of audio. It supports up to four distinct speakers. Turn-taking stays consistent over long sessions. ๐—œ๐˜ ๐˜„๐—ผ๐—ฟ๐—ธ๐˜€ ๐—ฏ๐˜† ๐—ฟ๐—ฒ๐—ฑ๐˜‚๐—ฐ๐—ถ๐—ป๐—ด ๐˜๐—ถ๐—บ๐—ฒ ๐—ฟ๐—ฒ๐˜€๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป. Audio compresses into semantic and acoustic tokens. They run at 7.5 Hz instead of frame-level audio. A language model predicts structure. A diffusion head restores acoustic detail. ๐—œ๐˜ ๐—ฎ๐—น๐—น๐—ผ๐˜„๐˜€ ๐—น๐—ผ๐˜„-๐—น๐—ฎ๐˜๐—ฒ๐—ป๐—ฐ๐˜† ๐˜€๐˜๐—ฟ๐—ฒ๐—ฎ๐—บ๐—ถ๐—ป๐—ด ๐—ฎ๐˜‚๐—ฑ๐—ถ๐—ผ. The real-time variant streams text incrementally. First speech arrives in ~300 ms. A WebSocket demo shows live generation. The code is MIT-licensed and research-only. The repo already passed 20k GitHub stars.

Lior Alexander

61,122 views โ€ข 7 months ago

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.๐ŸŽจ โœจ New in 2.1: ๐Ÿ”นAdvanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. ๐Ÿ”นPrecise Chinese and English Text Rendering with seamless imageโ€“text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. ๐Ÿ”นRich Styles and High Aesthetic: Capable of generating images in various stylesโ€”including photorealistic portraits, comics, and vinyl figuresโ€”it delivers outstanding visual appeal and artistic quality. ๐Ÿ”นHigh-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideasโ€”like posters with slogans or multi-panel comicsโ€”into visuals faster than ever. Weโ€™re just getting started. Stay tuned for our native multimodal image generation model coming soon. ๐ŸŒWebsite: ๐Ÿ”—Github: ๐Ÿค—Hugging Face: โœจHugging Face Demo:

Tencent Hy

89,257 views โ€ข 11 months ago

Weโ€™re excited to announce the release and open-source of HunyuanImage 3.0 โ€” the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industryโ€™s flagship closed-source model.๐Ÿš€๐Ÿš€๐Ÿš€ HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: โœ…Reason with world knowledge โœ…Understand complex, thousand-word prompts โœ…Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0โ€™s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. ๐Ÿ‘‰๐ŸปTry it now: ๐Ÿ”—GitHub: ๐Ÿค—Hugging Face:

Tencent Hy

412,880 views โ€ข 11 months ago