Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Gemini 3.8 Flash TTS 📣 for expressive, nuanced voice acting, and 3.8 Flash-Lite TTS ⚡️ for high-volume, cost-efficient scale! Our most expressive, multilingual (100+ languages) speech generation models bring powerful new features for developers and creators alike: • Voice Design: Craft custom voices across languages, dialects, and accents...

11,914 görüntüleme • 1 gün önce •via X (Twitter)

17 Yorum

Alisa Fortin profil fotoğrafı
Alisa Fortin1 gün önce

Congratulations! We did it! 🍾

Thor 雷神 ⚡️ profil fotoğrafı
Thor 雷神 ⚡️1 gün önce

Huuge kudos to the team, so many cool new features 💪

Thor 雷神 ⚡️ profil fotoğrafı
Thor 雷神 ⚡️1 gün önce

Read all about it 👇

Jason Stephen profil fotoğrafı
Jason Stephen1 gün önce

Thor voice bot coming soon

Thor 雷神 ⚡️ profil fotoğrafı
Thor 雷神 ⚡️1 gün önce

Finally I can give myself a beautiful German accent 🙌

Emily profil fotoğrafı
Emily20 saat önce

Congratulations 🔥, hopefully it will be implemented in Gemini Notebook soon 🤞.

Hermes ᯅ profil fotoğrafı
Hermes ᯅ23 saat önce

Congrats on the launch Thor! The voice cues make it feel so much more realistic and give depth to the agent output.

Angus Hardy-Francis profil fotoğrafı
Angus Hardy-Francis1 gün önce

very cool - another week, another release. congrats 🙌

Thor 雷神 ⚡️ profil fotoğrafı
Thor 雷神 ⚡️1 gün önce

Oh yeah, Gemini Audio is on a roll! 🤘

Rod profil fotoğrafı
Rod1 gün önce

It's actually **really good** 👏🏻😱

Thor 雷神 ⚡️ profil fotoğrafı
Thor 雷神 ⚡️1 gün önce

Yup, the team cooked! 🔥

Connie Leung 🇭🇰🇨🇦|GDE (Angular, Cloud AI, Web) profil fotoğrafı
Connie Leung 🇭🇰🇨🇦|GDE (Angular, Cloud AI, Web)22 saat önce

Waiting for Firebase support

Thor 雷神 ⚡️ profil fotoğrafı
Thor 雷神 ⚡️22 saat önce

You mean Firebase AI SDK?

Akriti Keswani profil fotoğrafı
Akriti Keswani1 gün önce

This is truly so cool

Michael Waitze profil fotoğrafı
Michael Waitze1 gün önce

The ability to craft custom voices from simple text prompts is a huge leap, and we actually went deeper on this here:

Alec Freudenstein profil fotoğrafı
Alec Freudenstein1 gün önce

#1 on Hume's Voice Design Benchmark (71.4) and Overall Quality Index!!!

Zico profil fotoğrafı
Zico1 gün önce

Congrats on the launch, Thor! Gemini 3.8 Flash TTS sounds fantastic, especially the two-speaker dialogue. I paired it with Agora RTC to build a live AI podcast where listeners can jump in and talk to the hosts:

Benzer Videolar

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,715 görüntüleme • 5 ay önce