Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Gemini 3.1 Flash TTS 🗣️, our latest text to speech model with scene direction, speaker level specificity, audio tags, more natural + expressive voices, and support for 70 different languages. Available via our new audio playground in AI Studio and in the Gemini API!

805,885 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 35

Фото профиля Logan Kilpatrick
Logan Kilpatrick5 месяцев назад

The progress from 2.5 to 3.1 has been super strong! Excited to see this much improvement for a Flash model. Learn more in our blog:

Фото профиля leo 🐾
leo 🐾5 месяцев назад

cool but gemini 3.1 pro is looking kinda mid these days wen 3.2? wen 3.5? 😮‍💨

Фото профиля Philipp Schmid
Philipp Schmid5 месяцев назад

Lets go!!

Фото профиля Darek Gusto
Darek Gusto5 месяцев назад

Tbh all those voices sound like AI. Is that the goal? I guess we'll never get the actual Gemini 3.1 Flash. I hope there's something good prepared for Google I/O.

Фото профиля baz
baz5 месяцев назад

still need improvement. the new model(3.1) is skipping words and the it doesn’t respect audio tag position. in this example, the whispering started from the beginning of the audio instead of at the point where the audio tag is. the test is carried with a custom UI using the gemini 3.1 tts preview also it is not yet available in the ai studio testing at:

Фото профиля Rami M
Rami M5 месяцев назад

I get this from ai studio: Http response at 400 or 500 level, error: , http status code: 429

Фото профиля M. Aziz Ulak
M. Aziz Ulak5 месяцев назад

Gemini 3.1 Flash TTS is now on Ollang, and it feels like a real upgrade. The instruction following is sharper and the stability is way better than 2.5 Flash TTS. Excited to keep playing with it ⚡

Фото профиля e
e5 месяцев назад

I've been using Gemma on a hardware appliance and would love to see google OSS their TTS. Any plans to opensource this?

Фото профиля bone
bone5 месяцев назад

hard to tell what scene is doing with the ad's music can it do stuff like AUDIENCE APPLAUSE or SPOOKY STORYTELLING

Фото профиля AIdriving
AIdriving5 месяцев назад

🙄

Фото профиля KITE AI
KITE AI5 месяцев назад

@JeffDean TTS with scene direction is a bigger unlock than it sounds. It collapses the audio production pipeline from studio → prompt. The real test will be consistency across long-form generation. Most models drift emotionally after 30 seconds.

Фото профиля Twlvone
Twlvone5 месяцев назад

scene direction in 70 languages. I'm about to turn my entire reading list into audiobooks with distinct character voices and I'm not even going to feel bad about it

Фото профиля Kory Mathewson
Kory Mathewson5 месяцев назад

I've been waiting years for text-to-speech that feels this natural and controllable - and today it's here for everyone in dozens of languages

Фото профиля Ivan Leo
Ivan Leo5 месяцев назад

Let’s gooooo

Фото профиля Abu Bakar Siddik
Abu Bakar Siddik5 месяцев назад

I am a big fan of the latest Live model. Using it in my product. Glad to see this. Looking forward to integrating with our product.

Фото профиля LGZhss
LGZhss5 месяцев назад

We NEED Gemini 3.1 Flash

Фото профиля Serene Trails
Serene Trails5 месяцев назад

bro, we've got all kind of 3.1 flash-xyz except for the main 3.1 flash! congrats on the launched.

Фото профиля Pato González 🦆
Pato González 🦆5 месяцев назад

I've been struggling for months with AI Studio trying to use 2.5 Flash Lite. Hopefuly this works better and gives me better quota 🤞

Фото профиля David Ondrej
David Ondrej5 месяцев назад

is it in NotebookLM?

Фото профиля Jordan
Jordan5 месяцев назад

What’s a 429 error? Is it my fault.

Фото профиля Danny Thompson
Danny Thompson5 месяцев назад

Whoa! I... Ok this is impressive as hell. Voice AI is the one thing I need out about and have been building a lot in this area. This, this is friggin cool.

Фото профиля M. Aziz Ulak
M. Aziz Ulak5 месяцев назад

I'm updating our backend API calls right now to start using 3.1 Flash immediately 👏

Фото профиля Evil Lord Falcon
Evil Lord Falcon5 месяцев назад

PLEASEEEEEEEE MAKE CHATS ‘TIME AWARE’ 🙏🙏🙏 We need Gemini to recognize when a new day has started!!

Фото профиля Shobit
Shobit5 месяцев назад

Every podcast, audiobook, and voice agent I was going to build this weekend just got a lot cheaper

Фото профиля Vivek Patel
Vivek Patel5 месяцев назад

Where is 3.1 flash as llm?

Фото профиля Sidra Miconi
Sidra Miconi5 месяцев назад

SynthID watermarking baked into every output is Google playing the long game on trust. When AI-generated audio becomes indistinguishable from human speech, the companies that embedded provenance from day one will have the regulatory advantage.

Фото профиля Linus ✦ Ekenstam
Linus ✦ Ekenstam5 месяцев назад

whoop whoop TTS

Фото профиля utkarsh apoorva
utkarsh apoorva5 месяцев назад

Scene direction and speaker-level specificity is the missing piece for real-time conversation AI. The hard part is controlling the context mid-stream when the speaker's intent shifts (someone asks a question vs makes a statement vs is being sarcastic). Interested to see if the API allows that level of runtime control

Фото профиля Hans Yadav
Hans Yadav5 месяцев назад

Nice! For the recent Gemini hackathon I ended up using elevenLabs for tts + stt because latency wasn’t very good on the other model. Any improvements on that front?

Фото профиля Farhan
Farhan5 месяцев назад

elevenlabs has basically had a monopoly on the highly expressive tts market for a while now. if this is priced anything like the other gemini flash models, it's going to instantly become the default for most dev projects.

Фото профиля Chris Sean Dabatos
Chris Sean Dabatos5 месяцев назад

Oh this is great will test it out right now

Фото профиля Muratcan Koylan
Muratcan Koylan5 месяцев назад

Congrats team!

Фото профиля Samuel
Samuel5 месяцев назад

😭😭😭 I thought it was Gemini 3.1 Flash

Фото профиля Colt Feltes
Colt Feltes5 месяцев назад

The sea of steps trying to figure out how to use Google's AI products is relentless. Can use on personal account, but not on business. "Here's the new tool!" --> 15 clicks to find where to actually use it. I usually go straight back to CMD + T in claude or terminal out of frustration.

Фото профиля Melvin Vivas
Melvin Vivas5 месяцев назад

wow, awesome

Похожие видео