正在加载视频...

视频加载失败

Introducing Gemini 3.1 Flash TTS 🗣️, our latest text to speech model with scene direction, speaker level specificity, audio tags, more natural + expressive voices, and support for 70 different languages. Available via our new audio playground in AI Studio and in the Gemini API!

805,885 次观看 • 5 个月前 •via X (Twitter)

35 条评论

Logan Kilpatrick 的头像
Logan Kilpatrick5 个月前

The progress from 2.5 to 3.1 has been super strong! Excited to see this much improvement for a Flash model. Learn more in our blog:

leo 🐾 的头像
leo 🐾5 个月前

cool but gemini 3.1 pro is looking kinda mid these days wen 3.2? wen 3.5? 😮‍💨

Philipp Schmid 的头像
Philipp Schmid5 个月前

Lets go!!

Darek Gusto 的头像
Darek Gusto5 个月前

Tbh all those voices sound like AI. Is that the goal? I guess we'll never get the actual Gemini 3.1 Flash. I hope there's something good prepared for Google I/O.

baz 的头像
baz5 个月前

still need improvement. the new model(3.1) is skipping words and the it doesn’t respect audio tag position. in this example, the whispering started from the beginning of the audio instead of at the point where the audio tag is. the test is carried with a custom UI using the gemini 3.1 tts preview also it is not yet available in the ai studio testing at:

Rami M 的头像
Rami M5 个月前

I get this from ai studio: Http response at 400 or 500 level, error: , http status code: 429

M. Aziz Ulak 的头像
M. Aziz Ulak5 个月前

Gemini 3.1 Flash TTS is now on Ollang, and it feels like a real upgrade. The instruction following is sharper and the stability is way better than 2.5 Flash TTS. Excited to keep playing with it ⚡

e 的头像
e5 个月前

I've been using Gemma on a hardware appliance and would love to see google OSS their TTS. Any plans to opensource this?

bone 的头像
bone5 个月前

hard to tell what scene is doing with the ad's music can it do stuff like AUDIENCE APPLAUSE or SPOOKY STORYTELLING

AIdriving 的头像
AIdriving5 个月前

🙄

KITE AI 的头像
KITE AI5 个月前

@JeffDean TTS with scene direction is a bigger unlock than it sounds. It collapses the audio production pipeline from studio → prompt. The real test will be consistency across long-form generation. Most models drift emotionally after 30 seconds.

Twlvone 的头像
Twlvone5 个月前

scene direction in 70 languages. I'm about to turn my entire reading list into audiobooks with distinct character voices and I'm not even going to feel bad about it

Kory Mathewson 的头像
Kory Mathewson5 个月前

I've been waiting years for text-to-speech that feels this natural and controllable - and today it's here for everyone in dozens of languages

Ivan Leo 的头像
Ivan Leo5 个月前

Let’s gooooo

Abu Bakar Siddik 的头像
Abu Bakar Siddik5 个月前

I am a big fan of the latest Live model. Using it in my product. Glad to see this. Looking forward to integrating with our product.

LGZhss 的头像
LGZhss5 个月前

We NEED Gemini 3.1 Flash

Serene Trails 的头像
Serene Trails5 个月前

bro, we've got all kind of 3.1 flash-xyz except for the main 3.1 flash! congrats on the launched.

Pato González 🦆 的头像
Pato González 🦆5 个月前

I've been struggling for months with AI Studio trying to use 2.5 Flash Lite. Hopefuly this works better and gives me better quota 🤞

David Ondrej 的头像
David Ondrej5 个月前

is it in NotebookLM?

Jordan 的头像
Jordan5 个月前

What’s a 429 error? Is it my fault.

Danny Thompson 的头像
Danny Thompson5 个月前

Whoa! I... Ok this is impressive as hell. Voice AI is the one thing I need out about and have been building a lot in this area. This, this is friggin cool.

M. Aziz Ulak 的头像
M. Aziz Ulak5 个月前

I'm updating our backend API calls right now to start using 3.1 Flash immediately 👏

Evil Lord Falcon 的头像
Evil Lord Falcon5 个月前

PLEASEEEEEEEE MAKE CHATS ‘TIME AWARE’ 🙏🙏🙏 We need Gemini to recognize when a new day has started!!

Shobit 的头像
Shobit5 个月前

Every podcast, audiobook, and voice agent I was going to build this weekend just got a lot cheaper

Vivek Patel 的头像
Vivek Patel5 个月前

Where is 3.1 flash as llm?

Sidra Miconi 的头像
Sidra Miconi5 个月前

SynthID watermarking baked into every output is Google playing the long game on trust. When AI-generated audio becomes indistinguishable from human speech, the companies that embedded provenance from day one will have the regulatory advantage.

Linus ✦ Ekenstam 的头像
Linus ✦ Ekenstam5 个月前

whoop whoop TTS

utkarsh apoorva 的头像
utkarsh apoorva5 个月前

Scene direction and speaker-level specificity is the missing piece for real-time conversation AI. The hard part is controlling the context mid-stream when the speaker's intent shifts (someone asks a question vs makes a statement vs is being sarcastic). Interested to see if the API allows that level of runtime control

Hans Yadav 的头像
Hans Yadav5 个月前

Nice! For the recent Gemini hackathon I ended up using elevenLabs for tts + stt because latency wasn’t very good on the other model. Any improvements on that front?

Farhan 的头像
Farhan5 个月前

elevenlabs has basically had a monopoly on the highly expressive tts market for a while now. if this is priced anything like the other gemini flash models, it's going to instantly become the default for most dev projects.

Chris Sean Dabatos 的头像
Chris Sean Dabatos5 个月前

Oh this is great will test it out right now

Muratcan Koylan 的头像
Muratcan Koylan5 个月前

Congrats team!

Samuel 的头像
Samuel5 个月前

😭😭😭 I thought it was Gemini 3.1 Flash

Colt Feltes 的头像
Colt Feltes5 个月前

The sea of steps trying to figure out how to use Google's AI products is relentless. Can use on personal account, but not on business. "Here's the new tool!" --> 15 clicks to find where to actually use it. I usually go straight back to CMD + T in claude or terminal out of frustration.

Melvin Vivas 的头像
Melvin Vivas5 个月前

wow, awesome

相关视频