正在加载视频...
视频加载失败
Introducing Gemini 3.1 Flash TTS 🗣️, our latest text to speech model with scene direction, speaker level specificity, audio tags, more natural + expressive voices, and support for 70 different languages. Available via our new audio playground in AI Studio and in the Gemini API!
35 条评论

The progress from 2.5 to 3.1 has been super strong! Excited to see this much improvement for a Flash model. Learn more in our blog:

cool but gemini 3.1 pro is looking kinda mid these days wen 3.2? wen 3.5? 😮💨

Lets go!!

Tbh all those voices sound like AI. Is that the goal? I guess we'll never get the actual Gemini 3.1 Flash. I hope there's something good prepared for Google I/O.

still need improvement. the new model(3.1) is skipping words and the it doesn’t respect audio tag position. in this example, the whispering started from the beginning of the audio instead of at the point where the audio tag is. the test is carried with a custom UI using the gemini 3.1 tts preview also it is not yet available in the ai studio testing at:

I get this from ai studio: Http response at 400 or 500 level, error: , http status code: 429

Gemini 3.1 Flash TTS is now on Ollang, and it feels like a real upgrade. The instruction following is sharper and the stability is way better than 2.5 Flash TTS. Excited to keep playing with it ⚡

I've been using Gemma on a hardware appliance and would love to see google OSS their TTS. Any plans to opensource this?

hard to tell what scene is doing with the ad's music can it do stuff like AUDIENCE APPLAUSE or SPOOKY STORYTELLING

🙄

@JeffDean TTS with scene direction is a bigger unlock than it sounds. It collapses the audio production pipeline from studio → prompt. The real test will be consistency across long-form generation. Most models drift emotionally after 30 seconds.

scene direction in 70 languages. I'm about to turn my entire reading list into audiobooks with distinct character voices and I'm not even going to feel bad about it

I've been waiting years for text-to-speech that feels this natural and controllable - and today it's here for everyone in dozens of languages

Let’s gooooo

I am a big fan of the latest Live model. Using it in my product. Glad to see this. Looking forward to integrating with our product.

We NEED Gemini 3.1 Flash

bro, we've got all kind of 3.1 flash-xyz except for the main 3.1 flash! congrats on the launched.

I've been struggling for months with AI Studio trying to use 2.5 Flash Lite. Hopefuly this works better and gives me better quota 🤞

is it in NotebookLM?

What’s a 429 error? Is it my fault.

Whoa! I... Ok this is impressive as hell. Voice AI is the one thing I need out about and have been building a lot in this area. This, this is friggin cool.

I'm updating our backend API calls right now to start using 3.1 Flash immediately 👏

PLEASEEEEEEEE MAKE CHATS ‘TIME AWARE’ 🙏🙏🙏 We need Gemini to recognize when a new day has started!!

Every podcast, audiobook, and voice agent I was going to build this weekend just got a lot cheaper

Where is 3.1 flash as llm?

SynthID watermarking baked into every output is Google playing the long game on trust. When AI-generated audio becomes indistinguishable from human speech, the companies that embedded provenance from day one will have the regulatory advantage.

whoop whoop TTS

Scene direction and speaker-level specificity is the missing piece for real-time conversation AI. The hard part is controlling the context mid-stream when the speaker's intent shifts (someone asks a question vs makes a statement vs is being sarcastic). Interested to see if the API allows that level of runtime control

Nice! For the recent Gemini hackathon I ended up using elevenLabs for tts + stt because latency wasn’t very good on the other model. Any improvements on that front?

elevenlabs has basically had a monopoly on the highly expressive tts market for a while now. if this is priced anything like the other gemini flash models, it's going to instantly become the default for most dev projects.

Oh this is great will test it out right now

Congrats team!

😭😭😭 I thought it was Gemini 3.1 Flash

The sea of steps trying to figure out how to use Google's AI products is relentless. Can use on personal account, but not on business. "Here's the new tool!" --> 15 clicks to find where to actually use it. I usually go straight back to CMD + T in claude or terminal out of frustration.

wow, awesome



