Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Gemini 3.1 Flash TTS 🗣️, our latest text to speech model with scene direction, speaker level specificity, audio tags, more natural + expressive voices, and support for 70 different languages. Available via our new audio playground in AI Studio and in the Gemini API!

805,885 Aufrufe • vor 5 Monaten •via X (Twitter)

35 Kommentare

Profilbild von Logan Kilpatrick
Logan Kilpatrickvor 5 Monaten

The progress from 2.5 to 3.1 has been super strong! Excited to see this much improvement for a Flash model. Learn more in our blog:

Profilbild von leo 🐾
leo 🐾vor 5 Monaten

cool but gemini 3.1 pro is looking kinda mid these days wen 3.2? wen 3.5? 😮‍💨

Profilbild von Philipp Schmid
Philipp Schmidvor 5 Monaten

Lets go!!

Profilbild von Darek Gusto
Darek Gustovor 5 Monaten

Tbh all those voices sound like AI. Is that the goal? I guess we'll never get the actual Gemini 3.1 Flash. I hope there's something good prepared for Google I/O.

Profilbild von baz
bazvor 5 Monaten

still need improvement. the new model(3.1) is skipping words and the it doesn’t respect audio tag position. in this example, the whispering started from the beginning of the audio instead of at the point where the audio tag is. the test is carried with a custom UI using the gemini 3.1 tts preview also it is not yet available in the ai studio testing at:

Profilbild von Rami M
Rami Mvor 5 Monaten

I get this from ai studio: Http response at 400 or 500 level, error: , http status code: 429

Profilbild von M. Aziz Ulak
M. Aziz Ulakvor 5 Monaten

Gemini 3.1 Flash TTS is now on Ollang, and it feels like a real upgrade. The instruction following is sharper and the stability is way better than 2.5 Flash TTS. Excited to keep playing with it ⚡

Profilbild von e
evor 5 Monaten

I've been using Gemma on a hardware appliance and would love to see google OSS their TTS. Any plans to opensource this?

Profilbild von bone
bonevor 5 Monaten

hard to tell what scene is doing with the ad's music can it do stuff like AUDIENCE APPLAUSE or SPOOKY STORYTELLING

Profilbild von AIdriving
AIdrivingvor 5 Monaten

🙄

Profilbild von KITE AI
KITE AIvor 5 Monaten

@JeffDean TTS with scene direction is a bigger unlock than it sounds. It collapses the audio production pipeline from studio → prompt. The real test will be consistency across long-form generation. Most models drift emotionally after 30 seconds.

Profilbild von Twlvone
Twlvonevor 5 Monaten

scene direction in 70 languages. I'm about to turn my entire reading list into audiobooks with distinct character voices and I'm not even going to feel bad about it

Profilbild von Kory Mathewson
Kory Mathewsonvor 5 Monaten

I've been waiting years for text-to-speech that feels this natural and controllable - and today it's here for everyone in dozens of languages

Profilbild von Ivan Leo
Ivan Leovor 5 Monaten

Let’s gooooo

Profilbild von Abu Bakar Siddik
Abu Bakar Siddikvor 5 Monaten

I am a big fan of the latest Live model. Using it in my product. Glad to see this. Looking forward to integrating with our product.

Profilbild von LGZhss
LGZhssvor 5 Monaten

We NEED Gemini 3.1 Flash

Profilbild von Serene Trails
Serene Trailsvor 5 Monaten

bro, we've got all kind of 3.1 flash-xyz except for the main 3.1 flash! congrats on the launched.

Profilbild von Pato González 🦆
Pato González 🦆vor 5 Monaten

I've been struggling for months with AI Studio trying to use 2.5 Flash Lite. Hopefuly this works better and gives me better quota 🤞

Profilbild von David Ondrej
David Ondrejvor 5 Monaten

is it in NotebookLM?

Profilbild von Jordan
Jordanvor 5 Monaten

What’s a 429 error? Is it my fault.

Profilbild von Danny Thompson
Danny Thompsonvor 5 Monaten

Whoa! I... Ok this is impressive as hell. Voice AI is the one thing I need out about and have been building a lot in this area. This, this is friggin cool.

Profilbild von M. Aziz Ulak
M. Aziz Ulakvor 5 Monaten

I'm updating our backend API calls right now to start using 3.1 Flash immediately 👏

Profilbild von Evil Lord Falcon
Evil Lord Falconvor 5 Monaten

PLEASEEEEEEEE MAKE CHATS ‘TIME AWARE’ 🙏🙏🙏 We need Gemini to recognize when a new day has started!!

Profilbild von Shobit
Shobitvor 5 Monaten

Every podcast, audiobook, and voice agent I was going to build this weekend just got a lot cheaper

Profilbild von Vivek Patel
Vivek Patelvor 5 Monaten

Where is 3.1 flash as llm?

Profilbild von Sidra Miconi
Sidra Miconivor 5 Monaten

SynthID watermarking baked into every output is Google playing the long game on trust. When AI-generated audio becomes indistinguishable from human speech, the companies that embedded provenance from day one will have the regulatory advantage.

Profilbild von Linus ✦ Ekenstam
Linus ✦ Ekenstamvor 5 Monaten

whoop whoop TTS

Profilbild von utkarsh apoorva
utkarsh apoorvavor 5 Monaten

Scene direction and speaker-level specificity is the missing piece for real-time conversation AI. The hard part is controlling the context mid-stream when the speaker's intent shifts (someone asks a question vs makes a statement vs is being sarcastic). Interested to see if the API allows that level of runtime control

Profilbild von Hans Yadav
Hans Yadavvor 5 Monaten

Nice! For the recent Gemini hackathon I ended up using elevenLabs for tts + stt because latency wasn’t very good on the other model. Any improvements on that front?

Profilbild von Farhan
Farhanvor 5 Monaten

elevenlabs has basically had a monopoly on the highly expressive tts market for a while now. if this is priced anything like the other gemini flash models, it's going to instantly become the default for most dev projects.

Profilbild von Chris Sean Dabatos
Chris Sean Dabatosvor 5 Monaten

Oh this is great will test it out right now

Profilbild von Muratcan Koylan
Muratcan Koylanvor 5 Monaten

Congrats team!

Profilbild von Samuel
Samuelvor 5 Monaten

😭😭😭 I thought it was Gemini 3.1 Flash

Profilbild von Colt Feltes
Colt Feltesvor 5 Monaten

The sea of steps trying to figure out how to use Google's AI products is relentless. Can use on personal account, but not on business. "Here's the new tool!" --> 15 clicks to find where to actually use it. I usually go straight back to CMD + T in claude or terminal out of frustration.

Profilbild von Melvin Vivas
Melvin Vivasvor 5 Monaten

wow, awesome

Ähnliche Videos