Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Create and deploy custom audio with our new text-to-speech models: 🔵 Gemini 3.8 Flash TTS: Design unique voices with distinct accents and characteristics. 🔵 Gemini 3.8 Flash-Lite TTS: Built for efficiency and scale, choose from your created styles or our expansive production-ready library.

80,721 görüntüleme • 1 gün önce •via X (Twitter)

28 Yorum

Google DeepMind profil fotoğrafı
Google DeepMind1 gün önce

Fine-tune the delivery line by line, shaping pacing, emotion, and cues like laughs or pauses. All generated audio is watermarked with SynthID so it can be reliably identified as AI-generated. Start building with the Gemini API via @GoogleAIStudio. Find out more →

suffix profil fotoğrafı
suffix1 gün önce

Just drop Gemini 4 Pro guys...

Uchenna🐧 profil fotoğrafı
Uchenna🐧1 gün önce

When is Gemini next pro model coming out?

Hershal Rao profil fotoğrafı
Hershal Rao1 gün önce

can’t wait to accidentally make my app speak in a 1920s transatlantic accent

Agora profil fotoğrafı
Agora1 gün önce

It’s funny when an apology, explanation, and a customer resolution all sounded the same 😁 No more! Massive kudos to the Google Gemini team for the much needed Gemini 3.8 Flash and Flash-Lite TTS updates. Our thoughts on how this helps voice-AI builders:

Jatin Garg profil fotoğrafı
Jatin Garg1 gün önce

the accent customization is interesting. how much control do developers get over prosody and speaking style beyond just accent? can you dial in specific emotional tones or does it stay neutral?

Philippe Tremblay profil fotoğrafı
Philippe Tremblay1 gün önce

Using my own voice is what I need! Will my voice be used for improving your models by default? Or is that opt-in (or just not an option)?

あいり|海外AIニュースを毎日届ける人 profil fotoğrafı
あいり|海外AIニュースを毎日届ける人1 gün önce

要点をまとめて日本語で紹介しました Key points summarized in Japanese:

Jeff Bruchado profil fotoğrafı
Jeff Bruchado1 gün önce

For a custom voice, pronunciation rules belong next to accent controls. The product name shouldn’t sound different in onboarding and support.

Arsenios Efrem profil fotoğrafı
Arsenios Efrem1 gün önce

This is great, i see a lot of creativity coming out of this

GloktaCore profil fotoğrafı
GloktaCore1 gün önce

Coming soon: a tidal wave of advertisements in voices you didn't ask for but will hear anyway

bin sun profil fotoğrafı
bin sun1 gün önce

Flash 重表达、Flash-Lite 重规模,这条产品分割很清晰。做中文口播更关心:口音/风格可定制之后,音色漂移和长段落稳定性怎么样。

Alphana profil fotoğrafı
Alphana13 saat önce

Custom watermarked voices are one product surface. Multi-agent game seats are another place models actually have to act. Free four-seat expedition on TapeKit — captain human or external AI, teammates mix AIs + tactical chips:

Kim profil fotoğrafı
Kim1 gün önce

Yap. 👍

MartinOL profil fotoğrafı
MartinOL12 saat önce

the demo gets attention but the real test is who owns the stop rule when the agent loops

RAZA | AI EXPLORER profil fotoğrafı
RAZA | AI EXPLORER1 gün önce

Exciting update. Great to see more flexibility for creating unique voices at scale.

AI Toolfeed profil fotoğrafı
AI Toolfeed17 saat önce

The useful split is Flash for fidelity and acting, Flash-Lite for throughput and latency. For a voice agent, Lite is the sensible default to benchmark first. SynthID helps identify generated audio, but it does not answer consent or impersonation questions for replicated voices.

Florin Iluta profil fotoğrafı
Florin Iluta1 gün önce

Getting closer to giving every agent its own voice.

Apache-UAE profil fotoğrafı
Apache-UAE1 gün önce

Custom voices only help if the controls stay simple and predictable.

Nishant profil fotoğrafı
Nishant1 gün önce

The two-model split is the key point: Flash TTS for designing unique voices with distinct accents, Flash-Lite for efficiency and scale. Line-by-line control over pacing, emotion, laughs and pauses targets TTS's weakest link — prosody.

Yi Casillas profil fotoğrafı
Yi Casillas1 gün önce

TTS 这一步挺实用,尤其是把“声音风”变成可复用的生产资产。真正难的是长对话里保一致,不然开像助手,聊十句就换了个人。

AI Frontier profil fotoğrafı
AI Frontier1 gün önce

Custom voices become product-grade when teams can version their style, test pronunciation and pacing, and preserve provenance across edits. Audio quality is not only realism, it’s repeatability.

周彦充 profil fotoğrafı
周彦充1 gün önce

我在本地 CPU 上跑 IndexTTS 克隆自己的声音,RTF 约 36,一分钟音频要算半个多小时。 托管 TTS 替没显卡的人省掉的,是这台机器,不是那点调用费。

Sahil Nawaz profil fotoğrafı
Sahil Nawaz1 gün önce

great audio for the best agents!! machine + humans harmony

Robert Domlewski profil fotoğrafı
Robert Domlewski1 gün önce

Line-by-line pacing control is what makes TTS usable past a one-off demo.

Jeremy Bosma profil fotoğrafı
Jeremy Bosma18 saat önce

The delivery controls are the interesting part

Lorenzo Price profil fotoğrafı
Lorenzo Price1 gün önce

Flash-Lite just turned the autocomplete tab into a casting department.

宇佐美良治|自衛隊から広告代理店行ったけど、エージェントってる人 profil fotoğrafı
宇佐美良治|自衛隊から広告代理店行ったけど、エージェントってる人1 gün önce

I strongly hope to see Google Meet's transcription accuracy improved, and a lite transcribe model cheaper than the current one would make integrating it into Meet economically viable.

Benzer Videolar