Загрузка видео...
Не удалось загрузить видео
We recently helped develop 2️⃣ AI tools: NotebookLM and Illuminate to narrate articles and papers, generate stories based on prompts, and even create multi-speaker audio discussions. 💬 A snapshot of how the technology works. 🧵
115,744 просмотров • 1 год назад •via X (Twitter)
Комментарии: 10

These tools build upon our previous research, which includes: ⚪ Creating a model that could produce 30 second segments of dialogue between multiple speakers ⚪ Technology that cast audio generation as a language modeling problem by mapping its generations to sequences of tokens. ↓

Using these advances, our latest speech generation technology can produce 2 minutes of dialogue with improved speaker consistency. 🔊 To generate longer segments, we created a new speech codec which compresses audio into a sequence of tokens in as low as 600 bits per second - without compromising the quality.

A 2 minute piece of dialogue still requires generating over 5000 tokens. 📈 That’s why we also developed a specialized Transformer, which can handle vast amounts of data, match the acoustic token structure, and decode them back into audio using the speech codec. ↓

The applications for advanced speech generation are vast. 🗣️ From improving accessibility and creating new educational experiences to combining these developments with our Gemini models, we’re excited to continue pushing the boundaries of what’s possible. ↓

One of the best things Google has shipped after

The most impressive parts are the speaker overlaps, realistic disfluencies, natural pauses, tone, and timing. I think the fact that the training audio was *unscripted* really helped!

@ValueAnalyst1 you could use this for your AI amatures if you want..

has API?

When new languages? Like Spanish

Can we please have a conversational AI with either 1 of the hosts voices please?
