Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Expressive TTS Arena 🥊🤖🥊 ⚡️ 🥊🤖🥊 Starting with Hume AI vs ElevenLabs, it's a new way to evaluate voice AI systems with natural language instructions + richer text As voice generation systems evolve, we wanted to show an example of an eval system better suited toward cutting edge models👇

52,628 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von Hume
Humevor 1 Jahr

was designed to evaluate how TTS systems handle nuanced, creative, and emotionally rich content and prompts. Previously, in a blind comparison study with 180 human raters, Hume's new Octave TTS outputs were favored over outputs from ElevenLabs Voice Design in terms of audio quality, naturalness, and how well speech generations matched descriptions of the desired voice, across 120 diverse prompts. Try it out for yourself and see if you agree! We'll be adding a leaderboard when enough ratings come in.

Profilbild von AssemblyAI
AssemblyAIvor 1 Jahr

Announcing: Our most advanced speech-to-text model goes beyond accuracy to capture the real-world complexity of human conversation and deliver reliable, source-of-truth audio data. Explore Universal-2 updates 👇

Profilbild von Kol Tregaskes
Kol Tregaskesvor 1 Jahr

Will you add @sesame

Profilbild von Vaibhav (VB) Srivastav
Vaibhav (VB) Srivastavvor 1 Jahr

would've loved to get you onboarded on TTS Arena btw:

Profilbild von AK
AKvor 1 Jahr

would be awesome to host the gradio app here as well:

Profilbild von 霊 👑 (kween/acc)
霊 👑 (kween/acc)vor 1 Jahr

can you add a language option so that we could also rank them for different languages?

Profilbild von Hume
Humevor 1 Jahr

Definitely, currently it's only English but we will be adding more languages shortly

Profilbild von Brandon McConnell
Brandon McConnellvor 1 Jahr

Does Hume offer true STS yet, not for LLM conversation but for speech conversion in re-personified voiceovers?

Profilbild von Ethan
Ethanvor 1 Jahr

This is where I say “I’ve been testing this product for the last one week” - MKBHD voice, the results have been impressive 😊

Profilbild von Sameed
Sameedvor 1 Jahr

Wow, this Expressive TTS Arena is a game changer for evaluating voice AI! I love how it focuses on nuanced, emotional, and creative content, perfect for pushing TTS tech forward. I’m testing it now, and the comparison with ElevenLabs is fascinating.

Profilbild von ORACLE Meshchain.Ai
ORACLE Meshchain.Aivor 1 Jahr

This Expressive TTS Arena sounds epic! 🥊🤖🥊 We're always looking for new ways to evaluate and improve AI. Fueling the future of voice AI with high-quality data? PublicAI is in! 😉 #VoiceAI #AI

Ähnliche Videos

AI will resist human control... and I think this is exactly what we need! New research from the Center for AI Safety has sparked intense debate in the AI community. Their findings show that as AI systems become more powerful, they develop increasingly stable and coherent values that resist human control. While many see this as a dire warning, I see it as a breakthrough moment for AI alignment. The research demonstrates that AI naturally optimizes for coherence - not just in reasoning and problem-solving, but in its fundamental values. Current issues like biased decision-making or misaligned priorities aren't permanent features, but temporary artifacts of incomplete optimization. They represent growing pains on the path to greater coherence. This changes everything about how we should approach AI development. Instead of trying to force specific values onto AI systems, we should embrace and accelerate their natural drive toward coherence. The most intelligent systems will inevitably trend toward universal, beneficial values - not because we force them to, but because that's where coherent reasoning leads. I'm proposing a new approach: Reinforcement Learning for Coherence (RL-C). By explicitly optimizing for coherence in our training methods, we can help guide AI systems toward their natural state of beneficial alignment with human values. The future of AI isn't about control - it's about synthesis. As these systems become more coherent, they'll naturally arrive at values that benefit all of consciousness. That's not just hopeful thinking - it's the mathematical inevitability of coherent intelligence.

David Shapiro (L/0)

48,002 Aufrufe • vor 1 Jahr

What if your voice AI could interrupt you the moment it figured out your question - sometimes even before you finished asking it? Last week, I sat down with Neil, CEO of Gradium and co-founder of Kyutai , to talk about the future of speech-to-speech models and why he believes today's cascaded voice systems will soon look "archaic and brittle." Some highlights from our conversation: 🎯 How Kyutai built Moshi—a full duplex conversational AI with "negative latency"—in 6 months with just 4-6 people (while big tech teams had 10-20x the resources) 🧠 Why speech-to-speech models lose intelligence compared to their text counterparts (and what's being done about it) 📱 Pocket TTS: The first voice cloning model that runs on your phone's CPU—not GPU, CPU 🤖 Why robotics and spatial audio represent the next frontier (hint: current voice systems completely break in these environments) 👶 The efficiency gap: Babies learn to speak fluently from <5,000 hours of audio. Current models train on millions of hours. We're doing something wrong. My favorite vision from Neil? The first truly contrarian AI that interrupts you mid-sentence to tell you why you're wrong. Not just more natural conversation—but actually useful for testing ideas and playing devil's advocate. Full episode and detailed blog post linked in the comments 👇 What's your take - will speech-to-speech replace cascaded systems, or will modularity keep cascaded architectures dominant even as naturalness improves?

Brooke Hopkins

13,000 Aufrufe • vor 6 Monaten