Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new...

666,014 görüntüleme • 1 ay önce •via X (Twitter)

54 Yorum

Rohan Paul profil fotoğrafı
Rohan Paul1 ay önce

Just tested some of the voice, its super realistic. And imo, the character-level timestamps for interruption handling might be the most practical feature here.

Soniox profil fotoğrafı
Soniox1 ay önce

Thank you. Timestamps were the most requested feature already in with our previous tts model.

Kartik Patel profil fotoğrafı
Kartik Patel1 ay önce

woohoo! bullish on soniox @ashubhgupta time to update our APIs

Soniox profil fotoğrafı
Soniox1 ay önce

@ashubhgupta 🐂

Kamran Wajdani profil fotoğrafı
Kamran Wajdani1 ay önce

The expressive audio tags paired with low-latency streaming really stand out—perfect for building responsive voice interfaces. At $0.70 per hour, it's incredibly cost-effective for high-volume deployments.

Soniox profil fotoğrafı
Soniox1 ay önce

Yes. It's built with high volume in mind. When you need to run voice agents at scale, cost is top tier consideration.

Tina Tavčar profil fotoğrafı
Tina Tavčar1 ay önce

This feels like a big step forward for TTS. Really cool to see it out in the world.

James profil fotoğrafı
James1 ay önce

This is just absurd. Great release!

Soniox profil fotoğrafı
Soniox1 ay önce

Thanks. Our work is not done yet. We keep on pushing.

Easwee profil fotoğrafı
Easwee1 ay önce

🎙️🎙️🎙️🎙️🎙️🚀

ghosty profil fotoğrafı
ghosty1 ay önce

0.70$ per HOUR? that's a steal

Soniox profil fotoğrafı
Soniox1 ay önce

@saascity_io 🥷

Alexey Fateev profil fotoğrafı
Alexey Fateev1 ay önce

HF LINK? 🔗 WHERE

Soniox profil fotoğrafı
Soniox1 ay önce

Not OSS unfortunately 😓

Aria Tech profil fotoğrafı
Aria Tech1 ay önce

60 languages and 70 cents per hour is wild

中崎工房 | AIで1時間の仕事を5分で終わらせる人 profil fotoğrafı
中崎工房 | AIで1時間の仕事を5分で終わらせる人1 ay önce

That sounds very realistic! I've been waiting for your team introducing emotions to your very fast, cheap and good quality model! I'll try and see how it works!

Soniox profil fotoğrafı
Soniox1 ay önce

Thank you, let us know which voice performed the best for you.

Vini Lana profil fotoğrafı
Vini Lana1 ay önce

Excited to see how will it perform when you support brazilian portuguese 🔥

Soniox profil fotoğrafı
Soniox1 ay önce

Olá! Anotado 👋

Aark profil fotoğrafı
Aark1 ay önce

insane pricing!

Soniox profil fotoğrafı
Soniox1 ay önce

Priced for scale!

Jacob profil fotoğrafı
Jacob1 ay önce

Amazing work! Any plans to support IPA phonemes?

Soniox profil fotoğrafı
Soniox1 ay önce

We will look into it - there seems to be real interest there. Thanks.

Mahimai Raja J ‎ profil fotoğrafı
Mahimai Raja J ‎1 ay önce

Wow! amazing

Soniox profil fotoğrafı
Soniox1 ay önce

Thanks!

Lingwei Wu profil fotoğrafı
Lingwei Wu1 ay önce

awesome tts model!

Soniox profil fotoğrafı
Soniox1 ay önce

@drunkpiano_me Thank you!

Lingwei Wu profil fotoğrafı
Lingwei Wu1 ay önce

I love it. Compared to other TTS models, this one is the cheapest yet still the best for my app.

Soniox profil fotoğrafı
Soniox1 ay önce

@drunkpiano_me We invested into accuracy of the spoken data - alphanumerics, emails, addresses, ids etc. need to be spoken correctly cross all supported languages. With the addition of emotion tags it also makes it sound more human.

Crypto Révolution 🇫🇷 profil fotoğrafı
Crypto Révolution 🇫🇷1 ay önce

You really should get integrated with @OpenRouter

Josh Peterson profil fotoğrafı
Josh Peterson1 ay önce

Very promising! But this currently rules it out for my stack unfortunately. Looking for maximally flexible nat lang guided nuance and reliability.

Naman Arora profil fotoğrafı
Naman Arora1 ay önce

Amazing product hope you give free test

Soniox profil fotoğrafı
Soniox1 ay önce

Please do, and send us feedback - we read all of it and consider in future releases.

Naman Arora profil fotoğrafı
Naman Arora1 ay önce

It's free ?

Aaliya profil fotoğrafı
Aaliya1 ay önce

The voice quality at this price is really impressive.

Infroy.dev profil fotoğrafı
Infroy.dev1 ay önce

Add it to

Paweł profil fotoğrafı
Paweł1 ay önce

Sounds great!

Soniox profil fotoğrafı
Soniox1 ay önce

Great - we will keep pushing the boundaries!

Pivi profil fotoğrafı
Pivi1 ay önce

@Mill3sim3 👀

Blaze (Balázs Galambosi) profil fotoğrafı
Blaze (Balázs Galambosi)1 ay önce

@ArtificialAnlys this and omnivoice surely missed from tts leaderboards (especially controlled).

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan1 ay önce

multilingual cloning often loses speaker identity during code switching, that is the production test

AIDB profil fotoğrafı
AIDB1 ay önce

Nice one. Checking this out…

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

$0.70 an hour is really cheap

Mehwish kiran profil fotoğrafı
Mehwish kiran1 ay önce

60+ languages with this level of quality is seriously impressive 🔥

Lily Anderson profil fotoğrafı
Lily Anderson1 ay önce

Soniox TTS v2 looks like a major step forward for realistic AI voices. Better expression, multilingual support, and voice control are pushing text-to-speech closer to human-level communication. 🚀

Hussain Fakhruddin profil fotoğrafı
Hussain Fakhruddin1 ay önce

Expressive control through audio tags is a nice level of detail.

luis profil fotoğrafı
luis1 ay önce

@Presidentlin

anml profil fotoğrafı
anml1 ay önce

guys, really impressive, would use it for internal radio ads built on demand for products on sale just today - but need a preset/presets for it

Soniox profil fotoğrafı
Soniox1 ay önce

Noted, we will look into expanding the tooling around it.

m profil fotoğrafı
m1 ay önce

Thanks of course, but I prefer open weights. 😄

eurema profil fotoğrafı
eurema1 ay önce

Demo version sounds cool. I'm waiting for a TTS that would add some randomness in the output, some vibe of a bored person speaking, cause every TTS right now sounds like a radio broadcasting professional with perfect mood and pronounciation.

Soniox profil fotoğrafı
Soniox1 ay önce

Audio tags could help you introduce part of that:

Eugenio Rdz profil fotoğrafı
Eugenio Rdz1 ay önce

Hello! We love using soniox STT, we tried the TTS in Spanish and even tried cloning our own local voices… as a feedback we are not using it because it still sounds robotic, the short 20 seconds of voice cloning sounds easy but if it could receive bigger audio for better quality!

Soniox profil fotoğrafı
Soniox1 ay önce

Sample quality impacts the instant voice clone a lot. A noisy/flat sample will sound robotic no matter what. The video you see in our post uses all voices generated through our TTS - voice cloning can be powerful when fed the right data.

Benzer Videolar