Loading video...

Video Failed to Load

Go Home

Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new...

666,014 views • 1 month ago •via X (Twitter)

54 Comments

Rohan Paul's profile picture
Rohan Paul1 month ago

Just tested some of the voice, its super realistic. And imo, the character-level timestamps for interruption handling might be the most practical feature here.

Soniox's profile picture
Soniox1 month ago

Thank you. Timestamps were the most requested feature already in with our previous tts model.

Kartik Patel's profile picture
Kartik Patel1 month ago

woohoo! bullish on soniox @ashubhgupta time to update our APIs

Soniox's profile picture
Soniox1 month ago

@ashubhgupta 🐂

Kamran Wajdani's profile picture
Kamran Wajdani1 month ago

The expressive audio tags paired with low-latency streaming really stand out—perfect for building responsive voice interfaces. At $0.70 per hour, it's incredibly cost-effective for high-volume deployments.

Soniox's profile picture
Soniox1 month ago

Yes. It's built with high volume in mind. When you need to run voice agents at scale, cost is top tier consideration.

Tina Tavčar's profile picture
Tina Tavčar1 month ago

This feels like a big step forward for TTS. Really cool to see it out in the world.

James's profile picture
James1 month ago

This is just absurd. Great release!

Soniox's profile picture
Soniox1 month ago

Thanks. Our work is not done yet. We keep on pushing.

Easwee's profile picture
Easwee1 month ago

🎙️🎙️🎙️🎙️🎙️🚀

ghosty's profile picture
ghosty1 month ago

0.70$ per HOUR? that's a steal

Soniox's profile picture
Soniox1 month ago

@saascity_io 🥷

Alexey Fateev's profile picture
Alexey Fateev1 month ago

HF LINK? 🔗 WHERE

Soniox's profile picture
Soniox1 month ago

Not OSS unfortunately 😓

Aria Tech's profile picture
Aria Tech1 month ago

60 languages and 70 cents per hour is wild

中崎工房 | AIで1時間の仕事を5分で終わらせる人's profile picture
中崎工房 | AIで1時間の仕事を5分で終わらせる人1 month ago

That sounds very realistic! I've been waiting for your team introducing emotions to your very fast, cheap and good quality model! I'll try and see how it works!

Soniox's profile picture
Soniox1 month ago

Thank you, let us know which voice performed the best for you.

Vini Lana's profile picture
Vini Lana1 month ago

Excited to see how will it perform when you support brazilian portuguese 🔥

Soniox's profile picture
Soniox1 month ago

Olá! Anotado 👋

Aark's profile picture
Aark1 month ago

insane pricing!

Soniox's profile picture
Soniox1 month ago

Priced for scale!

Jacob's profile picture
Jacob1 month ago

Amazing work! Any plans to support IPA phonemes?

Soniox's profile picture
Soniox1 month ago

We will look into it - there seems to be real interest there. Thanks.

Mahimai Raja J ‎'s profile picture
Mahimai Raja J ‎1 month ago

Wow! amazing

Soniox's profile picture
Soniox1 month ago

Thanks!

Lingwei Wu's profile picture
Lingwei Wu1 month ago

awesome tts model!

Soniox's profile picture
Soniox1 month ago

@drunkpiano_me Thank you!

Lingwei Wu's profile picture
Lingwei Wu1 month ago

I love it. Compared to other TTS models, this one is the cheapest yet still the best for my app.

Soniox's profile picture
Soniox1 month ago

@drunkpiano_me We invested into accuracy of the spoken data - alphanumerics, emails, addresses, ids etc. need to be spoken correctly cross all supported languages. With the addition of emotion tags it also makes it sound more human.

Crypto Révolution 🇫🇷's profile picture
Crypto Révolution 🇫🇷1 month ago

You really should get integrated with @OpenRouter

Josh Peterson's profile picture
Josh Peterson1 month ago

Very promising! But this currently rules it out for my stack unfortunately. Looking for maximally flexible nat lang guided nuance and reliability.

Naman Arora's profile picture
Naman Arora1 month ago

Amazing product hope you give free test

Soniox's profile picture
Soniox1 month ago

Please do, and send us feedback - we read all of it and consider in future releases.

Naman Arora's profile picture
Naman Arora1 month ago

It's free ?

Aaliya's profile picture
Aaliya1 month ago

The voice quality at this price is really impressive.

Infroy.dev's profile picture
Infroy.dev1 month ago

Add it to

Paweł's profile picture
Paweł1 month ago

Sounds great!

Soniox's profile picture
Soniox1 month ago

Great - we will keep pushing the boundaries!

Pivi's profile picture
Pivi1 month ago

@Mill3sim3 👀

Blaze (Balázs Galambosi)'s profile picture
Blaze (Balázs Galambosi)1 month ago

@ArtificialAnlys this and omnivoice surely missed from tts leaderboards (especially controlled).

Sebastian Buzdugan's profile picture
Sebastian Buzdugan1 month ago

multilingual cloning often loses speaker identity during code switching, that is the production test

AIDB's profile picture
AIDB1 month ago

Nice one. Checking this out…

AI Mastery Guide's profile picture
AI Mastery Guide1 month ago

$0.70 an hour is really cheap

Mehwish kiran's profile picture
Mehwish kiran1 month ago

60+ languages with this level of quality is seriously impressive 🔥

Lily Anderson's profile picture
Lily Anderson1 month ago

Soniox TTS v2 looks like a major step forward for realistic AI voices. Better expression, multilingual support, and voice control are pushing text-to-speech closer to human-level communication. 🚀

Hussain Fakhruddin's profile picture
Hussain Fakhruddin1 month ago

Expressive control through audio tags is a nice level of detail.

luis's profile picture
luis1 month ago

@Presidentlin

anml's profile picture
anml1 month ago

guys, really impressive, would use it for internal radio ads built on demand for products on sale just today - but need a preset/presets for it

Soniox's profile picture
Soniox1 month ago

Noted, we will look into expanding the tooling around it.

m's profile picture
m1 month ago

Thanks of course, but I prefer open weights. 😄

eurema's profile picture
eurema1 month ago

Demo version sounds cool. I'm waiting for a TTS that would add some randomness in the output, some vibe of a bored person speaking, cause every TTS right now sounds like a radio broadcasting professional with perfect mood and pronounciation.

Soniox's profile picture
Soniox1 month ago

Audio tags could help you introduce part of that:

Eugenio Rdz's profile picture
Eugenio Rdz1 month ago

Hello! We love using soniox STT, we tried the TTS in Spanish and even tried cloning our own local voices… as a feedback we are not using it because it still sounds robotic, the short 20 seconds of voice cloning sounds easy but if it could receive bigger audio for better quality!

Soniox's profile picture
Soniox1 month ago

Sample quality impacts the instant voice clone a lot. A noisy/flat sample will sound robotic no matter what. The video you see in our post uses all voices generated through our TTS - voice cloning can be powerful when fed the right data.

Related Videos