正在加载视频...

视频加载失败

Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new...

665,977 次观看 • 1 个月前 •via X (Twitter)

54 条评论

Rohan Paul 的头像
Rohan Paul1 个月前

Just tested some of the voice, its super realistic. And imo, the character-level timestamps for interruption handling might be the most practical feature here.

Soniox 的头像
Soniox1 个月前

Thank you. Timestamps were the most requested feature already in with our previous tts model.

Kartik Patel 的头像
Kartik Patel1 个月前

woohoo! bullish on soniox @ashubhgupta time to update our APIs

Soniox 的头像
Soniox1 个月前

@ashubhgupta 🐂

Kamran Wajdani 的头像
Kamran Wajdani1 个月前

The expressive audio tags paired with low-latency streaming really stand out—perfect for building responsive voice interfaces. At $0.70 per hour, it's incredibly cost-effective for high-volume deployments.

Soniox 的头像
Soniox1 个月前

Yes. It's built with high volume in mind. When you need to run voice agents at scale, cost is top tier consideration.

Tina Tavčar 的头像
Tina Tavčar1 个月前

This feels like a big step forward for TTS. Really cool to see it out in the world.

James 的头像
James1 个月前

This is just absurd. Great release!

Soniox 的头像
Soniox1 个月前

Thanks. Our work is not done yet. We keep on pushing.

Easwee 的头像
Easwee1 个月前

🎙️🎙️🎙️🎙️🎙️🚀

ghosty 的头像
ghosty1 个月前

0.70$ per HOUR? that's a steal

Soniox 的头像
Soniox1 个月前

@saascity_io 🥷

Alexey Fateev 的头像
Alexey Fateev1 个月前

HF LINK? 🔗 WHERE

Soniox 的头像
Soniox1 个月前

Not OSS unfortunately 😓

Aria Tech 的头像
Aria Tech1 个月前

60 languages and 70 cents per hour is wild

中崎工房 | AIで1時間の仕事を5分で終わらせる人 的头像
中崎工房 | AIで1時間の仕事を5分で終わらせる人1 个月前

That sounds very realistic! I've been waiting for your team introducing emotions to your very fast, cheap and good quality model! I'll try and see how it works!

Soniox 的头像
Soniox1 个月前

Thank you, let us know which voice performed the best for you.

Vini Lana 的头像
Vini Lana1 个月前

Excited to see how will it perform when you support brazilian portuguese 🔥

Soniox 的头像
Soniox1 个月前

Olá! Anotado 👋

Aark 的头像
Aark1 个月前

insane pricing!

Soniox 的头像
Soniox1 个月前

Priced for scale!

Jacob 的头像
Jacob1 个月前

Amazing work! Any plans to support IPA phonemes?

Soniox 的头像
Soniox1 个月前

We will look into it - there seems to be real interest there. Thanks.

Mahimai Raja J ‎ 的头像
Mahimai Raja J ‎1 个月前

Wow! amazing

Soniox 的头像
Soniox1 个月前

Thanks!

Lingwei Wu 的头像
Lingwei Wu1 个月前

awesome tts model!

Soniox 的头像
Soniox1 个月前

@drunkpiano_me Thank you!

Lingwei Wu 的头像
Lingwei Wu1 个月前

I love it. Compared to other TTS models, this one is the cheapest yet still the best for my app.

Soniox 的头像
Soniox1 个月前

@drunkpiano_me We invested into accuracy of the spoken data - alphanumerics, emails, addresses, ids etc. need to be spoken correctly cross all supported languages. With the addition of emotion tags it also makes it sound more human.

Crypto Révolution 🇫🇷 的头像
Crypto Révolution 🇫🇷1 个月前

You really should get integrated with @OpenRouter

Josh Peterson 的头像
Josh Peterson1 个月前

Very promising! But this currently rules it out for my stack unfortunately. Looking for maximally flexible nat lang guided nuance and reliability.

Naman Arora 的头像
Naman Arora1 个月前

Amazing product hope you give free test

Soniox 的头像
Soniox1 个月前

Please do, and send us feedback - we read all of it and consider in future releases.

Naman Arora 的头像
Naman Arora1 个月前

It's free ?

Aaliya 的头像
Aaliya1 个月前

The voice quality at this price is really impressive.

Infroy.dev 的头像
Infroy.dev1 个月前

Add it to

Paweł 的头像
Paweł1 个月前

Sounds great!

Soniox 的头像
Soniox1 个月前

Great - we will keep pushing the boundaries!

Pivi 的头像
Pivi1 个月前

@Mill3sim3 👀

Blaze (Balázs Galambosi) 的头像
Blaze (Balázs Galambosi)1 个月前

@ArtificialAnlys this and omnivoice surely missed from tts leaderboards (especially controlled).

Sebastian Buzdugan 的头像
Sebastian Buzdugan1 个月前

multilingual cloning often loses speaker identity during code switching, that is the production test

AIDB 的头像
AIDB1 个月前

Nice one. Checking this out…

AI Mastery Guide 的头像
AI Mastery Guide1 个月前

$0.70 an hour is really cheap

Mehwish kiran 的头像
Mehwish kiran1 个月前

60+ languages with this level of quality is seriously impressive 🔥

Lily Anderson 的头像
Lily Anderson1 个月前

Soniox TTS v2 looks like a major step forward for realistic AI voices. Better expression, multilingual support, and voice control are pushing text-to-speech closer to human-level communication. 🚀

Hussain Fakhruddin 的头像
Hussain Fakhruddin1 个月前

Expressive control through audio tags is a nice level of detail.

luis 的头像
luis1 个月前

@Presidentlin

anml 的头像
anml1 个月前

guys, really impressive, would use it for internal radio ads built on demand for products on sale just today - but need a preset/presets for it

Soniox 的头像
Soniox1 个月前

Noted, we will look into expanding the tooling around it.

m 的头像
m1 个月前

Thanks of course, but I prefer open weights. 😄

eurema 的头像
eurema1 个月前

Demo version sounds cool. I'm waiting for a TTS that would add some randomness in the output, some vibe of a bored person speaking, cause every TTS right now sounds like a radio broadcasting professional with perfect mood and pronounciation.

Soniox 的头像
Soniox1 个月前

Audio tags could help you introduce part of that:

Eugenio Rdz 的头像
Eugenio Rdz1 个月前

Hello! We love using soniox STT, we tried the TTS in Spanish and even tried cloning our own local voices… as a feedback we are not using it because it still sounds robotic, the short 20 seconds of voice cloning sounds easy but if it could receive bigger audio for better quality!

Soniox 的头像
Soniox1 个月前

Sample quality impacts the instant voice clone a lot. A noisy/flat sample will sound robotic no matter what. The video you see in our post uses all voices generated through our TTS - voice cloning can be powerful when fed the right data.

相关视频