正在加载视频...
视频加载失败
Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new... show more
54 条评论

Just tested some of the voice, its super realistic. And imo, the character-level timestamps for interruption handling might be the most practical feature here.

Thank you. Timestamps were the most requested feature already in with our previous tts model.

woohoo! bullish on soniox @ashubhgupta time to update our APIs

@ashubhgupta 🐂

The expressive audio tags paired with low-latency streaming really stand out—perfect for building responsive voice interfaces. At $0.70 per hour, it's incredibly cost-effective for high-volume deployments.

Yes. It's built with high volume in mind. When you need to run voice agents at scale, cost is top tier consideration.

This feels like a big step forward for TTS. Really cool to see it out in the world.

This is just absurd. Great release!

Thanks. Our work is not done yet. We keep on pushing.

🎙️🎙️🎙️🎙️🎙️🚀

0.70$ per HOUR? that's a steal

@saascity_io 🥷

HF LINK? 🔗 WHERE

Not OSS unfortunately 😓

60 languages and 70 cents per hour is wild

That sounds very realistic! I've been waiting for your team introducing emotions to your very fast, cheap and good quality model! I'll try and see how it works!

Thank you, let us know which voice performed the best for you.

Excited to see how will it perform when you support brazilian portuguese 🔥

Olá! Anotado 👋

insane pricing!

Priced for scale!

Amazing work! Any plans to support IPA phonemes?

We will look into it - there seems to be real interest there. Thanks.

Wow! amazing

Thanks!

awesome tts model!

@drunkpiano_me Thank you!

I love it. Compared to other TTS models, this one is the cheapest yet still the best for my app.

@drunkpiano_me We invested into accuracy of the spoken data - alphanumerics, emails, addresses, ids etc. need to be spoken correctly cross all supported languages. With the addition of emotion tags it also makes it sound more human.

You really should get integrated with @OpenRouter

Very promising! But this currently rules it out for my stack unfortunately. Looking for maximally flexible nat lang guided nuance and reliability.

Amazing product hope you give free test

Please do, and send us feedback - we read all of it and consider in future releases.

It's free ?

The voice quality at this price is really impressive.

Add it to

Sounds great!

Great - we will keep pushing the boundaries!

@Mill3sim3 👀

@ArtificialAnlys this and omnivoice surely missed from tts leaderboards (especially controlled).

multilingual cloning often loses speaker identity during code switching, that is the production test

Nice one. Checking this out…

$0.70 an hour is really cheap

60+ languages with this level of quality is seriously impressive 🔥

Soniox TTS v2 looks like a major step forward for realistic AI voices. Better expression, multilingual support, and voice control are pushing text-to-speech closer to human-level communication. 🚀

Expressive control through audio tags is a nice level of detail.

@Presidentlin

guys, really impressive, would use it for internal radio ads built on demand for products on sale just today - but need a preset/presets for it

Noted, we will look into expanding the tooling around it.

Thanks of course, but I prefer open weights. 😄

Demo version sounds cool. I'm waiting for a TTS that would add some randomness in the output, some vibe of a bored person speaking, cause every TTS right now sounds like a radio broadcasting professional with perfect mood and pronounciation.

Audio tags could help you introduce part of that:

Hello! We love using soniox STT, we tried the TTS in Spanish and even tried cloning our own local voices… as a feedback we are not using it because it still sounds robotic, the short 20 seconds of voice cloning sounds easy but if it could receive bigger audio for better quality!

Sample quality impacts the instant voice clone a lot. A noisy/flat sample will sound robotic no matter what. The video you see in our post uses all voices generated through our TTS - voice cloning can be powerful when fed the right data.
