Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new...

665,977 Aufrufe • vor 1 Monat •via X (Twitter)

54 Kommentare

Profilbild von Rohan Paul
Rohan Paulvor 1 Monat

Just tested some of the voice, its super realistic. And imo, the character-level timestamps for interruption handling might be the most practical feature here.

Profilbild von Soniox
Sonioxvor 1 Monat

Thank you. Timestamps were the most requested feature already in with our previous tts model.

Profilbild von Kartik Patel
Kartik Patelvor 1 Monat

woohoo! bullish on soniox @ashubhgupta time to update our APIs

Profilbild von Soniox
Sonioxvor 1 Monat

@ashubhgupta 🐂

Profilbild von Kamran Wajdani
Kamran Wajdanivor 1 Monat

The expressive audio tags paired with low-latency streaming really stand out—perfect for building responsive voice interfaces. At $0.70 per hour, it's incredibly cost-effective for high-volume deployments.

Profilbild von Soniox
Sonioxvor 1 Monat

Yes. It's built with high volume in mind. When you need to run voice agents at scale, cost is top tier consideration.

Profilbild von Tina Tavčar
Tina Tavčarvor 1 Monat

This feels like a big step forward for TTS. Really cool to see it out in the world.

Profilbild von James
Jamesvor 1 Monat

This is just absurd. Great release!

Profilbild von Soniox
Sonioxvor 1 Monat

Thanks. Our work is not done yet. We keep on pushing.

Profilbild von Easwee
Easweevor 1 Monat

🎙️🎙️🎙️🎙️🎙️🚀

Profilbild von ghosty
ghostyvor 1 Monat

0.70$ per HOUR? that's a steal

Profilbild von Soniox
Sonioxvor 1 Monat

@saascity_io 🥷

Profilbild von Alexey Fateev
Alexey Fateevvor 1 Monat

HF LINK? 🔗 WHERE

Profilbild von Soniox
Sonioxvor 1 Monat

Not OSS unfortunately 😓

Profilbild von Aria Tech
Aria Techvor 1 Monat

60 languages and 70 cents per hour is wild

Profilbild von 中崎工房 | AIで1時間の仕事を5分で終わらせる人
中崎工房 | AIで1時間の仕事を5分で終わらせる人vor 1 Monat

That sounds very realistic! I've been waiting for your team introducing emotions to your very fast, cheap and good quality model! I'll try and see how it works!

Profilbild von Soniox
Sonioxvor 1 Monat

Thank you, let us know which voice performed the best for you.

Profilbild von Vini Lana
Vini Lanavor 1 Monat

Excited to see how will it perform when you support brazilian portuguese 🔥

Profilbild von Soniox
Sonioxvor 1 Monat

Olá! Anotado 👋

Profilbild von Aark
Aarkvor 1 Monat

insane pricing!

Profilbild von Soniox
Sonioxvor 1 Monat

Priced for scale!

Profilbild von Jacob
Jacobvor 1 Monat

Amazing work! Any plans to support IPA phonemes?

Profilbild von Soniox
Sonioxvor 1 Monat

We will look into it - there seems to be real interest there. Thanks.

Profilbild von Mahimai Raja J ‎
Mahimai Raja J ‎vor 1 Monat

Wow! amazing

Profilbild von Soniox
Sonioxvor 1 Monat

Thanks!

Profilbild von Lingwei Wu
Lingwei Wuvor 1 Monat

awesome tts model!

Profilbild von Soniox
Sonioxvor 1 Monat

@drunkpiano_me Thank you!

Profilbild von Lingwei Wu
Lingwei Wuvor 1 Monat

I love it. Compared to other TTS models, this one is the cheapest yet still the best for my app.

Profilbild von Soniox
Sonioxvor 1 Monat

@drunkpiano_me We invested into accuracy of the spoken data - alphanumerics, emails, addresses, ids etc. need to be spoken correctly cross all supported languages. With the addition of emotion tags it also makes it sound more human.

Profilbild von Crypto Révolution 🇫🇷
Crypto Révolution 🇫🇷vor 1 Monat

You really should get integrated with @OpenRouter

Profilbild von Josh Peterson
Josh Petersonvor 1 Monat

Very promising! But this currently rules it out for my stack unfortunately. Looking for maximally flexible nat lang guided nuance and reliability.

Profilbild von Naman Arora
Naman Aroravor 1 Monat

Amazing product hope you give free test

Profilbild von Soniox
Sonioxvor 1 Monat

Please do, and send us feedback - we read all of it and consider in future releases.

Profilbild von Naman Arora
Naman Aroravor 1 Monat

It's free ?

Profilbild von Aaliya
Aaliyavor 1 Monat

The voice quality at this price is really impressive.

Profilbild von Infroy.dev
Infroy.devvor 1 Monat

Add it to

Profilbild von Paweł
Pawełvor 1 Monat

Sounds great!

Profilbild von Soniox
Sonioxvor 1 Monat

Great - we will keep pushing the boundaries!

Profilbild von Pivi
Pivivor 1 Monat

@Mill3sim3 👀

Profilbild von Blaze (Balázs Galambosi)
Blaze (Balázs Galambosi)vor 1 Monat

@ArtificialAnlys this and omnivoice surely missed from tts leaderboards (especially controlled).

Profilbild von Sebastian Buzdugan
Sebastian Buzduganvor 1 Monat

multilingual cloning often loses speaker identity during code switching, that is the production test

Profilbild von AIDB
AIDBvor 1 Monat

Nice one. Checking this out…

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

$0.70 an hour is really cheap

Profilbild von Mehwish kiran
Mehwish kiranvor 1 Monat

60+ languages with this level of quality is seriously impressive 🔥

Profilbild von Lily Anderson
Lily Andersonvor 1 Monat

Soniox TTS v2 looks like a major step forward for realistic AI voices. Better expression, multilingual support, and voice control are pushing text-to-speech closer to human-level communication. 🚀

Profilbild von Hussain Fakhruddin
Hussain Fakhruddinvor 1 Monat

Expressive control through audio tags is a nice level of detail.

Profilbild von luis
luisvor 1 Monat

@Presidentlin

Profilbild von anml
anmlvor 1 Monat

guys, really impressive, would use it for internal radio ads built on demand for products on sale just today - but need a preset/presets for it

Profilbild von Soniox
Sonioxvor 1 Monat

Noted, we will look into expanding the tooling around it.

Profilbild von m
mvor 1 Monat

Thanks of course, but I prefer open weights. 😄

Profilbild von eurema
euremavor 1 Monat

Demo version sounds cool. I'm waiting for a TTS that would add some randomness in the output, some vibe of a bored person speaking, cause every TTS right now sounds like a radio broadcasting professional with perfect mood and pronounciation.

Profilbild von Soniox
Sonioxvor 1 Monat

Audio tags could help you introduce part of that:

Profilbild von Eugenio Rdz
Eugenio Rdzvor 1 Monat

Hello! We love using soniox STT, we tried the TTS in Spanish and even tried cloning our own local voices… as a feedback we are not using it because it still sounds robotic, the short 20 seconds of voice cloning sounds easy but if it could receive bigger audio for better quality!

Profilbild von Soniox
Sonioxvor 1 Monat

Sample quality impacts the instant voice clone a lot. A noisy/flat sample will sound robotic no matter what. The video you see in our post uses all voices generated through our TTS - voice cloning can be powerful when fed the right data.

Ähnliche Videos