Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new...

666,014 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 54

Фото профиля Rohan Paul
Rohan Paul1 месяц назад

Just tested some of the voice, its super realistic. And imo, the character-level timestamps for interruption handling might be the most practical feature here.

Фото профиля Soniox
Soniox1 месяц назад

Thank you. Timestamps were the most requested feature already in with our previous tts model.

Фото профиля Kartik Patel
Kartik Patel1 месяц назад

woohoo! bullish on soniox @ashubhgupta time to update our APIs

Фото профиля Soniox
Soniox1 месяц назад

@ashubhgupta 🐂

Фото профиля Kamran Wajdani
Kamran Wajdani1 месяц назад

The expressive audio tags paired with low-latency streaming really stand out—perfect for building responsive voice interfaces. At $0.70 per hour, it's incredibly cost-effective for high-volume deployments.

Фото профиля Soniox
Soniox1 месяц назад

Yes. It's built with high volume in mind. When you need to run voice agents at scale, cost is top tier consideration.

Фото профиля Tina Tavčar
Tina Tavčar1 месяц назад

This feels like a big step forward for TTS. Really cool to see it out in the world.

Фото профиля James
James1 месяц назад

This is just absurd. Great release!

Фото профиля Soniox
Soniox1 месяц назад

Thanks. Our work is not done yet. We keep on pushing.

Фото профиля Easwee
Easwee1 месяц назад

🎙️🎙️🎙️🎙️🎙️🚀

Фото профиля ghosty
ghosty1 месяц назад

0.70$ per HOUR? that's a steal

Фото профиля Soniox
Soniox1 месяц назад

@saascity_io 🥷

Фото профиля Alexey Fateev
Alexey Fateev1 месяц назад

HF LINK? 🔗 WHERE

Фото профиля Soniox
Soniox1 месяц назад

Not OSS unfortunately 😓

Фото профиля Aria Tech
Aria Tech1 месяц назад

60 languages and 70 cents per hour is wild

Фото профиля 中崎工房 | AIで1時間の仕事を5分で終わらせる人
中崎工房 | AIで1時間の仕事を5分で終わらせる人1 месяц назад

That sounds very realistic! I've been waiting for your team introducing emotions to your very fast, cheap and good quality model! I'll try and see how it works!

Фото профиля Soniox
Soniox1 месяц назад

Thank you, let us know which voice performed the best for you.

Фото профиля Vini Lana
Vini Lana1 месяц назад

Excited to see how will it perform when you support brazilian portuguese 🔥

Фото профиля Soniox
Soniox1 месяц назад

Olá! Anotado 👋

Фото профиля Aark
Aark1 месяц назад

insane pricing!

Фото профиля Soniox
Soniox1 месяц назад

Priced for scale!

Фото профиля Jacob
Jacob1 месяц назад

Amazing work! Any plans to support IPA phonemes?

Фото профиля Soniox
Soniox1 месяц назад

We will look into it - there seems to be real interest there. Thanks.

Фото профиля Mahimai Raja J ‎
Mahimai Raja J ‎1 месяц назад

Wow! amazing

Фото профиля Soniox
Soniox1 месяц назад

Thanks!

Фото профиля Lingwei Wu
Lingwei Wu1 месяц назад

awesome tts model!

Фото профиля Soniox
Soniox1 месяц назад

@drunkpiano_me Thank you!

Фото профиля Lingwei Wu
Lingwei Wu1 месяц назад

I love it. Compared to other TTS models, this one is the cheapest yet still the best for my app.

Фото профиля Soniox
Soniox1 месяц назад

@drunkpiano_me We invested into accuracy of the spoken data - alphanumerics, emails, addresses, ids etc. need to be spoken correctly cross all supported languages. With the addition of emotion tags it also makes it sound more human.

Фото профиля Crypto Révolution 🇫🇷
Crypto Révolution 🇫🇷1 месяц назад

You really should get integrated with @OpenRouter

Фото профиля Josh Peterson
Josh Peterson1 месяц назад

Very promising! But this currently rules it out for my stack unfortunately. Looking for maximally flexible nat lang guided nuance and reliability.

Фото профиля Naman Arora
Naman Arora1 месяц назад

Amazing product hope you give free test

Фото профиля Soniox
Soniox1 месяц назад

Please do, and send us feedback - we read all of it and consider in future releases.

Фото профиля Naman Arora
Naman Arora1 месяц назад

It's free ?

Фото профиля Aaliya
Aaliya1 месяц назад

The voice quality at this price is really impressive.

Фото профиля Infroy.dev
Infroy.dev1 месяц назад

Add it to

Фото профиля Paweł
Paweł1 месяц назад

Sounds great!

Фото профиля Soniox
Soniox1 месяц назад

Great - we will keep pushing the boundaries!

Фото профиля Pivi
Pivi1 месяц назад

@Mill3sim3 👀

Фото профиля Blaze (Balázs Galambosi)
Blaze (Balázs Galambosi)1 месяц назад

@ArtificialAnlys this and omnivoice surely missed from tts leaderboards (especially controlled).

Фото профиля Sebastian Buzdugan
Sebastian Buzdugan1 месяц назад

multilingual cloning often loses speaker identity during code switching, that is the production test

Фото профиля AIDB
AIDB1 месяц назад

Nice one. Checking this out…

Фото профиля AI Mastery Guide
AI Mastery Guide1 месяц назад

$0.70 an hour is really cheap

Фото профиля Mehwish kiran
Mehwish kiran1 месяц назад

60+ languages with this level of quality is seriously impressive 🔥

Фото профиля Lily Anderson
Lily Anderson1 месяц назад

Soniox TTS v2 looks like a major step forward for realistic AI voices. Better expression, multilingual support, and voice control are pushing text-to-speech closer to human-level communication. 🚀

Фото профиля Hussain Fakhruddin
Hussain Fakhruddin1 месяц назад

Expressive control through audio tags is a nice level of detail.

Фото профиля luis
luis1 месяц назад

@Presidentlin

Фото профиля anml
anml1 месяц назад

guys, really impressive, would use it for internal radio ads built on demand for products on sale just today - but need a preset/presets for it

Фото профиля Soniox
Soniox1 месяц назад

Noted, we will look into expanding the tooling around it.

Фото профиля m
m1 месяц назад

Thanks of course, but I prefer open weights. 😄

Фото профиля eurema
eurema1 месяц назад

Demo version sounds cool. I'm waiting for a TTS that would add some randomness in the output, some vibe of a bored person speaking, cause every TTS right now sounds like a radio broadcasting professional with perfect mood and pronounciation.

Фото профиля Soniox
Soniox1 месяц назад

Audio tags could help you introduce part of that:

Фото профиля Eugenio Rdz
Eugenio Rdz1 месяц назад

Hello! We love using soniox STT, we tried the TTS in Spanish and even tried cloning our own local voices… as a feedback we are not using it because it still sounds robotic, the short 20 seconds of voice cloning sounds easy but if it could receive bigger audio for better quality!

Фото профиля Soniox
Soniox1 месяц назад

Sample quality impacts the instant voice clone a lot. A noisy/flat sample will sound robotic no matter what. The video you see in our post uses all voices generated through our TTS - voice cloning can be powerful when fed the right data.

Похожие видео