Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Drop 5/14: Introducing Bulbul V3, our latest text-to-speech model. It raises the bar for how human it sounds, while being super robust. In an independent third-party human listening study, Bulbul V3 delivers the highest listener preference, and low error rates across use-cases and languages. See details in our blog,...

424,383 görüntüleme • 7 ay önce •via X (Twitter)

38 Yorum

Pratyush Kumar profil fotoğrafı
Pratyush Kumar7 ay önce

In a blind study conducted by @JoshTalksLive, listeners compared Bulbul V3, ElevenLabs (v3 alpha and v2.5 flash), and Cartesia Sonic-3 with over 20,000 votes. Bulbul V3 tops the scores for 8kHz audio, setting a new benchmark for speech synthesis for voice agents.

Pratyush Kumar profil fotoğrafı
Pratyush Kumar7 ay önce

The listeners also tagged real failure modes allowing us to evaluate stability. Bulbul V3 comes out on top, with the lowest average error rates.

Pratyush Kumar profil fotoğrafı
Pratyush Kumar7 ay önce

We also evaluated for the long-tail of language challenges such as speaking numerics, technical content, and named entities. Bulbul V3 consistently has the lowest error rates across languages.

Pratyush Kumar profil fotoğrafı
Pratyush Kumar7 ay önce

Bulbul V3 is live. This is your month to go all in. Unlimited usage open through February. The mic is yours!

India in Pixels by Ashris 🔱 profil fotoğrafı
India in Pixels by Ashris 🔱7 ay önce

Incredible work, team Sarvam! Proud of you all!

akash profil fotoğrafı
akash7 ay önce

who ever makes these on brand world class animations 🫡

ezioAuditoreeeee profil fotoğrafı
ezioAuditoreeeee7 ay önce

Insaneee one this is... My Grandfather is kinda shocked at how good this turned out to be. He's a die hard fan of Ponniyin Selvan. Who let TTS Cook...🤯🤯

Ankit Khandelwal profil fotoğrafı
Ankit Khandelwal7 ay önce

Deep Tech is funny. Everyone thinks you're useless until the one day you drop something great and everyone loses their mind. Keeping morale high until that point is crucial.

Anshuman Mahapatra profil fotoğrafı
Anshuman Mahapatra7 ay önce

Excited for drop 14/14👀

Samuel Cardillo profil fotoğrafı
Samuel Cardillo7 ay önce

in hebrew it means penis v3

Rasputin profil fotoğrafı
Rasputin7 ay önce

Another one

Vansh Thakur profil fotoğrafı
Vansh Thakur7 ay önce

If Sarvam creates a big LLM chatbot like ChatGPT, Gemini, Perplexity then rest all these Sarvam models will automatically get a huge network effect boost.

Ella Ke profil fotoğrafı
Ella Ke7 ay önce

אני בהלם שקראתם לזה בולבול

Ankit Khandelwal profil fotoğrafı
Ankit Khandelwal7 ay önce

It is an inflection point where the world realizes that India-based deep learning labs can also deliver. This will change investors' views, and they will start taking riskier bets on Indian deep tech startups more confidently.

Ersan - beyninikullan.com profil fotoğrafı
Ersan - beyninikullan.com7 ay önce

Bülbül

akash profil fotoğrafı
akash7 ay önce

lesssgooo bulbul

dattasai profil fotoğrafı
dattasai7 ay önce

Amazing work guys!! We got TTS..! Sarvam Vision is also really impressive..!

Himanshu Kumar profil fotoğrafı
Himanshu Kumar7 ay önce

@pratykumar, that listener preference data is a strong indicator of Bulbul V3's enhanced human-like quality.

NS profil fotoğrafı
NS7 ay önce

Good going @pratykumar.

मनीष तिवारी profil fotoğrafı
मनीष तिवारी7 ay önce

ratan ji ki awaj me

UP 10T Economic Goal™ profil fotoğrafı
UP 10T Economic Goal™7 ay önce

Congratulations to the Sarvam Team for amazing work.

Vansh Thakur profil fotoğrafı
Vansh Thakur7 ay önce

Pratyush, I request you to make a single website for web users and an app for mobile users in which a user can easily navigate all 14 models of yours in a single place.

Gaurav Tewatia profil fotoğrafı
Gaurav Tewatia7 ay önce

I would suggest that u partner with indian influencers like podcasters or something so that they can automatically translate their entire podcast into other languages for greater visibility

Arjun Singh profil fotoğrafı
Arjun Singh7 ay önce

Damn 🔥🔥

🇮🇳 𐄳 इन्द्रजाल 𐄳 profil fotoğrafı
🇮🇳 𐄳 इन्द्रजाल 𐄳7 ay önce

One suggestion: can we add a pure Hindi mode? Don’t want to deal with that Hinglish crap.

Sanskar Pandey profil fotoğrafı
Sanskar Pandey7 ay önce

amazing work!

Priyankar Pal profil fotoğrafı
Priyankar Pal7 ay önce

Nice name

Niteesh Yadav ( नितीश ) profil fotoğrafı
Niteesh Yadav ( नितीश )7 ay önce

Huge kudos for building something that’s not just strong on the tech, but also so thoughtfully and tastefully designed.

MS profil fotoğrafı
MS7 ay önce

Can da model handle “Hinglish” ? That’s imp given our desi urban lifestyle …-also ur models can be used for Filmy world in subtitles n post prod edits etc ..that’s a niche market in itself

Yadu A profil fotoğrafı
Yadu A7 ay önce

Why no pure #Marathi language option given in text to speech @SarvamAI ?? Also need more voice options as many sound robotic. Need smooth and sobre voices, not blunt, harsh robotic ones.

KCAerospace🚀🚀 profil fotoğrafı
KCAerospace🚀🚀7 ay önce

YOOO this is so cool So glad to see Indian AI labs getting SOTA

Nio profil fotoğrafı
Nio7 ay önce

Great job guys i will add it to my movizai tool

Ashish Dogra profil fotoğrafı
Ashish Dogra7 ay önce

you guys are killing it

Sanat Kumar Mahapatra profil fotoğrafı
Sanat Kumar Mahapatra7 ay önce

@manojzxc

Krrish Agarwalla profil fotoğrafı
Krrish Agarwalla7 ay önce

Hi Pratyush! Why don't you build a full duplex model ? It is will more natural

TheRaven profil fotoğrafı
TheRaven7 ay önce

Huge leap 👏 If Bulbul V3 truly combines more human-like prosody + low error rates across languages, that’s a big milestone especially for multilingual markets like India.

Lokesh KR profil fotoğrafı
Lokesh KR7 ay önce

🔥🔥🔥

Yash profil fotoğrafı
Yash7 ay önce

Awesome job ! Just tried Telugu, its nice. Still Sounds a slightly robtic though, a bit like old time radio . Still better than V2. 👍 do you have plans to add intention/emotion style to the tone in the future releases so that it sounds more natural?

Benzer Videolar