Video yükleniyor...
Video Yüklenemedi
Drop 5/14: Introducing Bulbul V3, our latest text-to-speech model. It raises the bar for how human it sounds, while being super robust. In an independent third-party human listening study, Bulbul V3 delivers the highest listener preference, and low error rates across use-cases and languages. See details in our blog,... show more
424,383 görüntüleme • 7 ay önce •via X (Twitter)
38 Yorum

In a blind study conducted by @JoshTalksLive, listeners compared Bulbul V3, ElevenLabs (v3 alpha and v2.5 flash), and Cartesia Sonic-3 with over 20,000 votes. Bulbul V3 tops the scores for 8kHz audio, setting a new benchmark for speech synthesis for voice agents.

The listeners also tagged real failure modes allowing us to evaluate stability. Bulbul V3 comes out on top, with the lowest average error rates.

We also evaluated for the long-tail of language challenges such as speaking numerics, technical content, and named entities. Bulbul V3 consistently has the lowest error rates across languages.

Bulbul V3 is live. This is your month to go all in. Unlimited usage open through February. The mic is yours!

Incredible work, team Sarvam! Proud of you all!

who ever makes these on brand world class animations 🫡

Insaneee one this is... My Grandfather is kinda shocked at how good this turned out to be. He's a die hard fan of Ponniyin Selvan. Who let TTS Cook...🤯🤯

Deep Tech is funny. Everyone thinks you're useless until the one day you drop something great and everyone loses their mind. Keeping morale high until that point is crucial.

Excited for drop 14/14👀

in hebrew it means penis v3

Another one

If Sarvam creates a big LLM chatbot like ChatGPT, Gemini, Perplexity then rest all these Sarvam models will automatically get a huge network effect boost.

אני בהלם שקראתם לזה בולבול

It is an inflection point where the world realizes that India-based deep learning labs can also deliver. This will change investors' views, and they will start taking riskier bets on Indian deep tech startups more confidently.

Bülbül

lesssgooo bulbul

Amazing work guys!! We got TTS..! Sarvam Vision is also really impressive..!

@pratykumar, that listener preference data is a strong indicator of Bulbul V3's enhanced human-like quality.

Good going @pratykumar.

ratan ji ki awaj me

Congratulations to the Sarvam Team for amazing work.

Pratyush, I request you to make a single website for web users and an app for mobile users in which a user can easily navigate all 14 models of yours in a single place.

I would suggest that u partner with indian influencers like podcasters or something so that they can automatically translate their entire podcast into other languages for greater visibility

Damn 🔥🔥

One suggestion: can we add a pure Hindi mode? Don’t want to deal with that Hinglish crap.

amazing work!

Nice name

Huge kudos for building something that’s not just strong on the tech, but also so thoughtfully and tastefully designed.

Can da model handle “Hinglish” ? That’s imp given our desi urban lifestyle …-also ur models can be used for Filmy world in subtitles n post prod edits etc ..that’s a niche market in itself

Why no pure #Marathi language option given in text to speech @SarvamAI ?? Also need more voice options as many sound robotic. Need smooth and sobre voices, not blunt, harsh robotic ones.

YOOO this is so cool So glad to see Indian AI labs getting SOTA

Great job guys i will add it to my movizai tool

you guys are killing it

@manojzxc

Hi Pratyush! Why don't you build a full duplex model ? It is will more natural

Huge leap 👏 If Bulbul V3 truly combines more human-like prosody + low error rates across languages, that’s a big milestone especially for multilingual markets like India.

🔥🔥🔥

Awesome job ! Just tried Telugu, its nice. Still Sounds a slightly robtic though, a bit like old time radio . Still better than V2. 👍 do you have plans to add intention/emotion style to the tone in the future releases so that it sounds more natural?
