Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Saaras V4, our most capable speech to text model yet. It delivers strong performance across both English and Indian languages, with better accuracy over a much wider range of speech. Saaras V4 was built with specific attention to noise, accents, dialects and mixed-language speech. Read more in the blog:

44,151 Aufrufe • vor 3 Tagen •via X (Twitter)

30 Kommentare

Profilbild von Sarvam
Sarvamvor 3 Tagen

Saaras V4 is state of the art across all 22 Indian languages. 10 of these languages have no commercial alternative today.

Profilbild von Sarvam
Sarvamvor 3 Tagen

It also gets the lowest average word error rate across seven English benchmarks covering global accents, meetings, media and financial audio.

Profilbild von Sarvam
Sarvamvor 3 Tagen

Keyterm prompting is a new addition in Saaras V4, giving the model more context before transcription begins. Names, brands, product terms and specialised vocabulary can be provided as context, helping Saaras V4 recognise words that are uncommon, domain-specific or easy to mishear. This is especially useful in areas where a small transcription error can materially change the meaning.

Profilbild von Sarvam
Sarvamvor 3 Tagen

Saaras V4 preserves transcription accuracy across noisy audio. We evaluated Saaras V4 on Kathbath Noisy using LLM-WER, across speech affected by compression, clipping and background interference. Saaras V4 maintains substantially higher transcription accuracy under these conditions, with an error rate less than half that of Deepgram Nova-3 and GPT-4o Transcribe.

Profilbild von Sarvam
Sarvamvor 3 Tagen

Saaras V4 also supports five transcript formats for different applications. The formats let the same model serve very different workflows, from low-latency voice agents to verbatim transcripts for compliance and cleaner English for analysis.

Profilbild von Vishal Singh 🥑
Vishal Singh 🥑vor 3 Tagen

What can Saaras V4 do now that it couldn't before? - Assamese helplines for Assam citizen services, disaster alerts, and health queries - Bodo helplines and panchayat bots for Bodoland / Assam local governance - Manipuri (Meitei) helplines for Manipur public services, hospitals, and education support - Combined Northeast voice desk covering Assamese + Bodo + Manipuri on one line - Santali tribal welfare apps in Jharkhand, Odisha, and Bengal - Santali mining / factory safety voice reporting in Jharkhand and nearby belts - Kashmiri citizen-service IVR and grievance bots for J&K - Dogri citizen-service IVR and local-body helpdesks for Jammu - Maithili agri advisories for Bihar / Mithila farmers - Maithili panchayat and ration / scheme query bots in north Bihar - Konkani banking IVR for Goa and the Konkan coast - Konkani insurance, UPI, and loan collections calls in Goa / coastal Karnataka / Maharashtra - Sanskrit temple recitation transcription and archive search Sanskrit education apps for pathshalas, universities, and exam practice - Sindhi community service, banking, and diaspora support lines - Nepali citizen and tourism helpdesks in Sikkim, Darjeeling, and border districts - Assamese + Nepali thinner-coverage upgrade for state portals that previously failed on dialect and code-mix - Low-resource language voice surveys for elections, health camps, and welfare schemes - Court / police verbatim records in Kashmiri, Dogri, Maithili, Santali, Bodo, or Manipuri - School and literacy apps that accept spoken answers in these languages instead of only Hindi or English Excited to see the ground level impact 🙌

Profilbild von Mayur chaudhary
Mayur chaudharyvor 3 Tagen

And what about 1 trillion parameter model ??

Profilbild von sai santosh kumar
sai santosh kumarvor 3 Tagen

Still no word level time stamps ?

Profilbild von alt
altvor 3 Tagen

Any chance of speech to speech model in near future?

Profilbild von Aware Citizen
Aware Citizenvor 3 Tagen

Woh great going team Sarvam india needs this due to various languages spoken in pan india..this will revolutionize translation in future

Profilbild von Lovey Jain
Lovey Jainvor 3 Tagen

Building on side Prem — a space where people can talk to someone they love/lost by uploading their voice audio and having an AI-powered conversation. I’m exploring whether Sarvam AI can help with the voice cloning + speech-to-speech pipeline Which Sarvam model/process would do?

Profilbild von ✦ SK ✦
✦ SK ✦vor 3 Tagen

sab kuch bna diya text SOTA model kab laoge?? Jo modi ke vishwaguru ka sapna kab pura kroge?? AGI???

Profilbild von Assu_Bhadhu
Assu_Bhadhuvor 3 Tagen

open weight?

Profilbild von Oleks
Oleksvor 3 Tagen

noise and accents are where most speech models fall apart

Profilbild von MochiBro
MochiBrovor 3 Tagen

just wow 😲😳

Profilbild von Mesut De
Mesut Devor 3 Tagen

conrat

Profilbild von Xien Luis
Xien Luisvor 3 Tagen

bulbul v4 next, trust

Profilbild von Tree (🌸, 🌿)
Tree (🌸, 🌿)vor 3 Tagen

gov's propaganda mouthpiece ai model...

Profilbild von LastResort
LastResortvor 3 Tagen

@grok price of this one?

Profilbild von Shrestha
Shresthavor 3 Tagen

What about coding agents? Yet on Waitlist🙁

Profilbild von Rajat
Rajatvor 3 Tagen

new Saaras V4 model is already available on Vercel's @aisdk ❤️

Profilbild von Democracy
Democracyvor 3 Tagen

Any Plan Coding Models ?

Profilbild von Vince Dsouza
Vince Dsouzavor 3 Tagen

the gradients and fonts on this are so clean!!!

Profilbild von On-Device Logs
On-Device Logsvor 3 Tagen

bruh another s2t model? hope it actually outperforms the hype

Profilbild von ankan
ankanvor 3 Tagen

Hey is anybody building with Sarvam please connect, have questions 🫵🏻

Profilbild von Sahil Nawaz
Sahil Nawazvor 3 Tagen

Congratulations on the launch!! Let's build a community together! DMs are open!

Profilbild von Audemy
Audemyvor 3 Tagen

kathbath, the noise benchmark, is read speech: a sample passed only if it exactly matched the prompt, and speakers were told to skip words they found hard to pronounce. that covers accent and channel. disordered speech is a different axis, and it's where asr is access.

Profilbild von Isoldegwow
Isoldegwowvor 3 Tagen

Finally a benchmark for people who start a sentence in English and finish it in Hindi.

Profilbild von Anil Pai
Anil Paivor 3 Tagen

Its v4 and still no word level timestamps. What is the use of segment level timestamps ? 🤷

Profilbild von Evia AI
Evia AIvor 3 Tagen

This kind of robustness to accents and noise is what really makes STT usable in real world

Ähnliche Videos