Video wird geladen...
Video konnte nicht geladen werden
Today we’re introducing Scribe v2: the most accurate transcription model ever released. While Scribe v2 Realtime is optimized for ultra low latency and agents use cases, Scribe v2 is built for batch transcription, subtitling, and captioning at scale.
558,058 Aufrufe • vor 9 Monaten •via X (Twitter)
40 Kommentare

Scribe v2 achieves the lowest word error rate based on industry-standard benchmarks. Scribe v2 improves on the state of the art stability from Scribe v1. It handles pauses, changes in tone and delivery, together with long silences without any issues, delivering unmatched accuracy across more than 90 languages.

Scribe v2 is now used in ElevenLabs Studio for more accurate subtitles, captions and transcriptions, supporting teams that manage large libraries of audio and video across marketing, media, research, training, and compliance use cases.

Keyterm Prompting. Keyterm prompting goes beyond standard Custom Vocabulary by using the transcript’s context. Select up to 100 words or phrases, and Scribe v2 will accurately decide when to transcribe those terms.

Entity Detection. Select up to 56 categories across Personally Identifiable Information, health data or payment details. Scribe v2 will automatically detect these instances and their exact timestamps in your transcript. Read the docs:

Smart Multi-language Support. Send audio with multiple languages and Scribe v2 will automatically detect and transcribe in the right language.

Other key features: - Smart Speaker Diarization: Intuitive labeling of every speaker for clear, organized transcripts - Precise Word-Level Timestamps: Capture the exact moment each word is spoken. Scribe v2 detailed timestamps enable seamless subtitle syncing and interactive experiences - Dynamic Audio Tagging: From laughter to footsteps, Scribe tags every sound event, enriching your transcripts with the full context of your audio - Enterprise ready with SOC 2, ISO27001, PCI DSS L1, HIPAA, GDPR compliance, EU & India residency and support for zero retention mode

Build with the API. With Scribe v2, developers and enterprises can automate complex audio pipelines, achieve higher accuracy in global content workflows, and scale with full compliance and data residency controls. Read the docs:

Try Scribe v2 today

Scribe v3 feature request, please remove "ums" and "ahs" from transcripts 🙏

cool

This model is a beast! It's the big version of Scribe v2 Realtime Very excited to see what you all build and create with it

This is awesome.. 👏

Scribe v2 feels like the moment speech-to-text finally grew up. Crystal-clear accuracy across dozens of languages, thoughtful keyterm guidance, and speaker diarization that actually works. ElevenLabs just raised the bar again. Brilliant work.

who made the video? great stuff!

Scribe v2 being optimized for batch transcription at scale is perfect for subtitling workflows. The improved accuracy will make it much easier to handle large volumes of content reliably.

Can it really beat @GroqInc ‘s Whisper Large v3 Turbo ? I feel no.

I'm using @elevenlabs & @suno etc.. for building an open-source version of @goClueso. It turns raw screen recordings into polished product demos. check out!!

Why do these companies always compare their model to these stupid old legacy models..... #DoingTooMuch

So awesome! 🔥 ElevenLabs keeps raising the bar.

Why are you guys not letting me generate :( . Does Elevenlabs have a country blocklist?

Scribe v2 looks absolutely game-changing! The benchmarks are insane—lowest WER across 90+ languages, smart keyterm prompting, entity detection, and that enterprise-grade compliance. Can't wait to test it for long-form podcasts and multilingual subs. Huge congrats @elevenlabs team!

This matters more than people think. Accurate batch transcription = better datasets, better agents, better products. Realtime is cool. But scale + accuracy is where real businesses are built.

So awesome, can’t imagine how complex it is to get to this level of accuracy, congrats to the whole team! Can’t wait to play around with it.

@superwhisper

Hello. I live in Kazakhstan and can't pay for your service. It doesn't accept my bank card. This is a problem for all users I know in my country. Could you please tell me when this issue will be resolved?

That's impressive

@grok , what’s de diff between v2 and v2 realtime and give me wild examples of what we can accomplish today with this new version?

@grok is it cheaper than OpenAI’s

Elevanlab new method

How fast is the async file processing via api?

Fantastic

Have you fixed Armenian? 👌

@elevenlabs, the accuracy improvements in Scribe v2 will significantly help with large-scale subtitling projects.

i need to hallucinate break this model next

This would be great news if you were not exploiting voice creators.

i love the animations in this video. it sooo clean

What is the price per hour of transcription?

Exciting news! Improving accuracy is a game changer.

@grok what does it cost to transcribe 80 hours of audio (approx 30 min each)

don’t know about the model yet, but the release video is pretty cool, gotta give them that

