Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Scribe — the most accurate Speech to Text model. It has the highest accuracy on benchmarks, outperforming previous state-of-the-art models such as Gemini 2.0 and OpenAI Whisper v3. It’s now the leading model for English, Spanish, Italian, and many more. With support for 99 languages, speaker diarization, character-level...

464,743 просмотров • 1 год назад •via X (Twitter)

Комментарии: 11

Фото профиля ElevenLabs
ElevenLabs1 год назад

It achieves the highest accuracy for the most common languages. And it significantly improves the performance of previously underserved languages such as Serbian, Cantonese, and Gujarati.

Фото профиля ElevenLabs
ElevenLabs1 год назад

Learn more about the benchmarking and features in our blog post:

Фото профиля ElevenLabs
ElevenLabs1 год назад

We have a low-latency version of Scribe coming soon, extending Scribe to real-time use cases.

Фото профиля ElevenLabs
ElevenLabs1 год назад

Scribe is available today in both our UI and API. It’s priced at $0.40 per hour of input audio, with an additional 50% discount available for the next 6 weeks. Sign up for an account here:

Фото профиля ElevenLabs
ElevenLabs1 год назад

Hear from @flavioschneide, one of the lead researchers behind the launch.

Фото профиля ElevenLabs
ElevenLabs1 год назад

Join us next week for a virtual event with the team behind the launch:

Фото профиля AssemblyAI
AssemblyAI1 год назад

Our speech-to-text models are the most accurate on the market with top rankings across industry benchmarks. - The highest accuracy rates—up to 95% - Up to 30% fewer hallucinations than other leaders - Low latency—63 minutes converts in 35 seconds Try via API for free today 👇

Фото профиля Lex Fridman
Lex Fridman1 год назад

Awesome!

Фото профиля Kenrik March
Kenrik March1 год назад

Not hating but feedback so hopefully you take this correctly: I want locally run models. Whisper can do this and even run in near realtime on mobile devices. Also .40 per hour is sky high compared to compute unless it’s wildly inefficient. Open Source previous generation models.

Фото профиля Max Rovensky
Max Rovensky1 год назад

Yeah but have you considered thag I can run whisper locally instead of paying you

Фото профиля Alex Balfanz
Alex Balfanz1 год назад

whoa.

Похожие видео

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 просмотров • 6 месяцев назад