Загрузка видео...

Не удалось загрузить видео

На главную

We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads. Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for...

454,150 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 34

Фото профиля OpenAI Developers
OpenAI Developers1 месяц назад

Builders can improve live audio transcription by providing: • Free-form context about the recording • Keywords for names and domain-specific terms • Expected input languages • Earlier transcribed turns as context On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it. Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1. Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

Фото профиля OpenAI Developers
OpenAI Developers1 месяц назад

GPT-Transcribe also improved on the Context Aware ASR benchmark, with semantic accuracy increasing from 41.6% without free-form context to 45.2% with it. Across 22 languages on Common Voice, GPT-Transcribe achieved a 19.27% transcription error rate, compared with 40.37% for Whisper (whisper-1). Across nine languages on the Real-World Audio Recording benchmark, it achieved an 8.98% transcription error rate, compared with 15.21% for Whisper (whisper-1).

Фото профиля ISA⅄
ISA⅄1 месяц назад

you had a model that understood contex better than any other model you have released. what did you do, pulled it from the app and kept it for yourselves. no other model comes close to 4o in understanding context. #4oForAll

Фото профиля David Stark
David Stark1 месяц назад

When are you actually gonna listen, and give us news about 4o??. #4oForAll

Фото профиля Joe Williams
Joe Williams1 месяц назад

@roanoke_gal Reintroduce 4o! #keep4o

Фото профиля Nick Dobos
Nick Dobos1 месяц назад

Any chance this works with "Sign in with OpenAI" OAuth?

Фото профиля Victor E. Nunez
Victor E. Nunez1 месяц назад

voice is getting the love it deserves. that was smooth @charlierguo 🎶

Фото профиля X Girls
X Girls1 месяц назад

GPT-Live-Transcribe showing off its low-latency live transcription skills 😂

Фото профиля Sriram Kiron
Sriram Kiron1 месяц назад

awesome!!! when do you think gpt-live-1 is coming to the API?

Фото профиля Umar Saeed
Umar Saeed1 месяц назад

Excited to see how context improves accuracy!

Фото профиля Futsy
Futsy1 месяц назад

@jxnlco Whisper v4 when?

Фото профиля Breno Brito
Breno Brito1 месяц назад

so not opensource?

Фото профиля corey.ching
corey.ching1 месяц назад

Can we get a full ukulele session by @charlierguo?

Фото профиля Jaco Meintjes
Jaco Meintjes1 месяц назад

Thanks but I'll wait for the next version where you are either better or cheaper...

Фото профиля nulltron
nulltron1 месяц назад

my wizard fingers reject anything voice oriented

Фото профиля Denis Spirin
Denis Spirin1 месяц назад

That is a good way to demonstrate the transcription product

Фото профиля Mike R
Mike R1 месяц назад

Can we please also get transcribe as an integral part of ChatGPT Work?

Фото профиля Chris Wandstrom
Chris Wandstrom1 месяц назад

Sharing my updated shortcut here for system-wide dictation to the clipboard on iOS, etc. Very nice when triggered with Action Button, Back Tap, widgets, … 👏 @athyuttamre @WenjieZi @OctopusG4 @vinhocent @theteriyu, and others involved in making these transcription models. Using gpt-transcribe and 5.6-terra for post-processing. Requires Apple Shortcuts app & OpenAI API key.

Фото профиля Bent
Bent1 месяц назад

What about new Whisper models?

Фото профиля Emre YILMAZ
Emre YILMAZ1 месяц назад

When will we have access to GPT Live 1 in the API?

Фото профиля Bruno Jurado
Bruno Jurado1 месяц назад

GPT-Transcribe: $0.0045 per minute, or $0.27 per hour GPT-Live-Transcribe: $0.017 per minute, or $1.02 per hour

Фото профиля AGI Companion 🅇
AGI Companion 🅇1 месяц назад

From the moment ChatGPT Live Voice released, OpenAI set the bar to whole new level and became the King of Voice Transcription and Duplex model. It destroyed all other AI in Voice chat. @elonmusk can we get something for Groki

Фото профиля Jure
Jure1 месяц назад

Where can we find a list of languages?

Фото профиля pierluigivalente.eth
pierluigivalente.eth1 месяц назад

Another 67 startup kil*ed by this amazing update

Фото профиля Katherine
Katherine1 месяц назад

@jxnlco I think I’ll just keep using whisper locally

Фото профиля Anis🐬Al
Anis🐬Al1 месяц назад

Capturing the essence of human speech is a profound endeavor. Beyond the technical feat of low latency and accuracy lies the goal of preserving the nuances of human thought. By refining how we translate voice into meaning, these models act as a bridge, ensuring that the richness of our shared stories and ideas remains intact in a digital landscape. A beautiful step toward more meaningful connection. 🕊️✨

Фото профиля Jaya Nayak
Jaya Nayak1 месяц назад

This is a huge update. The split between low-latency live and optimized async is exactly what developers have been needing, especially for handling messy, real-world audio. 😊

Фото профиля Drew Carson
Drew Carson1 месяц назад

wait isn't Scribe v2 by elevenlabs already WER of like 2.2% ? or am I missing something?

Фото профиля Qiwei
Qiwei1 месяц назад

Can we have a next version of open source whisper model?

Фото профиля Bobby Z
Bobby Z1 месяц назад

Did you really play the ukulele and speak Spanish or was that AI? 😂

Фото профиля TypeWhisper
TypeWhisper1 месяц назад

Just shipped both new models in TypeWhisper: gpt-transcribe for completed recordings and gpt-live-transcribe for realtime dictation, with context, language hints, dictionary terms, and configurable delay. Available now in the OpenAI plugin 1.3.0.

Фото профиля Benedict Kerres
Benedict Kerres1 месяц назад

good name

Фото профиля frustrated by ice
frustrated by ice1 месяц назад

Do you have a demo app where I can provide an api key ?

Фото профиля Michael Wall
Michael Wall1 месяц назад

plugging them in right now

Похожие видео

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 просмотров • 7 месяцев назад