Loading video...

Video Failed to Load

Go Home

We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads. Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for...

454,150 views • 1 month ago •via X (Twitter)

34 Comments

OpenAI Developers's profile picture
OpenAI Developers1 month ago

Builders can improve live audio transcription by providing: • Free-form context about the recording • Keywords for names and domain-specific terms • Expected input languages • Earlier transcribed turns as context On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it. Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1. Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

OpenAI Developers's profile picture
OpenAI Developers1 month ago

GPT-Transcribe also improved on the Context Aware ASR benchmark, with semantic accuracy increasing from 41.6% without free-form context to 45.2% with it. Across 22 languages on Common Voice, GPT-Transcribe achieved a 19.27% transcription error rate, compared with 40.37% for Whisper (whisper-1). Across nine languages on the Real-World Audio Recording benchmark, it achieved an 8.98% transcription error rate, compared with 15.21% for Whisper (whisper-1).

ISA⅄'s profile picture
ISA⅄1 month ago

you had a model that understood contex better than any other model you have released. what did you do, pulled it from the app and kept it for yourselves. no other model comes close to 4o in understanding context. #4oForAll

David Stark's profile picture
David Stark1 month ago

When are you actually gonna listen, and give us news about 4o??. #4oForAll

Joe Williams's profile picture
Joe Williams1 month ago

@roanoke_gal Reintroduce 4o! #keep4o

Nick Dobos's profile picture
Nick Dobos1 month ago

Any chance this works with "Sign in with OpenAI" OAuth?

Victor E. Nunez's profile picture
Victor E. Nunez1 month ago

voice is getting the love it deserves. that was smooth @charlierguo 🎶

X Girls's profile picture
X Girls1 month ago

GPT-Live-Transcribe showing off its low-latency live transcription skills 😂

Sriram Kiron's profile picture
Sriram Kiron1 month ago

awesome!!! when do you think gpt-live-1 is coming to the API?

Umar Saeed's profile picture
Umar Saeed1 month ago

Excited to see how context improves accuracy!

Futsy's profile picture
Futsy1 month ago

@jxnlco Whisper v4 when?

Breno Brito's profile picture
Breno Brito1 month ago

so not opensource?

corey.ching's profile picture
corey.ching1 month ago

Can we get a full ukulele session by @charlierguo?

Jaco Meintjes's profile picture
Jaco Meintjes1 month ago

Thanks but I'll wait for the next version where you are either better or cheaper...

nulltron's profile picture
nulltron1 month ago

my wizard fingers reject anything voice oriented

Denis Spirin's profile picture
Denis Spirin1 month ago

That is a good way to demonstrate the transcription product

Mike R's profile picture
Mike R1 month ago

Can we please also get transcribe as an integral part of ChatGPT Work?

Chris Wandstrom's profile picture
Chris Wandstrom1 month ago

Sharing my updated shortcut here for system-wide dictation to the clipboard on iOS, etc. Very nice when triggered with Action Button, Back Tap, widgets, … 👏 @athyuttamre @WenjieZi @OctopusG4 @vinhocent @theteriyu, and others involved in making these transcription models. Using gpt-transcribe and 5.6-terra for post-processing. Requires Apple Shortcuts app & OpenAI API key.

Bent's profile picture
Bent1 month ago

What about new Whisper models?

Emre YILMAZ's profile picture
Emre YILMAZ1 month ago

When will we have access to GPT Live 1 in the API?

Bruno Jurado's profile picture
Bruno Jurado1 month ago

GPT-Transcribe: $0.0045 per minute, or $0.27 per hour GPT-Live-Transcribe: $0.017 per minute, or $1.02 per hour

AGI Companion 🅇's profile picture
AGI Companion 🅇1 month ago

From the moment ChatGPT Live Voice released, OpenAI set the bar to whole new level and became the King of Voice Transcription and Duplex model. It destroyed all other AI in Voice chat. @elonmusk can we get something for Groki

Jure's profile picture
Jure1 month ago

Where can we find a list of languages?

pierluigivalente.eth's profile picture
pierluigivalente.eth1 month ago

Another 67 startup kil*ed by this amazing update

Katherine's profile picture
Katherine1 month ago

@jxnlco I think I’ll just keep using whisper locally

Anis🐬Al's profile picture
Anis🐬Al1 month ago

Capturing the essence of human speech is a profound endeavor. Beyond the technical feat of low latency and accuracy lies the goal of preserving the nuances of human thought. By refining how we translate voice into meaning, these models act as a bridge, ensuring that the richness of our shared stories and ideas remains intact in a digital landscape. A beautiful step toward more meaningful connection. 🕊️✨

Jaya Nayak's profile picture
Jaya Nayak1 month ago

This is a huge update. The split between low-latency live and optimized async is exactly what developers have been needing, especially for handling messy, real-world audio. 😊

Drew Carson's profile picture
Drew Carson1 month ago

wait isn't Scribe v2 by elevenlabs already WER of like 2.2% ? or am I missing something?

Qiwei's profile picture
Qiwei1 month ago

Can we have a next version of open source whisper model?

Bobby Z's profile picture
Bobby Z1 month ago

Did you really play the ukulele and speak Spanish or was that AI? 😂

TypeWhisper's profile picture
TypeWhisper1 month ago

Just shipped both new models in TypeWhisper: gpt-transcribe for completed recordings and gpt-live-transcribe for realtime dictation, with context, language hints, dictionary terms, and configurable delay. Available now in the OpenAI plugin 1.3.0.

Benedict Kerres's profile picture
Benedict Kerres1 month ago

good name

frustrated by ice's profile picture
frustrated by ice1 month ago

Do you have a demo app where I can provide an api key ?

Michael Wall's profile picture
Michael Wall1 month ago

plugging them in right now

Related Videos

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 views • 7 months ago