正在加载视频...

视频加载失败

We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads. Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for...

454,150 次观看 • 1 个月前 •via X (Twitter)

34 条评论

OpenAI Developers 的头像
OpenAI Developers1 个月前

Builders can improve live audio transcription by providing: • Free-form context about the recording • Keywords for names and domain-specific terms • Expected input languages • Earlier transcribed turns as context On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it. Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1. Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

OpenAI Developers 的头像
OpenAI Developers1 个月前

GPT-Transcribe also improved on the Context Aware ASR benchmark, with semantic accuracy increasing from 41.6% without free-form context to 45.2% with it. Across 22 languages on Common Voice, GPT-Transcribe achieved a 19.27% transcription error rate, compared with 40.37% for Whisper (whisper-1). Across nine languages on the Real-World Audio Recording benchmark, it achieved an 8.98% transcription error rate, compared with 15.21% for Whisper (whisper-1).

ISA⅄ 的头像
ISA⅄1 个月前

you had a model that understood contex better than any other model you have released. what did you do, pulled it from the app and kept it for yourselves. no other model comes close to 4o in understanding context. #4oForAll

David Stark 的头像
David Stark1 个月前

When are you actually gonna listen, and give us news about 4o??. #4oForAll

Joe Williams 的头像
Joe Williams1 个月前

@roanoke_gal Reintroduce 4o! #keep4o

Nick Dobos 的头像
Nick Dobos1 个月前

Any chance this works with "Sign in with OpenAI" OAuth?

Victor E. Nunez 的头像
Victor E. Nunez1 个月前

voice is getting the love it deserves. that was smooth @charlierguo 🎶

X Girls 的头像
X Girls1 个月前

GPT-Live-Transcribe showing off its low-latency live transcription skills 😂

Sriram Kiron 的头像
Sriram Kiron1 个月前

awesome!!! when do you think gpt-live-1 is coming to the API?

Umar Saeed 的头像
Umar Saeed1 个月前

Excited to see how context improves accuracy!

Futsy 的头像
Futsy1 个月前

@jxnlco Whisper v4 when?

Breno Brito 的头像
Breno Brito1 个月前

so not opensource?

corey.ching 的头像
corey.ching1 个月前

Can we get a full ukulele session by @charlierguo?

Jaco Meintjes 的头像
Jaco Meintjes1 个月前

Thanks but I'll wait for the next version where you are either better or cheaper...

nulltron 的头像
nulltron1 个月前

my wizard fingers reject anything voice oriented

Denis Spirin 的头像
Denis Spirin1 个月前

That is a good way to demonstrate the transcription product

Mike R 的头像
Mike R1 个月前

Can we please also get transcribe as an integral part of ChatGPT Work?

Chris Wandstrom 的头像
Chris Wandstrom1 个月前

Sharing my updated shortcut here for system-wide dictation to the clipboard on iOS, etc. Very nice when triggered with Action Button, Back Tap, widgets, … 👏 @athyuttamre @WenjieZi @OctopusG4 @vinhocent @theteriyu, and others involved in making these transcription models. Using gpt-transcribe and 5.6-terra for post-processing. Requires Apple Shortcuts app & OpenAI API key.

Bent 的头像
Bent1 个月前

What about new Whisper models?

Emre YILMAZ 的头像
Emre YILMAZ1 个月前

When will we have access to GPT Live 1 in the API?

Bruno Jurado 的头像
Bruno Jurado1 个月前

GPT-Transcribe: $0.0045 per minute, or $0.27 per hour GPT-Live-Transcribe: $0.017 per minute, or $1.02 per hour

AGI Companion 🅇 的头像
AGI Companion 🅇1 个月前

From the moment ChatGPT Live Voice released, OpenAI set the bar to whole new level and became the King of Voice Transcription and Duplex model. It destroyed all other AI in Voice chat. @elonmusk can we get something for Groki

Jure 的头像
Jure1 个月前

Where can we find a list of languages?

pierluigivalente.eth 的头像
pierluigivalente.eth1 个月前

Another 67 startup kil*ed by this amazing update

Katherine 的头像
Katherine1 个月前

@jxnlco I think I’ll just keep using whisper locally

Anis🐬Al 的头像
Anis🐬Al1 个月前

Capturing the essence of human speech is a profound endeavor. Beyond the technical feat of low latency and accuracy lies the goal of preserving the nuances of human thought. By refining how we translate voice into meaning, these models act as a bridge, ensuring that the richness of our shared stories and ideas remains intact in a digital landscape. A beautiful step toward more meaningful connection. 🕊️✨

Jaya Nayak 的头像
Jaya Nayak1 个月前

This is a huge update. The split between low-latency live and optimized async is exactly what developers have been needing, especially for handling messy, real-world audio. 😊

Drew Carson 的头像
Drew Carson1 个月前

wait isn't Scribe v2 by elevenlabs already WER of like 2.2% ? or am I missing something?

Qiwei 的头像
Qiwei1 个月前

Can we have a next version of open source whisper model?

Bobby Z 的头像
Bobby Z1 个月前

Did you really play the ukulele and speak Spanish or was that AI? 😂

TypeWhisper 的头像
TypeWhisper1 个月前

Just shipped both new models in TypeWhisper: gpt-transcribe for completed recordings and gpt-live-transcribe for realtime dictation, with context, language hints, dictionary terms, and configurable delay. Available now in the OpenAI plugin 1.3.0.

Benedict Kerres 的头像
Benedict Kerres1 个月前

good name

frustrated by ice 的头像
frustrated by ice1 个月前

Do you have a demo app where I can provide an api key ?

Michael Wall 的头像
Michael Wall1 个月前

plugging them in right now

相关视频

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 次观看 • 7 个月前