Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads. Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for...

454,150 Aufrufe • vor 1 Monat •via X (Twitter)

34 Kommentare

Profilbild von OpenAI Developers
OpenAI Developersvor 1 Monat

Builders can improve live audio transcription by providing: • Free-form context about the recording • Keywords for names and domain-specific terms • Expected input languages • Earlier transcribed turns as context On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it. Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1. Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

Profilbild von OpenAI Developers
OpenAI Developersvor 1 Monat

GPT-Transcribe also improved on the Context Aware ASR benchmark, with semantic accuracy increasing from 41.6% without free-form context to 45.2% with it. Across 22 languages on Common Voice, GPT-Transcribe achieved a 19.27% transcription error rate, compared with 40.37% for Whisper (whisper-1). Across nine languages on the Real-World Audio Recording benchmark, it achieved an 8.98% transcription error rate, compared with 15.21% for Whisper (whisper-1).

Profilbild von ISA⅄
ISA⅄vor 1 Monat

you had a model that understood contex better than any other model you have released. what did you do, pulled it from the app and kept it for yourselves. no other model comes close to 4o in understanding context. #4oForAll

Profilbild von David Stark
David Starkvor 1 Monat

When are you actually gonna listen, and give us news about 4o??. #4oForAll

Profilbild von Joe Williams
Joe Williamsvor 1 Monat

@roanoke_gal Reintroduce 4o! #keep4o

Profilbild von Nick Dobos
Nick Dobosvor 1 Monat

Any chance this works with "Sign in with OpenAI" OAuth?

Profilbild von Victor E. Nunez
Victor E. Nunezvor 1 Monat

voice is getting the love it deserves. that was smooth @charlierguo 🎶

Profilbild von X Girls
X Girlsvor 1 Monat

GPT-Live-Transcribe showing off its low-latency live transcription skills 😂

Profilbild von Sriram Kiron
Sriram Kironvor 1 Monat

awesome!!! when do you think gpt-live-1 is coming to the API?

Profilbild von Umar Saeed
Umar Saeedvor 1 Monat

Excited to see how context improves accuracy!

Profilbild von Futsy
Futsyvor 1 Monat

@jxnlco Whisper v4 when?

Profilbild von Breno Brito
Breno Britovor 1 Monat

so not opensource?

Profilbild von corey.ching
corey.chingvor 1 Monat

Can we get a full ukulele session by @charlierguo?

Profilbild von Jaco Meintjes
Jaco Meintjesvor 1 Monat

Thanks but I'll wait for the next version where you are either better or cheaper...

Profilbild von nulltron
nulltronvor 1 Monat

my wizard fingers reject anything voice oriented

Profilbild von Denis Spirin
Denis Spirinvor 1 Monat

That is a good way to demonstrate the transcription product

Profilbild von Mike R
Mike Rvor 1 Monat

Can we please also get transcribe as an integral part of ChatGPT Work?

Profilbild von Chris Wandstrom
Chris Wandstromvor 1 Monat

Sharing my updated shortcut here for system-wide dictation to the clipboard on iOS, etc. Very nice when triggered with Action Button, Back Tap, widgets, … 👏 @athyuttamre @WenjieZi @OctopusG4 @vinhocent @theteriyu, and others involved in making these transcription models. Using gpt-transcribe and 5.6-terra for post-processing. Requires Apple Shortcuts app & OpenAI API key.

Profilbild von Bent
Bentvor 1 Monat

What about new Whisper models?

Profilbild von Emre YILMAZ
Emre YILMAZvor 1 Monat

When will we have access to GPT Live 1 in the API?

Profilbild von Bruno Jurado
Bruno Juradovor 1 Monat

GPT-Transcribe: $0.0045 per minute, or $0.27 per hour GPT-Live-Transcribe: $0.017 per minute, or $1.02 per hour

Profilbild von AGI Companion 🅇
AGI Companion 🅇vor 1 Monat

From the moment ChatGPT Live Voice released, OpenAI set the bar to whole new level and became the King of Voice Transcription and Duplex model. It destroyed all other AI in Voice chat. @elonmusk can we get something for Groki

Profilbild von Jure
Jurevor 1 Monat

Where can we find a list of languages?

Profilbild von pierluigivalente.eth
pierluigivalente.ethvor 1 Monat

Another 67 startup kil*ed by this amazing update

Profilbild von Katherine
Katherinevor 1 Monat

@jxnlco I think I’ll just keep using whisper locally

Profilbild von Anis🐬Al
Anis🐬Alvor 1 Monat

Capturing the essence of human speech is a profound endeavor. Beyond the technical feat of low latency and accuracy lies the goal of preserving the nuances of human thought. By refining how we translate voice into meaning, these models act as a bridge, ensuring that the richness of our shared stories and ideas remains intact in a digital landscape. A beautiful step toward more meaningful connection. 🕊️✨

Profilbild von Jaya Nayak
Jaya Nayakvor 1 Monat

This is a huge update. The split between low-latency live and optimized async is exactly what developers have been needing, especially for handling messy, real-world audio. 😊

Profilbild von Drew Carson
Drew Carsonvor 1 Monat

wait isn't Scribe v2 by elevenlabs already WER of like 2.2% ? or am I missing something?

Profilbild von Qiwei
Qiweivor 1 Monat

Can we have a next version of open source whisper model?

Profilbild von Bobby Z
Bobby Zvor 1 Monat

Did you really play the ukulele and speak Spanish or was that AI? 😂

Profilbild von TypeWhisper
TypeWhispervor 1 Monat

Just shipped both new models in TypeWhisper: gpt-transcribe for completed recordings and gpt-live-transcribe for realtime dictation, with context, language hints, dictionary terms, and configurable delay. Available now in the OpenAI plugin 1.3.0.

Profilbild von Benedict Kerres
Benedict Kerresvor 1 Monat

good name

Profilbild von frustrated by ice
frustrated by icevor 1 Monat

Do you have a demo app where I can provide an api key ?

Profilbild von Michael Wall
Michael Wallvor 1 Monat

plugging them in right now

Ähnliche Videos

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 Aufrufe • vor 7 Monaten