Video yükleniyor...
Video Yüklenemedi
We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads. Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for... show more
454,150 görüntüleme • 1 ay önce •via X (Twitter)
34 Yorum

Builders can improve live audio transcription by providing: • Free-form context about the recording • Keywords for names and domain-specific terms • Expected input languages • Earlier transcribed turns as context On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it. Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1. Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

GPT-Transcribe also improved on the Context Aware ASR benchmark, with semantic accuracy increasing from 41.6% without free-form context to 45.2% with it. Across 22 languages on Common Voice, GPT-Transcribe achieved a 19.27% transcription error rate, compared with 40.37% for Whisper (whisper-1). Across nine languages on the Real-World Audio Recording benchmark, it achieved an 8.98% transcription error rate, compared with 15.21% for Whisper (whisper-1).

you had a model that understood contex better than any other model you have released. what did you do, pulled it from the app and kept it for yourselves. no other model comes close to 4o in understanding context. #4oForAll

When are you actually gonna listen, and give us news about 4o??. #4oForAll

@roanoke_gal Reintroduce 4o! #keep4o

Any chance this works with "Sign in with OpenAI" OAuth?

voice is getting the love it deserves. that was smooth @charlierguo 🎶

GPT-Live-Transcribe showing off its low-latency live transcription skills 😂

awesome!!! when do you think gpt-live-1 is coming to the API?

Excited to see how context improves accuracy!

@jxnlco Whisper v4 when?

so not opensource?

Can we get a full ukulele session by @charlierguo?

Thanks but I'll wait for the next version where you are either better or cheaper...

my wizard fingers reject anything voice oriented

That is a good way to demonstrate the transcription product

Can we please also get transcribe as an integral part of ChatGPT Work?
Sharing my updated shortcut here for system-wide dictation to the clipboard on iOS, etc. Very nice when triggered with Action Button, Back Tap, widgets, … 👏 @athyuttamre @WenjieZi @OctopusG4 @vinhocent @theteriyu, and others involved in making these transcription models. Using gpt-transcribe and 5.6-terra for post-processing. Requires Apple Shortcuts app & OpenAI API key.

What about new Whisper models?

When will we have access to GPT Live 1 in the API?

GPT-Transcribe: $0.0045 per minute, or $0.27 per hour GPT-Live-Transcribe: $0.017 per minute, or $1.02 per hour

From the moment ChatGPT Live Voice released, OpenAI set the bar to whole new level and became the King of Voice Transcription and Duplex model. It destroyed all other AI in Voice chat. @elonmusk can we get something for Groki

Where can we find a list of languages?

Another 67 startup kil*ed by this amazing update

@jxnlco I think I’ll just keep using whisper locally

Capturing the essence of human speech is a profound endeavor. Beyond the technical feat of low latency and accuracy lies the goal of preserving the nuances of human thought. By refining how we translate voice into meaning, these models act as a bridge, ensuring that the richness of our shared stories and ideas remains intact in a digital landscape. A beautiful step toward more meaningful connection. 🕊️✨

This is a huge update. The split between low-latency live and optimized async is exactly what developers have been needing, especially for handling messy, real-world audio. 😊

wait isn't Scribe v2 by elevenlabs already WER of like 2.2% ? or am I missing something?

Can we have a next version of open source whisper model?

Did you really play the ukulele and speak Spanish or was that AI? 😂

Just shipped both new models in TypeWhisper: gpt-transcribe for completed recordings and gpt-live-transcribe for realtime dictation, with context, language hints, dictionary terms, and configurable delay. Available now in the OpenAI plugin 1.3.0.

good name

Do you have a demo app where I can provide an api key ?

plugging them in right now
