Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcription of completed audio files and batch workloads. Both models better understand context and deliver more accurate transcription on real world audio across accents and languages, including for...

454,150 görüntüleme • 1 ay önce •via X (Twitter)

34 Yorum

OpenAI Developers profil fotoğrafı
OpenAI Developers1 ay önce

Builders can improve live audio transcription by providing: • Free-form context about the recording • Keywords for names and domain-specific terms • Expected input languages • Earlier transcribed turns as context On our new Context Aware Automatic Speech Recognition benchmark, GPT-Live-Transcribe’s semantic accuracy increased from 38.5% without free-form context to 44.6% with it. Across 22 languages on Common Voice, GPT-Live-Transcribe achieved a 19.70% transcription error rate, compared with 20.33% for GPT-Realtime-Whisper-1. Across nine languages on the Real-World Audio Recording benchmark, it achieved a 9.60% transcription error rate, compared with 11.65% for GPT-Realtime-Whisper-1.

OpenAI Developers profil fotoğrafı
OpenAI Developers1 ay önce

GPT-Transcribe also improved on the Context Aware ASR benchmark, with semantic accuracy increasing from 41.6% without free-form context to 45.2% with it. Across 22 languages on Common Voice, GPT-Transcribe achieved a 19.27% transcription error rate, compared with 40.37% for Whisper (whisper-1). Across nine languages on the Real-World Audio Recording benchmark, it achieved an 8.98% transcription error rate, compared with 15.21% for Whisper (whisper-1).

ISA⅄ profil fotoğrafı
ISA⅄1 ay önce

you had a model that understood contex better than any other model you have released. what did you do, pulled it from the app and kept it for yourselves. no other model comes close to 4o in understanding context. #4oForAll

David Stark profil fotoğrafı
David Stark1 ay önce

When are you actually gonna listen, and give us news about 4o??. #4oForAll

Joe Williams profil fotoğrafı
Joe Williams1 ay önce

@roanoke_gal Reintroduce 4o! #keep4o

Nick Dobos profil fotoğrafı
Nick Dobos1 ay önce

Any chance this works with "Sign in with OpenAI" OAuth?

Victor E. Nunez profil fotoğrafı
Victor E. Nunez1 ay önce

voice is getting the love it deserves. that was smooth @charlierguo 🎶

X Girls profil fotoğrafı
X Girls1 ay önce

GPT-Live-Transcribe showing off its low-latency live transcription skills 😂

Sriram Kiron profil fotoğrafı
Sriram Kiron1 ay önce

awesome!!! when do you think gpt-live-1 is coming to the API?

Umar Saeed profil fotoğrafı
Umar Saeed1 ay önce

Excited to see how context improves accuracy!

Futsy profil fotoğrafı
Futsy1 ay önce

@jxnlco Whisper v4 when?

Breno Brito profil fotoğrafı
Breno Brito1 ay önce

so not opensource?

corey.ching profil fotoğrafı
corey.ching1 ay önce

Can we get a full ukulele session by @charlierguo?

Jaco Meintjes profil fotoğrafı
Jaco Meintjes1 ay önce

Thanks but I'll wait for the next version where you are either better or cheaper...

nulltron profil fotoğrafı
nulltron1 ay önce

my wizard fingers reject anything voice oriented

Denis Spirin profil fotoğrafı
Denis Spirin1 ay önce

That is a good way to demonstrate the transcription product

Mike R profil fotoğrafı
Mike R1 ay önce

Can we please also get transcribe as an integral part of ChatGPT Work?

Chris Wandstrom profil fotoğrafı
Chris Wandstrom1 ay önce

Sharing my updated shortcut here for system-wide dictation to the clipboard on iOS, etc. Very nice when triggered with Action Button, Back Tap, widgets, … 👏 @athyuttamre @WenjieZi @OctopusG4 @vinhocent @theteriyu, and others involved in making these transcription models. Using gpt-transcribe and 5.6-terra for post-processing. Requires Apple Shortcuts app & OpenAI API key.

Bent profil fotoğrafı
Bent1 ay önce

What about new Whisper models?

Emre YILMAZ profil fotoğrafı
Emre YILMAZ1 ay önce

When will we have access to GPT Live 1 in the API?

Bruno Jurado profil fotoğrafı
Bruno Jurado1 ay önce

GPT-Transcribe: $0.0045 per minute, or $0.27 per hour GPT-Live-Transcribe: $0.017 per minute, or $1.02 per hour

AGI Companion 🅇 profil fotoğrafı
AGI Companion 🅇1 ay önce

From the moment ChatGPT Live Voice released, OpenAI set the bar to whole new level and became the King of Voice Transcription and Duplex model. It destroyed all other AI in Voice chat. @elonmusk can we get something for Groki

Jure profil fotoğrafı
Jure1 ay önce

Where can we find a list of languages?

pierluigivalente.eth profil fotoğrafı
pierluigivalente.eth1 ay önce

Another 67 startup kil*ed by this amazing update

Katherine profil fotoğrafı
Katherine1 ay önce

@jxnlco I think I’ll just keep using whisper locally

Anis🐬Al profil fotoğrafı
Anis🐬Al1 ay önce

Capturing the essence of human speech is a profound endeavor. Beyond the technical feat of low latency and accuracy lies the goal of preserving the nuances of human thought. By refining how we translate voice into meaning, these models act as a bridge, ensuring that the richness of our shared stories and ideas remains intact in a digital landscape. A beautiful step toward more meaningful connection. 🕊️✨

Jaya Nayak profil fotoğrafı
Jaya Nayak1 ay önce

This is a huge update. The split between low-latency live and optimized async is exactly what developers have been needing, especially for handling messy, real-world audio. 😊

Drew Carson profil fotoğrafı
Drew Carson1 ay önce

wait isn't Scribe v2 by elevenlabs already WER of like 2.2% ? or am I missing something?

Qiwei profil fotoğrafı
Qiwei1 ay önce

Can we have a next version of open source whisper model?

Bobby Z profil fotoğrafı
Bobby Z1 ay önce

Did you really play the ukulele and speak Spanish or was that AI? 😂

TypeWhisper profil fotoğrafı
TypeWhisper1 ay önce

Just shipped both new models in TypeWhisper: gpt-transcribe for completed recordings and gpt-live-transcribe for realtime dictation, with context, language hints, dictionary terms, and configurable delay. Available now in the OpenAI plugin 1.3.0.

Benedict Kerres profil fotoğrafı
Benedict Kerres1 ay önce

good name

frustrated by ice profil fotoğrafı
frustrated by ice1 ay önce

Do you have a demo app where I can provide an api key ?

Michael Wall profil fotoğrafı
Michael Wall1 ay önce

plugging them in right now

Benzer Videolar

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 görüntüleme • 7 ay önce