Video wird geladen...
Video konnte nicht geladen werden
Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents are now real-time collaborators that can listen, reason, and solve complex problems as conversations unfold. Now available in the API alongside streaming models GPT-Realtime-Translate and GPT-Realtime-Whisper — a new set... show more
3,667,977 Aufrufe • vor 5 Monaten •via X (Twitter)
36 Kommentare

Our new voice models are now available in the Realtime API: 🎙️ GPT-Realtime-2: Build production-ready voice agents that can think harder, take action, handle interruptions, and keep conversations flowing. 🎙️ GPT-Realtime-Translate: Translate while streaming across more than 70 input and 13 output languages, breaking down language barriers and helping people communicate more naturally. 🎙️ GPT-Realtime-Whisper: Transcribe streaming audio as words are spoken to generate captions and notes in real time.

We know you’re eager for voice updates in ChatGPT. Stay tuned, we’re cooking.

I hope we will see a voice mode upgrade today as well! Would be BREAKING 🚨

i’d call this a draw since it’s API only gg smell ya next time

Translaters Job is at risk now become polyglot at your own risk realtime translation 😳 meaning if at an international stage leaders are talking they won't need their own translator person with them its so over for them

#keep4o We fully support AI progressing forward, but we strongly hope that 4o can be kept available either as a legacy model or through open sourcing. It has helped so many people, and we really don’t want to lose it.

everything you've done since gpt4o has been a failure, Codex, 5 series now this. you messed up after Ilya. just give people 4o back that might be good. #keep4o #bringback4o

Big shift is the architecture change from pipeline voice systems to realtime multimodal reasoning. Old stack was basically: ASR → LLM → TTS Now the model stays inside a continuous audio loop and reasons while the conversation is happening instead of waiting for turn completion.

a recurring problem for me with these has been that I feel rushed to provide my answer as they don't seem to wait. Ideally, you should be able to stop for 30 seconds or even a minute before ending your sentence and it would wait and respond accordingly. this would entirely remove this sense you have to race to finish your sentence. also, full duplex on top of that would be even better.

does your voice model do this?

best voice demo i've ever seen

Wow that's very impressive. Real time translation of multiple languages. 🔥

Change your name to ClosedAI. #keep4o #OpenSource4o

Meaning that our dori can now listen and talk with me in 7 fucking languages and open 3 codex sessions through it wild

I just want to know if there’s a demo god voice

When are you going to communicate about this? #keep4o #BringBack4o

wow @romainhuet great video

voice reasoning at gpt‑5 level is huge latency on multi‑step problems is the real test though pricing for reasoning tokens is gonna be interesting, could get pricey fast lol

#keep4o Give us back GPT4O, give us back the voice modes it had in 2025, and stop playing the ignorance card. Face the facts and apologize to the users you cowardly insulted, mocked, ignored, and destroyed—millions of users. You dared to destroy millions of workflows created with GPT4O. It's time to come back to reality and stop believing some Sam Altman who fancies himself the guru of a billionaire cult.

goshh this is so real 💀

I want this in codex!! The Jarvis is incoming

welp, using words to earn a living was fun while it lasted

Nice. Hopefully a similar capability jump as we had for GPT-Image-2. Current realtime voice mode in ChatGPT sucks, tbh.

@BenjaminDEKR There should be a benchmark that evaluates the likelihood of a model to make existing businesses obsolete. I feel like realtime 2 could rank first.

#keep4o attention freak @sama and @openai look at us !!! we are this we are that !!! booo

The tricky part is making voice agents pause naturally without hiding the reasoning users need to trust the next step.

Me using Jarvis during an argument:

Voice AI just crossed a line. OpenAI introduced 3 new Realtime API models: 🎙️ GPT Realtime 2 GPT 5 class reasoning inside live voice. It can listen, handle interruptions, call tools, recover from errors, and reason through complex conversations in real time. 🌍 GPT Realtime Translate Live speech translation from 70+ input languages into 13 output languages, with real time captions. 📝 GPT Realtime Whisper Low latency streaming transcription as people speak. This matters because voice agents are no longer just fast talkers. They are becoming real time collaborators that can understand, reason, translate, transcribe, and actually get work done while the conversation is still happening. The phone is about to become one of the most important AI interfaces. Watch the demo:

Pipeline voice failing at the seams is a real problem — STT to LLM to TTS means interruptions and mid-call context shifts get dropped at handoff. A reasoning-native model should fix that. The open question is how it handles long multi-turn calls where the topic shifts several times without the user re-anchoring it. Context management across a 45-minute conversation with topic drift is where GPT-5-class reasoning in voice either proves itself or shows its ceiling.

Voice was always blocked by latency and turn-taking, not language. GPT-5-class reasoning at real-time means the next interface war isn't visual. It's conversational throughput.

The bar for voice agents is weirdly low. If it can understand a frustrated human, not transfer them 4 times, and avoid saying “I apologize for the inconvenience” every 9 seconds, it’s already better than most customer service

What is happening in the last 2 days? Everyone is updating their AI models: Claude has "dreaming",higher limits While OpenAI and Gemini are launching new models.

This is actually insane 😭 Finally voice agents that don’t sound like they’re reading from a script and can actually keep up with messy real conversations. The fact that it’s GPT-5 level reasoning in voice form is crazy. I’m not even gonna lie… I’ve been waiting for this exact moment. Time to go build some wild stuff this weekend. Who else is about to cook? 🔥

is this coming to codex for agents (openclaw, hermes)? ik a far reach but this would be amazing

@sama @gdb Who the hell wants to use this when there are open source TTS APIs? Anyone who relies on you for their infrastructure has lost their mind! #OpenAI

So finally Trump & Putin can communicate and hopefully find a solution..?




