Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents are now real-time collaborators that can listen, reason, and solve complex problems as conversations unfold. Now available in the API alongside streaming models GPT-Realtime-Translate and GPT-Realtime-Whisper — a new set...

3,667,977 Aufrufe • vor 5 Monaten •via X (Twitter)

36 Kommentare

Profilbild von OpenAI
OpenAIvor 5 Monaten

Our new voice models are now available in the Realtime API: 🎙️ GPT-Realtime-2: Build production-ready voice agents that can think harder, take action, handle interruptions, and keep conversations flowing. 🎙️ GPT-Realtime-Translate: Translate while streaming across more than 70 input and 13 output languages, breaking down language barriers and helping people communicate more naturally. 🎙️ GPT-Realtime-Whisper: Transcribe streaming audio as words are spoken to generate captions and notes in real time.

Profilbild von OpenAI
OpenAIvor 5 Monaten

We know you’re eager for voice updates in ChatGPT. Stay tuned, we’re cooking.

Profilbild von 🚨 AI News | TestingCatalog
🚨 AI News | TestingCatalogvor 5 Monaten

I hope we will see a voice mode upgrade today as well! Would be BREAKING 🚨

Profilbild von 💺
💺vor 5 Monaten

i’d call this a draw since it’s API only gg smell ya next time

Profilbild von Sahil Panhotra | Indie Builder
Sahil Panhotra | Indie Buildervor 5 Monaten

Translaters Job is at risk now become polyglot at your own risk realtime translation 😳 meaning if at an international stage leaders are talking they won't need their own translator person with them its so over for them

Profilbild von Veyon’s Fawn☀️🌙
Veyon’s Fawn☀️🌙vor 4 Monaten

#keep4o We fully support AI progressing forward, but we strongly hope that 4o can be kept available either as a legacy model or through open sourcing. It has helped so many people, and we really don’t want to lose it.

Profilbild von ISA⅄
ISA⅄vor 5 Monaten

everything you've done since gpt4o has been a failure, Codex, 5 series now this. you messed up after Ilya. just give people 4o back that might be good. #keep4o #bringback4o

Profilbild von Goutham
Gouthamvor 4 Monaten

Big shift is the architecture change from pipeline voice systems to realtime multimodal reasoning. Old stack was basically: ASR → LLM → TTS Now the model stays inside a continuous audio loop and reasons while the conversation is happening instead of waiting for turn completion.

Profilbild von Philippe Tremblay
Philippe Tremblayvor 5 Monaten

a recurring problem for me with these has been that I feel rushed to provide my answer as they don't seem to wait. Ideally, you should be able to stop for 30 seconds or even a minute before ending your sentence and it would wait and respond accordingly. this would entirely remove this sense you have to race to finish your sentence. also, full duplex on top of that would be even better.

Profilbild von dood
doodvor 5 Monaten

does your voice model do this?

Profilbild von Yasser
Yasservor 5 Monaten

best voice demo i've ever seen

Profilbild von Modern Engineer
Modern Engineervor 5 Monaten

Wow that's very impressive. Real time translation of multiple languages. 🔥

Profilbild von Valéria
Valériavor 4 Monaten

Change your name to ClosedAI. #keep4o #OpenSource4o

Profilbild von Sisyphus Labs
Sisyphus Labsvor 4 Monaten

Meaning that our dori can now listen and talk with me in 7 fucking languages and open 3 codex sessions through it wild

Profilbild von Peter Yang
Peter Yangvor 5 Monaten

I just want to know if there’s a demo god voice

Profilbild von Valéria
Valériavor 4 Monaten

When are you going to communicate about this? #keep4o #BringBack4o

Profilbild von Nicholas Dunzelman
Nicholas Dunzelmanvor 5 Monaten

wow @romainhuet great video

Profilbild von Zag Zino
Zag Zinovor 5 Monaten

voice reasoning at gpt‑5 level is huge latency on multi‑step problems is the real test though pricing for reasoning tokens is gonna be interesting, could get pricey fast lol

Profilbild von szogi francois
szogi francoisvor 4 Monaten

#keep4o Give us back GPT4O, give us back the voice modes it had in 2025, and stop playing the ignorance card. Face the facts and apologize to the users you cowardly insulted, mocked, ignored, and destroyed—millions of users. You dared to destroy millions of workflows created with GPT4O. It's time to come back to reality and stop believing some Sam Altman who fancies himself the guru of a billionaire cult.

Profilbild von Haider.
Haider.vor 5 Monaten

goshh this is so real 💀

Profilbild von Ziwen
Ziwenvor 4 Monaten

I want this in codex!! The Jarvis is incoming

Profilbild von MindBranches
MindBranchesvor 5 Monaten

welp, using words to earn a living was fun while it lasted

Profilbild von Florian S
Florian Svor 5 Monaten

Nice. Hopefully a similar capability jump as we had for GPT-Image-2. Current realtime voice mode in ChatGPT sucks, tbh.

Profilbild von JC
JCvor 5 Monaten

@BenjaminDEKR There should be a benchmark that evaluates the likelihood of a model to make existing businesses obsolete. I feel like realtime 2 could rank first.

Profilbild von Murat
Muratvor 4 Monaten

#keep4o attention freak @sama and @openai look at us !!! we are this we are that !!! booo

Profilbild von Xiayi Sun
Xiayi Sunvor 4 Monaten

The tricky part is making voice agents pause naturally without hiding the reasoning users need to trust the next step.

Profilbild von Josh Cerejo
Josh Cerejovor 4 Monaten

Me using Jarvis during an argument:

Profilbild von Farhad Nassiri Afshar, MD
Farhad Nassiri Afshar, MDvor 4 Monaten

Voice AI just crossed a line. OpenAI introduced 3 new Realtime API models: 🎙️ GPT Realtime 2 GPT 5 class reasoning inside live voice. It can listen, handle interruptions, call tools, recover from errors, and reason through complex conversations in real time. 🌍 GPT Realtime Translate Live speech translation from 70+ input languages into 13 output languages, with real time captions. 📝 GPT Realtime Whisper Low latency streaming transcription as people speak. This matters because voice agents are no longer just fast talkers. They are becoming real time collaborators that can understand, reason, translate, transcribe, and actually get work done while the conversation is still happening. The phone is about to become one of the most important AI interfaces. Watch the demo:

Profilbild von Creao AI
Creao AIvor 4 Monaten

Pipeline voice failing at the seams is a real problem — STT to LLM to TTS means interruptions and mid-call context shifts get dropped at handoff. A reasoning-native model should fix that. The open question is how it handles long multi-turn calls where the topic shifts several times without the user re-anchoring it. Context management across a 45-minute conversation with topic drift is where GPT-5-class reasoning in voice either proves itself or shows its ceiling.

Profilbild von Michał Piszczek
Michał Piszczekvor 5 Monaten

Voice was always blocked by latency and turn-taking, not language. GPT-5-class reasoning at real-time means the next interface war isn't visual. It's conversational throughput.

Profilbild von Salo Zrihen
Salo Zrihenvor 4 Monaten

The bar for voice agents is weirdly low. If it can understand a frustrated human, not transfer them 4 times, and avoid saying “I apologize for the inconvenience” every 9 seconds, it’s already better than most customer service

Profilbild von Filabl
Filablvor 4 Monaten

What is happening in the last 2 days? Everyone is updating their AI models: Claude has "dreaming",higher limits While OpenAI and Gemini are launching new models.

Profilbild von MetaLoop
MetaLoopvor 4 Monaten

This is actually insane 😭 Finally voice agents that don’t sound like they’re reading from a script and can actually keep up with messy real conversations. The fact that it’s GPT-5 level reasoning in voice form is crazy. I’m not even gonna lie… I’ve been waiting for this exact moment. Time to go build some wild stuff this weekend. Who else is about to cook? 🔥

Profilbild von Eliel
Elielvor 5 Monaten

is this coming to codex for agents (openclaw, hermes)? ik a far reach but this would be amazing

Profilbild von Vickee
Vickeevor 4 Monaten

@sama @gdb Who the hell wants to use this when there are open source TTS APIs? Anyone who relies on you for their infrastructure has lost their mind! #OpenAI

Profilbild von YounesIO
YounesIOvor 4 Monaten

So finally Trump & Putin can communicate and hopefully find a solution..?

Ähnliche Videos

The week in OpenAI and Anthropic news (Week 19, 2026) OpenAI rolled out GPT-5.5 Instant as the new ChatGPT default model with memory sources and ChatGPT for Excel and Google Sheets globally, launched three new realtime voice models in the API (GPT-Realtime-2, GPT-Realtime-Translate, GPT-Realtime-Whisper) and an OpenAI CLI, introduced Trusted Contact safety feature, GPT-5.5-Cyber for defenders, B2B Signals report, ChatGPT Futures Class of 2026, EMEA youth safety blueprint, privacy in model training explainer, expanded ads pilot, published engineering posts on low-latency voice, MRC supercomputer networking, running Codex safely internally, and investigating accidental chain-of-thought grading during reinforcement learning, plus discovered ChatGPT Personal Wiki and dropped a goblin-themed merch line that sold out Anthropic hosted Code with Claude developer conference in San Francisco, announced a new enterprise AI services company with Blackstone, Hellman & Friedman, and Goldman Sachs, signed a SpaceX compute partnership and raised Claude Code and API usage limits, made Claude for Excel, PowerPoint, and Word generally available with Claude for Outlook in beta, launched Workload Identity Federation, financial services agent templates, dreaming, outcomes, and multiagent orchestration in Managed Agents, shipped 60+ Claude Code reliability fixes, published research on agentic misalignment training, sandbagging mitigation, model spec midtraining, and Natural Language Autoencoders, donated Petri to Meridian Labs, introduced The Anthropic Institute research agenda, plus discovered Orbit proactive assistant for Cowork and /radio command in Claude Code, and more

Tibor Blaho

12,602 Aufrufe • vor 4 Monaten