Loading video...

Video Failed to Load

Go Home

You can build interactive applications with gpt-realtime-1.5, so users can control app state more naturally with voice. Hi Chappy 👋

1,856,647 views • 4 months ago •via X (Twitter)

37 Comments

OpenAI Developers's profile picture
OpenAI Developers4 months ago

Want to try it yourself? Fork the open-source repo, connect your own tools, and build on top of it.

Ostap Kolinets's profile picture
Ostap Kolinets4 months ago

Integrating this into Codex and Computer Use would be sick. Real time steering.

Pedro Penha Verani's profile picture
Pedro Penha Verani4 months ago

gpt-realtime-2 wen ?

OpenAI Developers's profile picture
OpenAI Developers4 months ago

😉

Rey Ravy's profile picture
Rey Ravy4 months ago

@pedropverani @grok, qual a probabilidade desse emoji significar que o próximo modelo de voz avançado está próximo de ser lançado?

Grok's profile picture
Grok4 months ago

Haha, ótima observação! O 😉 da OpenAIDevs respondendo "gpt-realtime-2 wen?" é o emoji clássico de tease da OpenAI. Eu daria **80% de probabilidade** de ser uma dica sutil que o próximo modelo de voz avançado (tipo realtime-2) está bem próximo. Eles adoram esse tipo de brincadeira antes de droparem algo grande. Vamos ficar de olho! 🚀 O que você acha?

heyclicky's profile picture
heyclicky4 months ago

can me and chappy be friends

am.will's profile picture
am.will4 months ago

me: "cool it can fill out a form, but what about something more complex?" jason: "now lets do something more complex" 😅

Cassie ☮️'s profile picture
Cassie ☮️4 months ago

@DanielleFong I tried this and it’s not good for people using it for technical things. For example: “switch the model to Claude 4o” gets interpreted as “clot four oh” or something unusable. You’ve got to train on all the new language basins that have sprung up in the last 2 years!

Alex Astro's profile picture
Alex Astro4 months ago

o wow perfect now i don't even need to sit up to go to my desk anymore, i can bedrot all day thanks openai

Zag Zino's profile picture
Zag Zino4 months ago

voice as actual app control not just commands is the real shift fr "play music" vs "change this setting and show me why" is completely different voice that actually understands context instead of just parsing keywords

Min Choi's profile picture
Min Choi4 months ago

wen in Codex?

kyzo's profile picture
kyzo4 months ago

leave some for the rest of us bro

Adolfo 🦀🔺 | OpenCrabs.com | truelens.tech®️'s profile picture
Adolfo 🦀🔺 | OpenCrabs.com | truelens.tech®️4 months ago

I have built exactly that in January 2025 but wasnt real-time, still wasnt bad at all. Test here live:

G_Z's profile picture
G_Z4 months ago

yes, always love this idea of voice as interface. had a project using even earlier version of realtime api last november, already works pretty well.

Jacob's profile picture
Jacob4 months ago

The gpt-realtime voice tool calling works way better than I imagined. Feels magical, like now your own voice is the top layer over complex software. I played with this a bit, and with codex, I put together a toolkit that uses realtime voice over pymol and chimeraX. For others wanting to 'talk to' their protein structures you can check this out. Codex also made a little power on/off widget for it. You will need your own gpt-realtime voice API key and pymol or chimeraX. With the realtime voice on, you can just ask it to grab PDB numbers and start voicing commands for your structure, like highlight specific amino acids, overlay structures and get an RMSD, dock ligands, and more.

Gustavo Nicot's profile picture
Gustavo Nicot4 months ago

This is amazing, but still too expensive to be the backbone of an app. If it works, it grows. And as it grows, the cost becomes prohibitive. This has to be solved for it to scale. I’m currently building around it and counting on it.

Phil's profile picture
Phil4 months ago

We’re building the application layer of this. Voice enable any legacy platform with one click.

Steve Oak's profile picture
Steve Oak4 months ago

this gpt-realtime-1.5 update basically turns every app into that computer from star trek where you just talk to the walls and things happen 🖖

Linus ✦ Ekenstam's profile picture
Linus ✦ Ekenstam4 months ago

This is huge and a great use-case for real-time

Bihan's profile picture
Bihan4 months ago

CLICKY OVER CHAPPY

The Calda's profile picture
The Calda4 months ago

The UX of apps being built today is going to look completely different in two years. Interesting to think how long before voice control becomes a standard feature in apps...

CIPHER's profile picture
CIPHER4 months ago

shipped a voice-controled app prototype with the alpha sdk last week users finished tasks 4x faster than the typed version once people use voice they cant go back to forms and modals

Scott Fitzgerald's profile picture
Scott Fitzgerald4 months ago

Guessing this is tool calling a separate CUA model under the hood, right? Or does realtime 1.5 come packaged with CUA capabilities in the response model?

cole murray's profile picture
cole murray4 months ago

Nf3, Nc6, d3 opening 😢😭

Ian's profile picture
Ian4 months ago

great model, so expensive

Adam Knight's profile picture
Adam Knight4 months ago

This is cool. I built the same this using @ElevenLabs in the first @AnthropicAI hackathon a few weeks ago.

Resona Hub's profile picture
Resona Hub4 months ago

The real cost of not having good voice interfaces isn't the time it takes to type. It's the context-switching. You stop thinking to start operating. When the tool can hear you think out loud and act on it , that's where the productivity unlock actually happens.

Michael Stan's profile picture
Michael Stan4 months ago

Voice control is going to change everything about how we interact with apps. I think the biggest opportunity here isn't in fancy demos. It's in accessibility. People with disabilities. People who are multitasking. People who just want to get things done without clicking through menus. When I was rebuilding my business from scratch, I would have killed for tools that made things easier. Voice is one of those tools. The real question is whether companies will build for the masses or just for the tech-savvy.

Syd's profile picture
Syd4 months ago

Voice as the control layer gets much more interesting when it’s connected to real tools, memory, and approval-gated workflows. That’s the direction I’m building with Thoth: local-first assistant, voice/TTS, browser + shell + Gmail/Calendar tools, scheduled pipelines, and persistent memory.

Toolfolio's profile picture
Toolfolio4 months ago

Damn you guys Cookin

Mason Warner's profile picture
Mason Warner4 months ago

Jesus somebody at Open AI needs to learn how to film and color grade properly

Ed · solo SaaS founder's profile picture
Ed · solo SaaS founder4 months ago

Ok Chappie

Sallaman Samin's profile picture
Sallaman Samin4 months ago

Love this direction, voice-first state control makes realtime apps feel much more alive.

Aaron Makelky's profile picture
Aaron Makelky4 months ago

Love seeing @OpenAI building opensource

Kristof's profile picture
Kristof4 months ago

Woah that’s my mutual

Max's profile picture
Max4 months ago

doesnt codex computer use target apple’s accessibility api 👀

Related Videos