正在加载视频...
视频加载失败
You can build interactive applications with gpt-realtime-1.5, so users can control app state more naturally with voice. Hi Chappy 👋
37 条评论

Want to try it yourself? Fork the open-source repo, connect your own tools, and build on top of it.

Integrating this into Codex and Computer Use would be sick. Real time steering.

gpt-realtime-2 wen ?

😉

@pedropverani @grok, qual a probabilidade desse emoji significar que o próximo modelo de voz avançado está próximo de ser lançado?

Haha, ótima observação! O 😉 da OpenAIDevs respondendo "gpt-realtime-2 wen?" é o emoji clássico de tease da OpenAI. Eu daria **80% de probabilidade** de ser uma dica sutil que o próximo modelo de voz avançado (tipo realtime-2) está bem próximo. Eles adoram esse tipo de brincadeira antes de droparem algo grande. Vamos ficar de olho! 🚀 O que você acha?

can me and chappy be friends

me: "cool it can fill out a form, but what about something more complex?" jason: "now lets do something more complex" 😅

@DanielleFong I tried this and it’s not good for people using it for technical things. For example: “switch the model to Claude 4o” gets interpreted as “clot four oh” or something unusable. You’ve got to train on all the new language basins that have sprung up in the last 2 years!

o wow perfect now i don't even need to sit up to go to my desk anymore, i can bedrot all day thanks openai

voice as actual app control not just commands is the real shift fr "play music" vs "change this setting and show me why" is completely different voice that actually understands context instead of just parsing keywords

wen in Codex?

leave some for the rest of us bro

I have built exactly that in January 2025 but wasnt real-time, still wasnt bad at all. Test here live:

yes, always love this idea of voice as interface. had a project using even earlier version of realtime api last november, already works pretty well.

The gpt-realtime voice tool calling works way better than I imagined. Feels magical, like now your own voice is the top layer over complex software. I played with this a bit, and with codex, I put together a toolkit that uses realtime voice over pymol and chimeraX. For others wanting to 'talk to' their protein structures you can check this out. Codex also made a little power on/off widget for it. You will need your own gpt-realtime voice API key and pymol or chimeraX. With the realtime voice on, you can just ask it to grab PDB numbers and start voicing commands for your structure, like highlight specific amino acids, overlay structures and get an RMSD, dock ligands, and more.

This is amazing, but still too expensive to be the backbone of an app. If it works, it grows. And as it grows, the cost becomes prohibitive. This has to be solved for it to scale. I’m currently building around it and counting on it.

We’re building the application layer of this. Voice enable any legacy platform with one click.

this gpt-realtime-1.5 update basically turns every app into that computer from star trek where you just talk to the walls and things happen 🖖

This is huge and a great use-case for real-time

CLICKY OVER CHAPPY

The UX of apps being built today is going to look completely different in two years. Interesting to think how long before voice control becomes a standard feature in apps...

shipped a voice-controled app prototype with the alpha sdk last week users finished tasks 4x faster than the typed version once people use voice they cant go back to forms and modals

Guessing this is tool calling a separate CUA model under the hood, right? Or does realtime 1.5 come packaged with CUA capabilities in the response model?

Nf3, Nc6, d3 opening 😢😭

great model, so expensive

This is cool. I built the same this using @ElevenLabs in the first @AnthropicAI hackathon a few weeks ago.

The real cost of not having good voice interfaces isn't the time it takes to type. It's the context-switching. You stop thinking to start operating. When the tool can hear you think out loud and act on it , that's where the productivity unlock actually happens.

Voice control is going to change everything about how we interact with apps. I think the biggest opportunity here isn't in fancy demos. It's in accessibility. People with disabilities. People who are multitasking. People who just want to get things done without clicking through menus. When I was rebuilding my business from scratch, I would have killed for tools that made things easier. Voice is one of those tools. The real question is whether companies will build for the masses or just for the tech-savvy.

Voice as the control layer gets much more interesting when it’s connected to real tools, memory, and approval-gated workflows. That’s the direction I’m building with Thoth: local-first assistant, voice/TTS, browser + shell + Gmail/Calendar tools, scheduled pipelines, and persistent memory.

Damn you guys Cookin

Jesus somebody at Open AI needs to learn how to film and color grade properly

Ok Chappie

Love this direction, voice-first state control makes realtime apps feel much more alive.

Love seeing @OpenAI building opensource

Woah that’s my mutual

doesnt codex computer use target apple’s accessibility api 👀





