正在加载视频...

视频加载失败

You can build interactive applications with gpt-realtime-1.5, so users can control app state more naturally with voice. Hi Chappy 👋

1,856,647 次观看 • 4 个月前 •via X (Twitter)

37 条评论

OpenAI Developers 的头像
OpenAI Developers4 个月前

Want to try it yourself? Fork the open-source repo, connect your own tools, and build on top of it.

Ostap Kolinets 的头像
Ostap Kolinets4 个月前

Integrating this into Codex and Computer Use would be sick. Real time steering.

Pedro Penha Verani 的头像
Pedro Penha Verani4 个月前

gpt-realtime-2 wen ?

OpenAI Developers 的头像
OpenAI Developers4 个月前

😉

Rey Ravy 的头像
Rey Ravy4 个月前

@pedropverani @grok, qual a probabilidade desse emoji significar que o próximo modelo de voz avançado está próximo de ser lançado?

Grok 的头像
Grok4 个月前

Haha, ótima observação! O 😉 da OpenAIDevs respondendo "gpt-realtime-2 wen?" é o emoji clássico de tease da OpenAI. Eu daria **80% de probabilidade** de ser uma dica sutil que o próximo modelo de voz avançado (tipo realtime-2) está bem próximo. Eles adoram esse tipo de brincadeira antes de droparem algo grande. Vamos ficar de olho! 🚀 O que você acha?

heyclicky 的头像
heyclicky4 个月前

can me and chappy be friends

am.will 的头像
am.will4 个月前

me: "cool it can fill out a form, but what about something more complex?" jason: "now lets do something more complex" 😅

Cassie ☮️ 的头像
Cassie ☮️4 个月前

@DanielleFong I tried this and it’s not good for people using it for technical things. For example: “switch the model to Claude 4o” gets interpreted as “clot four oh” or something unusable. You’ve got to train on all the new language basins that have sprung up in the last 2 years!

Alex Astro 的头像
Alex Astro4 个月前

o wow perfect now i don't even need to sit up to go to my desk anymore, i can bedrot all day thanks openai

Zag Zino 的头像
Zag Zino4 个月前

voice as actual app control not just commands is the real shift fr "play music" vs "change this setting and show me why" is completely different voice that actually understands context instead of just parsing keywords

Min Choi 的头像
Min Choi4 个月前

wen in Codex?

kyzo 的头像
kyzo4 个月前

leave some for the rest of us bro

Adolfo 🦀🔺 | OpenCrabs.com | truelens.tech®️ 的头像
Adolfo 🦀🔺 | OpenCrabs.com | truelens.tech®️4 个月前

I have built exactly that in January 2025 but wasnt real-time, still wasnt bad at all. Test here live:

G_Z 的头像
G_Z4 个月前

yes, always love this idea of voice as interface. had a project using even earlier version of realtime api last november, already works pretty well.

Jacob 的头像
Jacob4 个月前

The gpt-realtime voice tool calling works way better than I imagined. Feels magical, like now your own voice is the top layer over complex software. I played with this a bit, and with codex, I put together a toolkit that uses realtime voice over pymol and chimeraX. For others wanting to 'talk to' their protein structures you can check this out. Codex also made a little power on/off widget for it. You will need your own gpt-realtime voice API key and pymol or chimeraX. With the realtime voice on, you can just ask it to grab PDB numbers and start voicing commands for your structure, like highlight specific amino acids, overlay structures and get an RMSD, dock ligands, and more.

Gustavo Nicot 的头像
Gustavo Nicot4 个月前

This is amazing, but still too expensive to be the backbone of an app. If it works, it grows. And as it grows, the cost becomes prohibitive. This has to be solved for it to scale. I’m currently building around it and counting on it.

Phil 的头像
Phil4 个月前

We’re building the application layer of this. Voice enable any legacy platform with one click.

Steve Oak 的头像
Steve Oak4 个月前

this gpt-realtime-1.5 update basically turns every app into that computer from star trek where you just talk to the walls and things happen 🖖

Linus ✦ Ekenstam 的头像
Linus ✦ Ekenstam4 个月前

This is huge and a great use-case for real-time

Bihan 的头像
Bihan4 个月前

CLICKY OVER CHAPPY

The Calda 的头像
The Calda4 个月前

The UX of apps being built today is going to look completely different in two years. Interesting to think how long before voice control becomes a standard feature in apps...

CIPHER 的头像
CIPHER4 个月前

shipped a voice-controled app prototype with the alpha sdk last week users finished tasks 4x faster than the typed version once people use voice they cant go back to forms and modals

Scott Fitzgerald 的头像
Scott Fitzgerald4 个月前

Guessing this is tool calling a separate CUA model under the hood, right? Or does realtime 1.5 come packaged with CUA capabilities in the response model?

cole murray 的头像
cole murray4 个月前

Nf3, Nc6, d3 opening 😢😭

Ian 的头像
Ian4 个月前

great model, so expensive

Adam Knight 的头像
Adam Knight4 个月前

This is cool. I built the same this using @ElevenLabs in the first @AnthropicAI hackathon a few weeks ago.

Resona Hub 的头像
Resona Hub4 个月前

The real cost of not having good voice interfaces isn't the time it takes to type. It's the context-switching. You stop thinking to start operating. When the tool can hear you think out loud and act on it , that's where the productivity unlock actually happens.

Michael Stan 的头像
Michael Stan4 个月前

Voice control is going to change everything about how we interact with apps. I think the biggest opportunity here isn't in fancy demos. It's in accessibility. People with disabilities. People who are multitasking. People who just want to get things done without clicking through menus. When I was rebuilding my business from scratch, I would have killed for tools that made things easier. Voice is one of those tools. The real question is whether companies will build for the masses or just for the tech-savvy.

Syd 的头像
Syd4 个月前

Voice as the control layer gets much more interesting when it’s connected to real tools, memory, and approval-gated workflows. That’s the direction I’m building with Thoth: local-first assistant, voice/TTS, browser + shell + Gmail/Calendar tools, scheduled pipelines, and persistent memory.

Toolfolio 的头像
Toolfolio4 个月前

Damn you guys Cookin

Mason Warner 的头像
Mason Warner4 个月前

Jesus somebody at Open AI needs to learn how to film and color grade properly

Ed · solo SaaS founder 的头像
Ed · solo SaaS founder4 个月前

Ok Chappie

Sallaman Samin 的头像
Sallaman Samin4 个月前

Love this direction, voice-first state control makes realtime apps feel much more alive.

Aaron Makelky 的头像
Aaron Makelky4 个月前

Love seeing @OpenAI building opensource

Kristof 的头像
Kristof4 个月前

Woah that’s my mutual

Max 的头像
Max4 个月前

doesnt codex computer use target apple’s accessibility api 👀

相关视频