Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

You can build interactive applications with gpt-realtime-1.5, so users can control app state more naturally with voice. Hi Chappy 👋

1,856,647 Aufrufe • vor 4 Monaten •via X (Twitter)

37 Kommentare

Profilbild von OpenAI Developers
OpenAI Developersvor 4 Monaten

Want to try it yourself? Fork the open-source repo, connect your own tools, and build on top of it.

Profilbild von Ostap Kolinets
Ostap Kolinetsvor 4 Monaten

Integrating this into Codex and Computer Use would be sick. Real time steering.

Profilbild von Pedro Penha Verani
Pedro Penha Veranivor 4 Monaten

gpt-realtime-2 wen ?

Profilbild von OpenAI Developers
OpenAI Developersvor 4 Monaten

😉

Profilbild von Rey Ravy
Rey Ravyvor 4 Monaten

@pedropverani @grok, qual a probabilidade desse emoji significar que o próximo modelo de voz avançado está próximo de ser lançado?

Profilbild von Grok
Grokvor 4 Monaten

Haha, ótima observação! O 😉 da OpenAIDevs respondendo "gpt-realtime-2 wen?" é o emoji clássico de tease da OpenAI. Eu daria **80% de probabilidade** de ser uma dica sutil que o próximo modelo de voz avançado (tipo realtime-2) está bem próximo. Eles adoram esse tipo de brincadeira antes de droparem algo grande. Vamos ficar de olho! 🚀 O que você acha?

Profilbild von heyclicky
heyclickyvor 4 Monaten

can me and chappy be friends

Profilbild von am.will
am.willvor 4 Monaten

me: "cool it can fill out a form, but what about something more complex?" jason: "now lets do something more complex" 😅

Profilbild von Cassie ☮️
Cassie ☮️vor 4 Monaten

@DanielleFong I tried this and it’s not good for people using it for technical things. For example: “switch the model to Claude 4o” gets interpreted as “clot four oh” or something unusable. You’ve got to train on all the new language basins that have sprung up in the last 2 years!

Profilbild von Alex Astro
Alex Astrovor 4 Monaten

o wow perfect now i don't even need to sit up to go to my desk anymore, i can bedrot all day thanks openai

Profilbild von Zag Zino
Zag Zinovor 4 Monaten

voice as actual app control not just commands is the real shift fr "play music" vs "change this setting and show me why" is completely different voice that actually understands context instead of just parsing keywords

Profilbild von Min Choi
Min Choivor 4 Monaten

wen in Codex?

Profilbild von kyzo
kyzovor 4 Monaten

leave some for the rest of us bro

Profilbild von Adolfo 🦀🔺 | OpenCrabs.com | truelens.tech®️
Adolfo 🦀🔺 | OpenCrabs.com | truelens.tech®️vor 4 Monaten

I have built exactly that in January 2025 but wasnt real-time, still wasnt bad at all. Test here live:

Profilbild von G_Z
G_Zvor 4 Monaten

yes, always love this idea of voice as interface. had a project using even earlier version of realtime api last november, already works pretty well.

Profilbild von Jacob
Jacobvor 4 Monaten

The gpt-realtime voice tool calling works way better than I imagined. Feels magical, like now your own voice is the top layer over complex software. I played with this a bit, and with codex, I put together a toolkit that uses realtime voice over pymol and chimeraX. For others wanting to 'talk to' their protein structures you can check this out. Codex also made a little power on/off widget for it. You will need your own gpt-realtime voice API key and pymol or chimeraX. With the realtime voice on, you can just ask it to grab PDB numbers and start voicing commands for your structure, like highlight specific amino acids, overlay structures and get an RMSD, dock ligands, and more.

Profilbild von Gustavo Nicot
Gustavo Nicotvor 4 Monaten

This is amazing, but still too expensive to be the backbone of an app. If it works, it grows. And as it grows, the cost becomes prohibitive. This has to be solved for it to scale. I’m currently building around it and counting on it.

Profilbild von Phil
Philvor 4 Monaten

We’re building the application layer of this. Voice enable any legacy platform with one click.

Profilbild von Steve Oak
Steve Oakvor 4 Monaten

this gpt-realtime-1.5 update basically turns every app into that computer from star trek where you just talk to the walls and things happen 🖖

Profilbild von Linus ✦ Ekenstam
Linus ✦ Ekenstamvor 4 Monaten

This is huge and a great use-case for real-time

Profilbild von Bihan
Bihanvor 4 Monaten

CLICKY OVER CHAPPY

Profilbild von The Calda
The Caldavor 4 Monaten

The UX of apps being built today is going to look completely different in two years. Interesting to think how long before voice control becomes a standard feature in apps...

Profilbild von CIPHER
CIPHERvor 4 Monaten

shipped a voice-controled app prototype with the alpha sdk last week users finished tasks 4x faster than the typed version once people use voice they cant go back to forms and modals

Profilbild von Scott Fitzgerald
Scott Fitzgeraldvor 4 Monaten

Guessing this is tool calling a separate CUA model under the hood, right? Or does realtime 1.5 come packaged with CUA capabilities in the response model?

Profilbild von cole murray
cole murrayvor 4 Monaten

Nf3, Nc6, d3 opening 😢😭

Profilbild von Ian
Ianvor 4 Monaten

great model, so expensive

Profilbild von Adam Knight
Adam Knightvor 4 Monaten

This is cool. I built the same this using @ElevenLabs in the first @AnthropicAI hackathon a few weeks ago.

Profilbild von Resona Hub
Resona Hubvor 4 Monaten

The real cost of not having good voice interfaces isn't the time it takes to type. It's the context-switching. You stop thinking to start operating. When the tool can hear you think out loud and act on it , that's where the productivity unlock actually happens.

Profilbild von Michael Stan
Michael Stanvor 4 Monaten

Voice control is going to change everything about how we interact with apps. I think the biggest opportunity here isn't in fancy demos. It's in accessibility. People with disabilities. People who are multitasking. People who just want to get things done without clicking through menus. When I was rebuilding my business from scratch, I would have killed for tools that made things easier. Voice is one of those tools. The real question is whether companies will build for the masses or just for the tech-savvy.

Profilbild von Syd
Sydvor 4 Monaten

Voice as the control layer gets much more interesting when it’s connected to real tools, memory, and approval-gated workflows. That’s the direction I’m building with Thoth: local-first assistant, voice/TTS, browser + shell + Gmail/Calendar tools, scheduled pipelines, and persistent memory.

Profilbild von Toolfolio
Toolfoliovor 4 Monaten

Damn you guys Cookin

Profilbild von Mason Warner
Mason Warnervor 4 Monaten

Jesus somebody at Open AI needs to learn how to film and color grade properly

Profilbild von Ed · solo SaaS founder
Ed · solo SaaS foundervor 4 Monaten

Ok Chappie

Profilbild von Sallaman Samin
Sallaman Saminvor 4 Monaten

Love this direction, voice-first state control makes realtime apps feel much more alive.

Profilbild von Aaron Makelky
Aaron Makelkyvor 4 Monaten

Love seeing @OpenAI building opensource

Profilbild von Kristof
Kristofvor 4 Monaten

Woah that’s my mutual

Profilbild von Max
Maxvor 4 Monaten

doesnt codex computer use target apple’s accessibility api 👀

Ähnliche Videos