Video yükleniyor...
Video Yüklenemedi
GPT-Realtime 2 is the future of the operating system. I've been experimenting with it for a couple weeks now, and I gotta say, it's pretty gosh darn incredible. Opening apps, searching the web, even editing in Premiere. All with just my voice. And it only takes a few prompts... show more
607,765 görüntüleme • 3 ay önce •via X (Twitter)
35 Yorum

It's way harder for me to talk than to click.

Nice for people who like using voice to control… (not me)… Will be good when @neuralink allows us to think the action though

The thing is we won’t use voice to control existing user interfaces. The user interfaces will change completely. But people keep bolting voice onto existing UX. Why? Because designing the new operating system that’s voice is much harder. It will happen but it requires rethinking the computer OS.

Pretty neeto, but way too slow and token gobbly to be practical.

Still way too expensive in the stock config. I did a deep dive her tho which optimizes it down to cents per hour instead of tens of dollars.

voice-first OS is gonna break the moment you need precision edits. but cool that it works at all

@grok how much does it cost

we are pretty close to this😂

13 years later. Real life imitating art.

*navigates web via voice looks insanely time consuming and difficult* “Now that we’ve seen how easy it is to browse the web” Pretty cool though that the voice to action is so fast.

You’re making tons of video cuts here to make it look like it’s responding at significantly faster speeds than it actually is

This looks awesome, the future of OS. Now we just need someone to invent a periferal to use our brainwaves to speak to ai instead of voice.

Found it.

Never got into using voice for coding tho, the problem is my voice stream is forward only - not editable - compared to text, and any tts solution that allows editing is text in disguise. Also writing is auto reflective as I’m staring at what I’m writing and thinking about it.

I am so intrigued

awesome

Am I seeing @superwhisper here?

love it!!!

18:38 - I doubt the little latency is atleast seconds when it's faster if you do it rather than the back and forth and the anxiety to check what the agent is trying to do. But agreed it could become near instant soon (and curious how the tokenomics be then)

This is badass. Followed

@_EverythingAI had this already! Check it out its super cool

how is this different from claude cowork !?

Oh we know. All of our platforms (desktop and mobile) utilize Realtime-2. Computer Use isn’t practical for most users but access to their data is. ReSono Labs Voice is a full privacy ecosystem (local storage) for a Voice First platform (mobile/desktop)

yes, once can process locally both vision and voice. otherwise too expensive so ironically, ai bubble burst will make this viable as RAM/chip costs plummet

so many many many layers of translations supple for surrogates.

Your video literally unlocked my ability to build a voice controlled MLB dashboard Lots of bugs and annoyances and generally disobedient agent, but it went from 0->1 thanks @per_simmons_

Freeing up both hands completely? That would mean productivity would skyrocket!

Dude, you were like you were on cocaine! So fast! Very good video, thanks!

Getting celery man vibes

Insane!!!

can it MEOW??

The demo is cool, but it also highlights the gap for me: the commands have to be extremely explicit, like “open Spotify and play X artist.” That feels closer to dictating AppleScript/osascript than talking naturally. Impressive, but still more party trick than new OS.

You're a professional dawg

@grok how does GPT-Realtime 2 compare to Grok Voice Think Fast 1.0

probably the future for those who prefer this type of UX
