Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

GPT-Realtime 2 is the future of the operating system. I've been experimenting with it for a couple weeks now, and I gotta say, it's pretty gosh darn incredible. Opening apps, searching the web, even editing in Premiere. All with just my voice. And it only takes a few prompts...

607,765 görüntüleme • 3 ay önce •via X (Twitter)

35 Yorum

Masih profil fotoğrafı
Masih3 ay önce

It's way harder for me to talk than to click.

Jimmy profil fotoğrafı
Jimmy3 ay önce

Nice for people who like using voice to control… (not me)… Will be good when @neuralink allows us to think the action though

COURT profil fotoğrafı
COURT3 ay önce

The thing is we won’t use voice to control existing user interfaces. The user interfaces will change completely. But people keep bolting voice onto existing UX. Why? Because designing the new operating system that’s voice is much harder. It will happen but it requires rethinking the computer OS.

Raspy Camper🌲 profil fotoğrafı
Raspy Camper🌲3 ay önce

Pretty neeto, but way too slow and token gobbly to be practical.

Slopware Engineer profil fotoğrafı
Slopware Engineer3 ay önce

Still way too expensive in the stock config. I did a deep dive her tho which optimizes it down to cents per hour instead of tens of dollars.

Samian profil fotoğrafı
Samian3 ay önce

voice-first OS is gonna break the moment you need precision edits. but cool that it works at all

𝑯𝒊𝒃𝒓𝒂 profil fotoğrafı
𝑯𝒊𝒃𝒓𝒂3 ay önce

@grok how much does it cost

Steve Li profil fotoğrafı
Steve Li3 ay önce

we are pretty close to this😂

pystar profil fotoğrafı
pystar3 ay önce

13 years later. Real life imitating art.

JMT profil fotoğrafı
JMT3 ay önce

*navigates web via voice looks insanely time consuming and difficult* “Now that we’ve seen how easy it is to browse the web” Pretty cool though that the voice to action is so fast.

Steve Bauman profil fotoğrafı
Steve Bauman3 ay önce

You’re making tons of video cuts here to make it look like it’s responding at significantly faster speeds than it actually is

Lorenzo Nuvoletta profil fotoğrafı
Lorenzo Nuvoletta3 ay önce

This looks awesome, the future of OS. Now we just need someone to invent a periferal to use our brainwaves to speak to ai instead of voice.

Stephen Turner 🇬🇧🇺🇦 profil fotoğrafı
Stephen Turner 🇬🇧🇺🇦3 ay önce

Found it.

George Ramonov profil fotoğrafı
George Ramonov3 ay önce

Never got into using voice for coding tho, the problem is my voice stream is forward only - not editable - compared to text, and any tts solution that allows editing is text in disguise. Also writing is auto reflective as I’m staring at what I’m writing and thinking about it.

Ant's AI Lab profil fotoğrafı
Ant's AI Lab2 ay önce

I am so intrigued

Michael Wall profil fotoğrafı
Michael Wall3 ay önce

awesome

Raf V. profil fotoğrafı
Raf V.3 ay önce

Am I seeing @superwhisper here?

Mouad profil fotoğrafı
Mouad3 ay önce

love it!!!

Pavan profil fotoğrafı
Pavan3 ay önce

18:38 - I doubt the little latency is atleast seconds when it's faster if you do it rather than the back and forth and the anxiety to check what the agent is trying to do. But agreed it could become near instant soon (and curious how the tokenomics be then)

JP profil fotoğrafı
JP3 ay önce

This is badass. Followed

Nobody profil fotoğrafı
Nobody3 ay önce

@_EverythingAI had this already! Check it out its super cool

🅾️K profil fotoğrafı
🅾️K3 ay önce

how is this different from claude cowork !?

Christian McGivern profil fotoğrafı
Christian McGivern3 ay önce

Oh we know. All of our platforms (desktop and mobile) utilize Realtime-2. Computer Use isn’t practical for most users but access to their data is. ReSono Labs Voice is a full privacy ecosystem (local storage) for a Voice First platform (mobile/desktop)

nftsasha profil fotoğrafı
nftsasha3 ay önce

yes, once can process locally both vision and voice. otherwise too expensive so ironically, ai bubble burst will make this viable as RAM/chip costs plummet

Art Seabra 🫆 profil fotoğrafı
Art Seabra 🫆3 ay önce

so many many many layers of translations supple for surrogates.

DesignedByDogs profil fotoğrafı
DesignedByDogs2 ay önce

Your video literally unlocked my ability to build a voice controlled MLB dashboard Lots of bugs and annoyances and generally disobedient agent, but it went from 0->1 thanks @per_simmons_

Rolanddaan profil fotoğrafı
Rolanddaan3 ay önce

Freeing up both hands completely? That would mean productivity would skyrocket!

Andreas T. profil fotoğrafı
Andreas T.3 ay önce

Dude, you were like you were on cocaine! So fast! Very good video, thanks!

gino profil fotoğrafı
gino3 ay önce

Getting celery man vibes

Rafael Grossmann, MD, MSHS, FACS 🇻🇪🇺🇸 profil fotoğrafı
Rafael Grossmann, MD, MSHS, FACS 🇻🇪🇺🇸3 ay önce

Insane!!!

Kylan profil fotoğrafı
Kylan3 ay önce

can it MEOW??

nick profil fotoğrafı
nick3 ay önce

The demo is cool, but it also highlights the gap for me: the commands have to be extremely explicit, like “open Spotify and play X artist.” That feels closer to dictating AppleScript/osascript than talking naturally. Impressive, but still more party trick than new OS.

Io profil fotoğrafı
Io3 ay önce

You're a professional dawg

Kenny profil fotoğrafı
Kenny3 ay önce

@grok how does GPT-Realtime 2 compare to Grok Voice Think Fast 1.0

Will Haver profil fotoğrafı
Will Haver3 ay önce

probably the future for those who prefer this type of UX

Benzer Videolar

Nat Eliason’s (Nat Eliason) career arc is borderline absurd—but it works. He’ll spot a new tool or trend, master it, build a business around it, and move on. Nat’s pulled it off with the note-taking wave ($600k in sales from a Roam Research course), real estate (6x return flipping property in Austin), and crypto (published his insider story with Random House). Now it’s AI: he’s running a viral course on building apps with AI—$200k in pre-sales in just a week, 800 students and counting. I’ve known Nat for a long time and I think he has a great sense for where the puck is headed. He was one of the first guests I had on the podcast and I was delighted to have him on again. Here are a few takeaways from our conversation: - Coding with AI has become orders of magnitude easier for non-technical people over the last 2 years—Nat rarely has to help students fix bugs; they troubleshoot in Cursor on their own. - AI coding assistants are creating new behaviours in programming, like using a speech-to-text model to talk to an agent and having it write code for you. - The traditional learning curve of coding is flattening because AI tools let beginners build and iterate in faster feedback loops. - AI has given Nat leverage in spades—it increases his ability to be a creator while also building a robust business with as few people to manage as possible. He demos an AI book editor he coded for his sci-fi novel. - In the age of AI, software is becoming content and the barriers to create are lower than ever—but custom software for everything isn’t the answer. Nat’s model is that personalized tools make sense for that one thing you care the most about. - Nat believes that the future of writing with AI is a Cursor-style interface with a model that’s trained on your style and voice. This episode is a must-watch for writers, creators, and anyone interested in the future of product building. Watch below! Timestamps: Introduction: 00:01:45 The origins of Nat’s viral course on building apps with AI: 00:11:45 How coding with AI has evolved over the last two years: 00:18:46 Nat creates an app using Composer, Cursor’s AI assistant: 00:22:22 Tactical tips for coding with Cursor: 00:26:06 How coding with AI is creating new behaviours in programming: 00:29:06 What excites Nat the most about the future of AI: 00:32:41 A demo of Hubbard, the AI editor Nat built for his science fiction writing: 00:38:58 When does it makes sense to build custom software: 00:44:52 Nat’s take on the future of writing with AI: 00:49:18

Dan Shipper 📧

27,207 görüntüleme • 1 yıl önce