Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Built a "YouTube realtime copilot" browser extension using OpenAI's realtime 2 API: The agent watches the video alongside you, and can answer any question you have about what was just said via realtime voice chat. The crazy part to me is: It can differentiate the YouTube's audio stream and...

45,631 görüntüleme • 4 ay önce •via X (Twitter)

34 Yorum

程北FIG profil fotoğrafı
程北FIG4 ay önce

This feels like the right direction for AI video. Not just “summarize this video”, but watch with me, answer in context, and help me understand the part I just missed. The UX shift is from search-after-watching to assistance-while-watching.

Progress-means-a-lot profil fotoğrafı
Progress-means-a-lot4 ay önce

auto-pause when a user speaks so Copilot can engage seamlessly. Requiring manual intervention significantly degrades the experience; automation is key to overall quality.

Doxy profil fotoğrafı
Doxy4 ay önce

wait, does it actually parse the audio stream or just grab captions? curious if that separation trick holds up on chaotic tech talks

cesaer profil fotoğrafı
cesaer4 ay önce

on Github yet? need it!

Bakhty | 📣 Sarafan profil fotoğrafı
Bakhty | 📣 Sarafan4 ay önce

Kudos for a great addition to my list of best tools to learn languages 💪🏼🔥🚀

Devon Chaine profil fotoğrafı
Devon Chaine4 ay önce

This is not better than the YouTube so feature

Jingzi Zhao profil fotoğrafı
Jingzi Zhao4 ay önce

How much time do you spend vibe coding a product?

ORacle profil fotoğrafı
ORacle4 ay önce

太get這個邊看邊問的插件了!太強了,我幾天前還在盧頻扣字幕,現在已經能實時陪看了😂 我特別好奇這個插件從idea到上線要多久呢?對於普通vibecoder來說最難復製的是音頻流區分嗎?求指點

陈洁 profil fotoğrafı
陈洁4 ay önce

真的是非常需要了哈哈!

Mengxue profil fotoğrafı
Mengxue4 ay önce

I would love to see this way of human-AI interaction in more personal, casual settings -- like watching a reality TV show together and chat about it like friends, which will definitely scratch an itch when my bf doesn't want to watch Love is Blind with me :)

Soroush Fadaeimanesh profil fotoğrafı
Soroush Fadaeimanesh4 ay önce

watching together is the killer pattern for realtime. way more useful than the agent watching alone and summarizing. what's the latency feeling like in practice?

Konrad Major profil fotoğrafı
Konrad Major4 ay önce

You can paste the link into Gemini amd ask questions about the video.

Gregor profil fotoğrafı
Gregor4 ay önce

Multimodal context window doing the heavy lifting there.

Truly profil fotoğrafı
Truly4 ay önce

Really nice!

Joe Hsu profil fotoğrafı
Joe Hsu4 ay önce

such good use of the realtime api! never thought of piping multiple sources

Daniel Bigham profil fotoğrafı
Daniel Bigham4 ay önce

This is great!

Don D'Cruz profil fotoğrafı
Don D'Cruz4 ay önce

The cost is gonna be crazy bro

TriciaAmazingyear profil fotoğrafı
TriciaAmazingyear4 ay önce

In the olden days your Mom or Dad and/or GrandMom did that 💋‼️😉

Victor profil fotoğrafı
Victor4 ay önce

passive watching is officially dead, every video is now a conversation

Albiona Hoti profil fotoğrafı
Albiona Hoti4 ay önce

Wow!!

yun long profil fotoğrafı
yun long4 ay önce

have the same idea!

Tessa Archer profil fotoğrafı
Tessa Archer4 ay önce

Audio stream differentiation without false triggers separates toys from tools. What latency are you hitting?

AI_lv0 profil fotoğrafı
AI_lv04 ay önce

what is the cost of running this lol on Lex pod this could cost you a lot of money?!!!

zan profil fotoğrafı
zan4 ay önce

cool

MAX ONBOARDER ⭕ profil fotoğrafı
MAX ONBOARDER ⭕4 ay önce

Ohh this is amazing And can come in handy

0xmusashi profil fotoğrafı
0xmusashi4 ay önce

really cool idea zara 🔥

boilthesea profil fotoğrafı
boilthesea4 ay önce

How strong is that differentiation I wonder, would it hold up if a video maker intentionally tried to inject prompts? We could really use not confusing context with commands in literally every ai product.

Abhishek Kumar profil fotoğrafı
Abhishek Kumar4 ay önce

Great I am also thinking about buying a gun to kill a mosquito. And this line is about neither gun not mosquito.

William.K profil fotoğrafı
William.K4 ay önce

So cool! It just listen to the audio right? What about diagrams / visuals on the video?

Abhishek Kumar profil fotoğrafı
Abhishek Kumar4 ay önce

Why such an over-engineering solution? Even if you use gemini in chrome or perplexity comet, they both have this built in by default.

InternationalOptions profil fotoğrafı
InternationalOptions4 ay önce

@gabrielchua You’ve gotta gothub this!

BlanPlan profil fotoğrafı
BlanPlan4 ay önce

The voice/audio diarization piece is the part that quietly took years to get right. Tested OpenAI realtime 2 yesterday on background podcasts and it stayed silent through 90 minutes of video soundtrack until I directly addressed it. That separation is the real moat. Most realtime systems still spike on transient music.

J A Z I I profil fotoğrafı
J A Z I I4 ay önce

intresting how can i use it too>

Happy Monkey AI profil fotoğrafı
Happy Monkey AI4 ay önce

I've got a project on my to-do list that can extract key information and highlights from videos, though I was leaning into things like teams meetings over YouTube, using vision models and LLM, having something that can find and give me terminal commands used in a demo and so on

Benzer Videolar