正在加载视频...

视频加载失败

Built a "YouTube realtime copilot" browser extension using OpenAI's realtime 2 API: The agent watches the video alongside you, and can answer any question you have about what was just said via realtime voice chat. The crazy part to me is: It can differentiate the YouTube's audio stream and...

45,631 次观看 • 4 个月前 •via X (Twitter)

34 条评论

程北FIG 的头像
程北FIG4 个月前

This feels like the right direction for AI video. Not just “summarize this video”, but watch with me, answer in context, and help me understand the part I just missed. The UX shift is from search-after-watching to assistance-while-watching.

Progress-means-a-lot 的头像
Progress-means-a-lot4 个月前

auto-pause when a user speaks so Copilot can engage seamlessly. Requiring manual intervention significantly degrades the experience; automation is key to overall quality.

Doxy 的头像
Doxy4 个月前

wait, does it actually parse the audio stream or just grab captions? curious if that separation trick holds up on chaotic tech talks

cesaer 的头像
cesaer4 个月前

on Github yet? need it!

Bakhty | 📣 Sarafan 的头像
Bakhty | 📣 Sarafan4 个月前

Kudos for a great addition to my list of best tools to learn languages 💪🏼🔥🚀

Devon Chaine 的头像
Devon Chaine4 个月前

This is not better than the YouTube so feature

Jingzi Zhao 的头像
Jingzi Zhao4 个月前

How much time do you spend vibe coding a product?

ORacle 的头像
ORacle4 个月前

太get這個邊看邊問的插件了!太強了,我幾天前還在盧頻扣字幕,現在已經能實時陪看了😂 我特別好奇這個插件從idea到上線要多久呢?對於普通vibecoder來說最難復製的是音頻流區分嗎?求指點

陈洁 的头像
陈洁4 个月前

真的是非常需要了哈哈!

Mengxue 的头像
Mengxue4 个月前

I would love to see this way of human-AI interaction in more personal, casual settings -- like watching a reality TV show together and chat about it like friends, which will definitely scratch an itch when my bf doesn't want to watch Love is Blind with me :)

Soroush Fadaeimanesh 的头像
Soroush Fadaeimanesh4 个月前

watching together is the killer pattern for realtime. way more useful than the agent watching alone and summarizing. what's the latency feeling like in practice?

Konrad Major 的头像
Konrad Major4 个月前

You can paste the link into Gemini amd ask questions about the video.

Gregor 的头像
Gregor4 个月前

Multimodal context window doing the heavy lifting there.

Truly 的头像
Truly4 个月前

Really nice!

Joe Hsu 的头像
Joe Hsu4 个月前

such good use of the realtime api! never thought of piping multiple sources

Daniel Bigham 的头像
Daniel Bigham4 个月前

This is great!

Don D'Cruz 的头像
Don D'Cruz4 个月前

The cost is gonna be crazy bro

TriciaAmazingyear 的头像
TriciaAmazingyear4 个月前

In the olden days your Mom or Dad and/or GrandMom did that 💋‼️😉

Victor 的头像
Victor4 个月前

passive watching is officially dead, every video is now a conversation

Albiona Hoti 的头像
Albiona Hoti4 个月前

Wow!!

yun long 的头像
yun long4 个月前

have the same idea!

Tessa Archer 的头像
Tessa Archer4 个月前

Audio stream differentiation without false triggers separates toys from tools. What latency are you hitting?

AI_lv0 的头像
AI_lv04 个月前

what is the cost of running this lol on Lex pod this could cost you a lot of money?!!!

zan 的头像
zan4 个月前

cool

MAX ONBOARDER ⭕ 的头像
MAX ONBOARDER ⭕4 个月前

Ohh this is amazing And can come in handy

0xmusashi 的头像
0xmusashi4 个月前

really cool idea zara 🔥

boilthesea 的头像
boilthesea4 个月前

How strong is that differentiation I wonder, would it hold up if a video maker intentionally tried to inject prompts? We could really use not confusing context with commands in literally every ai product.

Abhishek Kumar 的头像
Abhishek Kumar4 个月前

Great I am also thinking about buying a gun to kill a mosquito. And this line is about neither gun not mosquito.

William.K 的头像
William.K4 个月前

So cool! It just listen to the audio right? What about diagrams / visuals on the video?

Abhishek Kumar 的头像
Abhishek Kumar4 个月前

Why such an over-engineering solution? Even if you use gemini in chrome or perplexity comet, they both have this built in by default.

InternationalOptions 的头像
InternationalOptions4 个月前

@gabrielchua You’ve gotta gothub this!

BlanPlan 的头像
BlanPlan4 个月前

The voice/audio diarization piece is the part that quietly took years to get right. Tested OpenAI realtime 2 yesterday on background podcasts and it stayed silent through 90 minutes of video soundtrack until I directly addressed it. That separation is the real moat. Most realtime systems still spike on transient music.

J A Z I I 的头像
J A Z I I4 个月前

intresting how can i use it too>

Happy Monkey AI 的头像
Happy Monkey AI4 个月前

I've got a project on my to-do list that can extract key information and highlights from videos, though I was leaning into things like teams meetings over YouTube, using vision models and LLM, having something that can find and give me terminal commands used in a demo and so on

相关视频