Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Built a "YouTube realtime copilot" browser extension using OpenAI's realtime 2 API: The agent watches the video alongside you, and can answer any question you have about what was just said via realtime voice chat. The crazy part to me is: It can differentiate the YouTube's audio stream and...

45,631 Aufrufe • vor 4 Monaten •via X (Twitter)

34 Kommentare

Profilbild von 程北FIG
程北FIGvor 4 Monaten

This feels like the right direction for AI video. Not just “summarize this video”, but watch with me, answer in context, and help me understand the part I just missed. The UX shift is from search-after-watching to assistance-while-watching.

Profilbild von Progress-means-a-lot
Progress-means-a-lotvor 4 Monaten

auto-pause when a user speaks so Copilot can engage seamlessly. Requiring manual intervention significantly degrades the experience; automation is key to overall quality.

Profilbild von Doxy
Doxyvor 4 Monaten

wait, does it actually parse the audio stream or just grab captions? curious if that separation trick holds up on chaotic tech talks

Profilbild von cesaer
cesaervor 4 Monaten

on Github yet? need it!

Profilbild von Bakhty | 📣 Sarafan
Bakhty | 📣 Sarafanvor 4 Monaten

Kudos for a great addition to my list of best tools to learn languages 💪🏼🔥🚀

Profilbild von Devon Chaine
Devon Chainevor 4 Monaten

This is not better than the YouTube so feature

Profilbild von Jingzi Zhao
Jingzi Zhaovor 4 Monaten

How much time do you spend vibe coding a product?

Profilbild von ORacle
ORaclevor 4 Monaten

太get這個邊看邊問的插件了!太強了,我幾天前還在盧頻扣字幕,現在已經能實時陪看了😂 我特別好奇這個插件從idea到上線要多久呢?對於普通vibecoder來說最難復製的是音頻流區分嗎?求指點

Profilbild von 陈洁
陈洁vor 4 Monaten

真的是非常需要了哈哈!

Profilbild von Mengxue
Mengxuevor 4 Monaten

I would love to see this way of human-AI interaction in more personal, casual settings -- like watching a reality TV show together and chat about it like friends, which will definitely scratch an itch when my bf doesn't want to watch Love is Blind with me :)

Profilbild von Soroush Fadaeimanesh
Soroush Fadaeimaneshvor 4 Monaten

watching together is the killer pattern for realtime. way more useful than the agent watching alone and summarizing. what's the latency feeling like in practice?

Profilbild von Konrad Major
Konrad Majorvor 4 Monaten

You can paste the link into Gemini amd ask questions about the video.

Profilbild von Gregor
Gregorvor 4 Monaten

Multimodal context window doing the heavy lifting there.

Profilbild von Truly
Trulyvor 4 Monaten

Really nice!

Profilbild von Joe Hsu
Joe Hsuvor 4 Monaten

such good use of the realtime api! never thought of piping multiple sources

Profilbild von Daniel Bigham
Daniel Bighamvor 4 Monaten

This is great!

Profilbild von Don D'Cruz
Don D'Cruzvor 4 Monaten

The cost is gonna be crazy bro

Profilbild von TriciaAmazingyear
TriciaAmazingyearvor 4 Monaten

In the olden days your Mom or Dad and/or GrandMom did that 💋‼️😉

Profilbild von Victor
Victorvor 4 Monaten

passive watching is officially dead, every video is now a conversation

Profilbild von Albiona Hoti
Albiona Hotivor 4 Monaten

Wow!!

Profilbild von yun long
yun longvor 4 Monaten

have the same idea!

Profilbild von Tessa Archer
Tessa Archervor 4 Monaten

Audio stream differentiation without false triggers separates toys from tools. What latency are you hitting?

Profilbild von AI_lv0
AI_lv0vor 4 Monaten

what is the cost of running this lol on Lex pod this could cost you a lot of money?!!!

Profilbild von zan
zanvor 4 Monaten

cool

Profilbild von MAX ONBOARDER ⭕
MAX ONBOARDER ⭕vor 4 Monaten

Ohh this is amazing And can come in handy

Profilbild von 0xmusashi
0xmusashivor 4 Monaten

really cool idea zara 🔥

Profilbild von boilthesea
boiltheseavor 4 Monaten

How strong is that differentiation I wonder, would it hold up if a video maker intentionally tried to inject prompts? We could really use not confusing context with commands in literally every ai product.

Profilbild von Abhishek Kumar
Abhishek Kumarvor 4 Monaten

Great I am also thinking about buying a gun to kill a mosquito. And this line is about neither gun not mosquito.

Profilbild von William.K
William.Kvor 4 Monaten

So cool! It just listen to the audio right? What about diagrams / visuals on the video?

Profilbild von Abhishek Kumar
Abhishek Kumarvor 4 Monaten

Why such an over-engineering solution? Even if you use gemini in chrome or perplexity comet, they both have this built in by default.

Profilbild von InternationalOptions
InternationalOptionsvor 4 Monaten

@gabrielchua You’ve gotta gothub this!

Profilbild von BlanPlan
BlanPlanvor 4 Monaten

The voice/audio diarization piece is the part that quietly took years to get right. Tested OpenAI realtime 2 yesterday on background podcasts and it stayed silent through 90 minutes of video soundtrack until I directly addressed it. That separation is the real moat. Most realtime systems still spike on transient music.

Profilbild von J A Z I I
J A Z I Ivor 4 Monaten

intresting how can i use it too>

Profilbild von Happy Monkey AI
Happy Monkey AIvor 4 Monaten

I've got a project on my to-do list that can extract key information and highlights from videos, though I was leaning into things like teams meetings over YouTube, using vision models and LLM, having something that can find and give me terminal commands used in a demo and so on

Ähnliche Videos