Загрузка видео...

Не удалось загрузить видео

На главную

Built a "YouTube realtime copilot" browser extension using OpenAI's realtime 2 API: The agent watches the video alongside you, and can answer any question you have about what was just said via realtime voice chat. The crazy part to me is: It can differentiate the YouTube's audio stream and...

45,631 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 34

Фото профиля 程北FIG
程北FIG4 месяцев назад

This feels like the right direction for AI video. Not just “summarize this video”, but watch with me, answer in context, and help me understand the part I just missed. The UX shift is from search-after-watching to assistance-while-watching.

Фото профиля Progress-means-a-lot
Progress-means-a-lot4 месяцев назад

auto-pause when a user speaks so Copilot can engage seamlessly. Requiring manual intervention significantly degrades the experience; automation is key to overall quality.

Фото профиля Doxy
Doxy4 месяцев назад

wait, does it actually parse the audio stream or just grab captions? curious if that separation trick holds up on chaotic tech talks

Фото профиля cesaer
cesaer4 месяцев назад

on Github yet? need it!

Фото профиля Bakhty | 📣 Sarafan
Bakhty | 📣 Sarafan4 месяцев назад

Kudos for a great addition to my list of best tools to learn languages 💪🏼🔥🚀

Фото профиля Devon Chaine
Devon Chaine4 месяцев назад

This is not better than the YouTube so feature

Фото профиля Jingzi Zhao
Jingzi Zhao4 месяцев назад

How much time do you spend vibe coding a product?

Фото профиля ORacle
ORacle4 месяцев назад

太get這個邊看邊問的插件了!太強了,我幾天前還在盧頻扣字幕,現在已經能實時陪看了😂 我特別好奇這個插件從idea到上線要多久呢?對於普通vibecoder來說最難復製的是音頻流區分嗎?求指點

Фото профиля 陈洁
陈洁4 месяцев назад

真的是非常需要了哈哈!

Фото профиля Mengxue
Mengxue4 месяцев назад

I would love to see this way of human-AI interaction in more personal, casual settings -- like watching a reality TV show together and chat about it like friends, which will definitely scratch an itch when my bf doesn't want to watch Love is Blind with me :)

Фото профиля Soroush Fadaeimanesh
Soroush Fadaeimanesh4 месяцев назад

watching together is the killer pattern for realtime. way more useful than the agent watching alone and summarizing. what's the latency feeling like in practice?

Фото профиля Konrad Major
Konrad Major4 месяцев назад

You can paste the link into Gemini amd ask questions about the video.

Фото профиля Gregor
Gregor4 месяцев назад

Multimodal context window doing the heavy lifting there.

Фото профиля Truly
Truly4 месяцев назад

Really nice!

Фото профиля Joe Hsu
Joe Hsu4 месяцев назад

such good use of the realtime api! never thought of piping multiple sources

Фото профиля Daniel Bigham
Daniel Bigham4 месяцев назад

This is great!

Фото профиля Don D'Cruz
Don D'Cruz4 месяцев назад

The cost is gonna be crazy bro

Фото профиля TriciaAmazingyear
TriciaAmazingyear4 месяцев назад

In the olden days your Mom or Dad and/or GrandMom did that 💋‼️😉

Фото профиля Victor
Victor4 месяцев назад

passive watching is officially dead, every video is now a conversation

Фото профиля Albiona Hoti
Albiona Hoti4 месяцев назад

Wow!!

Фото профиля yun long
yun long4 месяцев назад

have the same idea!

Фото профиля Tessa Archer
Tessa Archer4 месяцев назад

Audio stream differentiation without false triggers separates toys from tools. What latency are you hitting?

Фото профиля AI_lv0
AI_lv04 месяцев назад

what is the cost of running this lol on Lex pod this could cost you a lot of money?!!!

Фото профиля zan
zan4 месяцев назад

cool

Фото профиля MAX ONBOARDER ⭕
MAX ONBOARDER ⭕4 месяцев назад

Ohh this is amazing And can come in handy

Фото профиля 0xmusashi
0xmusashi4 месяцев назад

really cool idea zara 🔥

Фото профиля boilthesea
boilthesea4 месяцев назад

How strong is that differentiation I wonder, would it hold up if a video maker intentionally tried to inject prompts? We could really use not confusing context with commands in literally every ai product.

Фото профиля Abhishek Kumar
Abhishek Kumar4 месяцев назад

Great I am also thinking about buying a gun to kill a mosquito. And this line is about neither gun not mosquito.

Фото профиля William.K
William.K4 месяцев назад

So cool! It just listen to the audio right? What about diagrams / visuals on the video?

Фото профиля Abhishek Kumar
Abhishek Kumar4 месяцев назад

Why such an over-engineering solution? Even if you use gemini in chrome or perplexity comet, they both have this built in by default.

Фото профиля InternationalOptions
InternationalOptions4 месяцев назад

@gabrielchua You’ve gotta gothub this!

Фото профиля BlanPlan
BlanPlan4 месяцев назад

The voice/audio diarization piece is the part that quietly took years to get right. Tested OpenAI realtime 2 yesterday on background podcasts and it stayed silent through 90 minutes of video soundtrack until I directly addressed it. That separation is the real moat. Most realtime systems still spike on transient music.

Фото профиля J A Z I I
J A Z I I4 месяцев назад

intresting how can i use it too>

Фото профиля Happy Monkey AI
Happy Monkey AI4 месяцев назад

I've got a project on my to-do list that can extract key information and highlights from videos, though I was leaning into things like teams meetings over YouTube, using vision models and LLM, having something that can find and give me terminal commands used in a demo and so on

Похожие видео