正在加载视频...

视频加载失败

Multimodality and streaming is hard. I've been building something that allows you connect streaming devices like screen capture, microphone, camera, and text easily to craft generative streaming pipelines. It works well with Gemini 2.0 models. Happy to open source if anyone wants it.

17,851 次观看 • 1 年前 •via X (Twitter)

10 条评论

Jaana Dogan ヤナ ドガン 的头像
Jaana Dogan ヤナ ドガン1 年前

It allows building pipelines with a little bit of configuration. It's super easy to quickly see what the models are capable of given a mixture of different modalities and context including custom components that can augment the context.

khaled (another one) 的头像
khaled (another one)1 年前

I’d like to use that

Sina Nejati 的头像
Sina Nejati1 年前

please do!

Samuel Navarro 的头像
Samuel Navarro1 年前

Really interested in something like this

Fakey McFakerson 的头像
Fakey McFakerson1 年前

👀

Kaushalya 的头像
Kaushalya1 年前

eggcellent! I'd like to try it.

Justin Collery 的头像
Justin Collery1 年前

Yes please!🙏

S4mpl3r 的头像
S4mpl3r1 年前

Very cool! I built a similar thing in python a couple months ago before all the realtime stuff came around. It's already open-source here:

Mike D 的头像
Mike D1 年前

+1 👆

Kesav 的头像
Kesav1 年前

Yes, please.

相关视频