正在加载视频...

视频加载失败

DiffusionGemma-as-Jev (aka djev) running near-real-time vision detection from a mobile phone using its native vision tower. Please don't fall down the stairs!

104,402 次观看 • 17 天前 •via X (Twitter)

15 条评论

Matt Mastracci 的头像
Matt Mastracci17 天前

You can even try it yourself! This is the shipping docker container with a packaged server and examples:

Matt Mastracci 的头像
Matt Mastracci17 天前

Each image is captured from the webcam, tokenized by the model, and the structured query gets answered by DiffusionGemma in about ~300-400ms. This is the stock DiffusionGemma model, no fine tuning, running practically stock vLLM with only a few patches.

Mike Petch 的头像
Mike Petch17 天前

It didn't do very well identifying the owner of the house (the furry one). Pet????? I'm sure th cat has something to say about that ;-)

Inch 的头像
Inch17 天前

it could be useful for blind

Richelle🧬 的头像
Richelle🧬17 天前

We just shipped open weights jev:

Matt Mastracci 的头像
Matt Mastracci17 天前

We should chat - would be great to share a single server implementation and make sure we align on multimodal Jev (mmJev?)! Feel free to reach out (DM or email)

Richelle🧬 的头像
Richelle🧬17 天前

dm me!

JB 的头像
JB17 天前

Was setting up a pipeline with YOLO26 + Jev but going to run this and compare.

Chalkers 的头像
Chalkers17 天前

I’m going to put it in to my robot dog

johnny, INFINITE FUN ENTHUSIAST 的头像
johnny, INFINITE FUN ENTHUSIAST17 天前

keep going, Matt!

unsever 的头像
unsever17 天前

Please give it perception of depth and try this again I'll wait

Latent 的头像
Latent17 天前

I have been following your posts since yesterday. Amazing

AD3 Digital Marketing 的头像
AD3 Digital Marketing17 天前

Very cool but to big for my 12gb vram Ç_Ç

Simplicity 的头像
Simplicity17 天前

Do you want T-800s? Because that's how you get T-800s.

ego 的头像
ego17 天前

I really need to know how much vram jev uses

相关视频