Loading video...
Video Failed to Load
DiffusionGemma-as-Jev (aka djev) running near-real-time vision detection from a mobile phone using its native vision tower. Please don't fall down the stairs!
104,402 views • 17 days ago •via X (Twitter)
15 Comments

You can even try it yourself! This is the shipping docker container with a packaged server and examples:

Each image is captured from the webcam, tokenized by the model, and the structured query gets answered by DiffusionGemma in about ~300-400ms. This is the stock DiffusionGemma model, no fine tuning, running practically stock vLLM with only a few patches.

It didn't do very well identifying the owner of the house (the furry one). Pet????? I'm sure th cat has something to say about that ;-)

it could be useful for blind

We just shipped open weights jev:

We should chat - would be great to share a single server implementation and make sure we align on multimodal Jev (mmJev?)! Feel free to reach out (DM or email)

dm me!

Was setting up a pipeline with YOLO26 + Jev but going to run this and compare.

I’m going to put it in to my robot dog

keep going, Matt!

Please give it perception of depth and try this again I'll wait

I have been following your posts since yesterday. Amazing

Very cool but to big for my 12gb vram Ç_Ç

Do you want T-800s? Because that's how you get T-800s.

I really need to know how much vram jev uses



