Video yükleniyor...
Video Yüklenemedi
glance-vlm speedlab is now open source! read: try: turn any webcam into multiple live AI detectors (emotion, count, object), running locally ⚡ Recorded live on an Apple M5, one-question loop: PyTorch/MPS FP16: ~210 ms p50 MLX 8-bit: ~160 ms p50 🧪 Controlled nine-question, fresh-frame benchmark: 358.5 → 259.6 ms... show more
41,369 görüntüleme • 11 gün önce •via X (Twitter)
27 Yorum

glance-vlm is a harness for measuring what happens when you chop off the generation from an open source vlm and use it as a classifier/decision model:

ICYMI that’s 161 ms latency on general vision true/false Qs at the end, running locally on webcam

compare glance 4B to laya-vision

Nice. One use case I see is to monitor elderly. My grandma is quite sick and it takes a lot of time, energy and attention to make sure she is good. Would be awesome to get an alert once the camera notices something weird in case we are not there. What do you think? Can be done?

I think worth playing around with! Once you know a good use case you might want to go specialized but this is an easy way to prototype something quickly

for the temporal gate, compare against the last scored frame, not the previous camera frame. otherwise slow drift can stay below the threshold forever while the answer gets stale.

hm thanks will take a look at this

Really neat . love it

thx :)

This is game changer for vision AI

🐐

This is good!

This is awesome! I need to fork it and turn it to a meeting companion that helps neurodivergent people with reading emotions and social queues.

this is so cool!

Thx :)

210ms locally is the bit that jumps out, that puts weird webcam workflows in reach without renting a gpu every time. curious how hot the m 5 gets after 20 minutes though

Agents got hands in 2023. Eyes took until 2026, at 260ms a glance. Counting people in a frame is arithmetic, a face is a guess. Which of your 21 failures surprised you most?

Great work Yohei!

The interesting bit isn’t the webcam demo—it’s the loop shape. Turning a VLM into a cheap local decision engine is a much more useful product primitive than another chat wrapper.

glance-vlmのspeedlab OSS、僕は触りたい。 webcamをローカル検出に、って相手の切り口は受け取る。 自分では未開。まずは

Local webcam AI detectors that stay on-device? That's the useful kind of open source. Bookmarking this.

Local VLM speedlab goes open source with Mac numbers that favor MLX over PyTorch MPS and a reproducible bench.

160ms p50 on MLX 8-bit is the number that matters - same loop on MPS carries the memory bandwidth tax. The webcam-as-sensor framing is right: local loops only work when inference stays on the unified memory side of the machine.

did the 84/84 hold once the frames got messier, or is that on the clean benchmark set

open sourcing the speedlab is the move. local multi-detectors under a couple hundred ms is the latency window game NPCs actually live in 👾

Local, low-latency detectors open up a different interaction model: the system can observe continuously without every glance becoming a cloud request. The UX challenge is making that visibility feel useful, not creepy.

Yohei - what broke first past nine-question batches?
