Loading video...
Video Failed to Load
simultaneous live detection of: 😡 emotion 📕 object 🖐️ finger count single call for 3 detections, 0.2 sec inference each
15,279 views • 14 days ago •via X (Twitter)
31 Comments

same thing, this time running locally. inference is slower, but no network latency (above was via replicate) inference: 1.1 sec seems inference time goes up with number of possible answers, and here's that's about 14 across three questions

just object detection is much faster with .5 sec latency much snappier

live true/false on object detection is at 0.4 seconds latency locally on an M5 macbook pro

got object detection down to below 0.35 sec latency locally

all three at a time, down to less than 1 sec latency - expression - object detection - finger count

all three at a time sub 0.5 second latency! - expression - object detection - finger count switched to qwen 8B via glance i think direct on MLX

is this using Jev?

no but a similar approach on qwen3 vl 4b

ok, that's super cool, you should legit come presenting at our next AI Socratic in SF. I already shared this, so I feel like I'm spamming you... but maybe you want to present at the Day online:

i would but weekends are tough as i usually got three kids hanging on my arms

yea, tbh I kinda regret having planned this for the weekend… especially with 8 more events in the planning for next month🥲 I’ve got pulled in by the Jev excitement myself.

This feels almost like watching someone juggle :)

this was my third take of the demo 😅

very cool. is it just literally qwen3-VL-4B or have you done something clever around it?

skipped the generation and read logits, which i've measured can cut up about ~85% of compute with multiple questions per call

would detection always be faster than generation? wondering if this could work for proof of human use cases

basically yes. but if it's a specialized enough use case that you do often AND it's important, i would guess it's almost always better to do a specialized model

Messerschmidt reborn

the pirate flag for the boat 😭

it detected a pirate ship yo 🏴☠️

ive wanted to play with this kind of idea so badly! think about the pattern detections/associative patterns a mental health-forward system could accomplish. "ive noticed that your distress signals fire at 3x the average every time youre working on _____" "you're joy and happiness scale rises ten fold everytime we talk about ____ or work on ____, you seem a bit down, would you like to _____ and see if it makes you feel any better?"

fast enough for live detection and correct enough for live detection are two different bars, only one got shown

it's correct enough for some use cases, not for others

0.2s each for emotion, object, and finger count in one call. are the 3 heads sharing a backbone, or 3 sequential passes packed into one request?

Yeah, the interesting part isn’t the demo speed — it’s collapsing three perception calls into one loop. If that holds off the happy-path video, this is the kind of latency shift that makes real-time feel real.

i keep building real‑time pipelines like yours and then waste time hunting where to talk about them so i'm building that finds relevant threads and drops a ready reply!

Love it, but I know you weren’t confused by Moonwalking!

That’s a nice example of multimodal inference becoming a systems problem, not a model demo. One call doing three detections at 0.2 seconds is the kind of small latency win that makes an interaction feel alive.

0.2 seconds for three detections is the kind of number that changes the product, not just the benchmark. The real test is whether it stays boringly reliable outside the demo.

one call three detections is the flex

Three detectors in one call is the kind of boring integration win that actually ships. The flashy demo is the model; the useful part is the latency staying predictable.
