Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

all three at a time, down to less than 1 sec latency - expression - object detection - finger count

19,486 Aufrufe • vor 13 Tagen •via X (Twitter)

9 Kommentare

Profilbild von Yohei
Yoheivor 13 Tagen

got object detection down to below 0.35 sec latency locally

Profilbild von Yohei
Yoheivor 14 Tagen

live true/false on object detection is at 0.4 seconds latency locally on an M5 macbook pro

Profilbild von Yohei
Yoheivor 14 Tagen

just object detection is much faster with .5 sec latency much snappier

Profilbild von Yohei
Yoheivor 14 Tagen

same thing, this time running locally. inference is slower, but no network latency (above was via replicate) inference: 1.1 sec seems inference time goes up with number of possible answers, and here's that's about 14 across three questions

Profilbild von Yohei
Yoheivor 14 Tagen

simultaneous live detection of: 😡 emotion 📕 object 🖐️ finger count single call for 3 detections, 0.2 sec inference each

Profilbild von Yohei
Yoheivor 13 Tagen

all three at a time sub 0.5 second latency! - expression - object detection - finger count switched to qwen 8B via glance i think direct on MLX

Profilbild von Mo ShaRaf (sharaf.eth)
Mo ShaRaf (sharaf.eth)vor 13 Tagen

Which model is this?

Profilbild von Yohei
Yoheivor 13 Tagen

qwen3-vl-4b via glance (no generation)

Profilbild von ethereagle · building
ethereagle · buildingvor 13 Tagen

sub 1s for expression + objects + fingers. did you pack all three into one call, or are they still separate and the model just got faster?

Ähnliche Videos

Check out our #ECCV2026 paper "Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention", where we make linear attention sparse in space, recurrent in time, and parallel in training, enabling the first purely-linear-attention-based neural network for asynchronous object detection with #EventCameras, outperforming the previous best asynchronous method with 20x less computation with truly event-by-event inference on CPU! Code released! Paper: Code: Video: Event cameras promise extremely low-latency vision, but to fully exploit them, the neural network must be low-latency too. We introduce #SpatiallySparseLinearAttention (#SSLA) for asynchronous object detection directly from raw events. Linear attention is particularly appealing for event cameras: it can be trained efficiently in parallel on long event sequences, while at inference it operates recurrently, updating its prediction every time a new event arrives. The problem is that conventional linear attention updates its entire state for every event. For object detection, where fine spatial resolution matters, this quickly becomes expensive. Our key idea is simple: an event only carries information about a small spatial region, so why update the entire spatial state? SSLA updates only the relevant parts of the state, enabling fine-grained spatial representations while keeping per-event computation low. We achieve: - >20× lower per-event computation than the strongest prior asynchronous baseline - State-of-the-art accuracy among asynchronous object detection methods - Truly event-by-event inference on CPU, designed to preserve the latency advantage of event cameras Come to our poster on Friday September 11, 2026 from 4-6pm at ExHall #389 Reference: Haiqing Hao, Zhipeng Sui, Rong Zou, Zijia Dai, Nikola Zubić, Davide Scaramuzza, Wenhui Wang Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention ECCV, 2026 Prophesee SynSense University of Zurich UZH Science European Research Council (ERC) UZHai UZH IfI Tesla BYD #EventCameras #ComputerVision #Robotics #DeepLearning #NeuromorphicVision #AI

Davide Scaramuzza

52,799 Aufrufe • vor 1 Monat