Video yükleniyor...
Video Yüklenemedi
NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately. - Parallel box decoding - 10× faster than Qwen3-VL - Trained on 138M queries/785M boxes - GUI, OCR, and document layout, dense detection - Open source Useful for computer-use agents and Physical AI,
131,447 görüntüleme • 4 gün önce •via X (Twitter)
14 Yorum

Worth a look if you work on agents, robotics, or document AI. -

shit is 4 months old, fucking clickbait

Does the 10× speedup hold on GUI workloads, or mainly dense grounding?

@skalskip92

@grok find this model to download

Whats the optimal packing of 17 boxes in a square?

Parallel box decoding sounds like the real trick. How does it hold up on dense GUI screens versus Qwen3-VL?

Locate Anything?

The final form of this will be an api service like jev.

That scene in The Matrix Reloaded could use some DLSS5 treatment tho

The throughput on that box decoding is wild

3B 还能跑这么快,GUI 场景应该挺香,等实测了

É mais rapido que o YOLO 26?

why i am not able to get simple 4 corner flat planer tracker which can run in web
