Loading video...
Video Failed to Load
NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately. - Parallel box decoding - 10× faster than Qwen3-VL - Trained on 138M queries/785M boxes - GUI, OCR, and document layout, dense detection - Open source Useful for computer-use agents and Physical AI,
131,447 views • 4 days ago •via X (Twitter)
14 Comments

Md Ismail Šojal 🕷️4 days ago
Worth a look if you work on agents, robotics, or document AI. -

falky4 days ago
shit is 4 months old, fucking clickbait

Fajar M Reza4 days ago
Does the 10× speedup hold on GUI workloads, or mainly dense grounding?

Samurai.AI4 days ago
@skalskip92

SHK4 days ago
@grok find this model to download

Dylan Colquhoun4 days ago
Whats the optimal packing of 17 boxes in a square?

Derek Chia4 days ago
Parallel box decoding sounds like the real trick. How does it hold up on dense GUI screens versus Qwen3-VL?

CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO4 days ago
Locate Anything?

Pies Avalon4 days ago
The final form of this will be an api service like jev.

Jim Wallace4 days ago
That scene in The Matrix Reloaded could use some DLSS5 treatment tho

Artzy4 days ago
The throughput on that box decoding is wild

安叫兽|Bird🕊️ 🔶 BNB4 days ago
3B 还能跑这么快,GUI 场景应该挺香,等实测了

Antonio Rocha4 days ago
É mais rapido que o YOLO 26?

Ramesh Devasi4 days ago
why i am not able to get simple 4 corner flat planer tracker which can run in web
