Loading video...

Video Failed to Load

Go Home

NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately. - Parallel box decoding - 10× faster than Qwen3-VL - Trained on 138M queries/785M boxes - GUI, OCR, and document layout, dense detection - Open source Useful for computer-use agents and Physical AI,

131,447 views • 4 days ago •via X (Twitter)

14 Comments

Md Ismail Šojal 🕷️'s profile picture
Md Ismail Šojal 🕷️4 days ago

Worth a look if you work on agents, robotics, or document AI. -

falky's profile picture
falky4 days ago

shit is 4 months old, fucking clickbait

Fajar M Reza's profile picture
Fajar M Reza4 days ago

Does the 10× speedup hold on GUI workloads, or mainly dense grounding?

Samurai.AI's profile picture
Samurai.AI4 days ago

@skalskip92

SHK's profile picture
SHK4 days ago

@grok find this model to download

Dylan Colquhoun's profile picture
Dylan Colquhoun4 days ago

Whats the optimal packing of 17 boxes in a square?

Derek Chia's profile picture
Derek Chia4 days ago

Parallel box decoding sounds like the real trick. How does it hold up on dense GUI screens versus Qwen3-VL?

CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO's profile picture
CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO4 days ago

Locate Anything?

Pies Avalon's profile picture
Pies Avalon4 days ago

The final form of this will be an api service like jev.

Jim Wallace's profile picture
Jim Wallace4 days ago

That scene in The Matrix Reloaded could use some DLSS5 treatment tho

Artzy's profile picture
Artzy4 days ago

The throughput on that box decoding is wild

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB4 days ago

3B 还能跑这么快,GUI 场景应该挺香,等实测了

Antonio Rocha's profile picture
Antonio Rocha4 days ago

É mais rapido que o YOLO 26?

Ramesh Devasi's profile picture
Ramesh Devasi4 days ago

why i am not able to get simple 4 corner flat planer tracker which can run in web

Related Videos