正在加载视频...

视频加载失败

NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately. - Parallel box decoding - 10× faster than Qwen3-VL - Trained on 138M queries/785M boxes - GUI, OCR, and document layout, dense detection - Open source Useful for computer-use agents and Physical AI,

131,447 次观看 • 4 天前 •via X (Twitter)

14 条评论

Md Ismail Šojal 🕷️ 的头像
Md Ismail Šojal 🕷️4 天前

Worth a look if you work on agents, robotics, or document AI. -

falky 的头像
falky4 天前

shit is 4 months old, fucking clickbait

Fajar M Reza 的头像
Fajar M Reza4 天前

Does the 10× speedup hold on GUI workloads, or mainly dense grounding?

Samurai.AI 的头像
Samurai.AI4 天前

@skalskip92

SHK 的头像
SHK4 天前

@grok find this model to download

Dylan Colquhoun 的头像
Dylan Colquhoun4 天前

Whats the optimal packing of 17 boxes in a square?

Derek Chia 的头像
Derek Chia4 天前

Parallel box decoding sounds like the real trick. How does it hold up on dense GUI screens versus Qwen3-VL?

CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO 的头像
CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO4 天前

Locate Anything?

Pies Avalon 的头像
Pies Avalon4 天前

The final form of this will be an api service like jev.

Jim Wallace 的头像
Jim Wallace4 天前

That scene in The Matrix Reloaded could use some DLSS5 treatment tho

Artzy 的头像
Artzy4 天前

The throughput on that box decoding is wild

安叫兽|Bird🕊️ 🔶 BNB 的头像
安叫兽|Bird🕊️ 🔶 BNB4 天前

3B 还能跑这么快,GUI 场景应该挺香,等实测了

Antonio Rocha 的头像
Antonio Rocha4 天前

É mais rapido que o YOLO 26?

Ramesh Devasi 的头像
Ramesh Devasi4 天前

why i am not able to get simple 4 corner flat planer tracker which can run in web

相关视频