Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately. - Parallel box decoding - 10× faster than Qwen3-VL - Trained on 138M queries/785M boxes - GUI, OCR, and document layout, dense detection - Open source Useful for computer-use agents and Physical AI,

131,447 görüntüleme • 4 gün önce •via X (Twitter)

14 Yorum

Md Ismail Šojal 🕷️ profil fotoğrafı
Md Ismail Šojal 🕷️4 gün önce

Worth a look if you work on agents, robotics, or document AI. -

falky profil fotoğrafı
falky4 gün önce

shit is 4 months old, fucking clickbait

Fajar M Reza profil fotoğrafı
Fajar M Reza4 gün önce

Does the 10× speedup hold on GUI workloads, or mainly dense grounding?

Samurai.AI profil fotoğrafı
Samurai.AI4 gün önce

@skalskip92

SHK profil fotoğrafı
SHK4 gün önce

@grok find this model to download

Dylan Colquhoun profil fotoğrafı
Dylan Colquhoun4 gün önce

Whats the optimal packing of 17 boxes in a square?

Derek Chia profil fotoğrafı
Derek Chia4 gün önce

Parallel box decoding sounds like the real trick. How does it hold up on dense GUI screens versus Qwen3-VL?

CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO profil fotoğrafı
CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO4 gün önce

Locate Anything?

Pies Avalon profil fotoğrafı
Pies Avalon4 gün önce

The final form of this will be an api service like jev.

Jim Wallace profil fotoğrafı
Jim Wallace4 gün önce

That scene in The Matrix Reloaded could use some DLSS5 treatment tho

Artzy profil fotoğrafı
Artzy4 gün önce

The throughput on that box decoding is wild

安叫兽|Bird🕊️ 🔶 BNB profil fotoğrafı
安叫兽|Bird🕊️ 🔶 BNB4 gün önce

3B 还能跑这么快,GUI 场景应该挺香,等实测了

Antonio Rocha profil fotoğrafı
Antonio Rocha4 gün önce

É mais rapido que o YOLO 26?

Ramesh Devasi profil fotoğrafı
Ramesh Devasi4 gün önce

why i am not able to get simple 4 corner flat planer tracker which can run in web

Benzer Videolar