Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately. - Parallel box decoding - 10× faster than Qwen3-VL - Trained on 138M queries/785M boxes - GUI, OCR, and document layout, dense detection - Open source Useful for computer-use agents and Physical AI,

131,447 Aufrufe • vor 4 Tagen •via X (Twitter)

14 Kommentare

Profilbild von Md Ismail Šojal 🕷️
Md Ismail Šojal 🕷️vor 4 Tagen

Worth a look if you work on agents, robotics, or document AI. -

Profilbild von falky
falkyvor 4 Tagen

shit is 4 months old, fucking clickbait

Profilbild von Fajar M Reza
Fajar M Rezavor 4 Tagen

Does the 10× speedup hold on GUI workloads, or mainly dense grounding?

Profilbild von Samurai.AI
Samurai.AIvor 4 Tagen

@skalskip92

Profilbild von SHK
SHKvor 4 Tagen

@grok find this model to download

Profilbild von Dylan Colquhoun
Dylan Colquhounvor 4 Tagen

Whats the optimal packing of 17 boxes in a square?

Profilbild von Derek Chia
Derek Chiavor 4 Tagen

Parallel box decoding sounds like the real trick. How does it hold up on dense GUI screens versus Qwen3-VL?

Profilbild von CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEO
CEO del Socialismo de Mercado🌹🕊️ 市场社会主义CEOvor 4 Tagen

Locate Anything?

Profilbild von Pies Avalon
Pies Avalonvor 4 Tagen

The final form of this will be an api service like jev.

Profilbild von Jim Wallace
Jim Wallacevor 4 Tagen

That scene in The Matrix Reloaded could use some DLSS5 treatment tho

Profilbild von Artzy
Artzyvor 4 Tagen

The throughput on that box decoding is wild

Profilbild von 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNBvor 4 Tagen

3B 还能跑这么快,GUI 场景应该挺香,等实测了

Profilbild von Antonio Rocha
Antonio Rochavor 4 Tagen

É mais rapido que o YOLO 26?

Profilbild von Ramesh Devasi
Ramesh Devasivor 4 Tagen

why i am not able to get simple 4 corner flat planer tracker which can run in web

Ähnliche Videos