Video yükleniyor...
Video Yüklenemedi
We present VLM-3R: a Vision-Language Model capable of 3D spatial reasoning from monocular video, grounding visual cues, geometry, and camera motion. ✅ No depth sensor ✅ No pre-built 3D maps ✅ End-to-end spatial + temporal reasoning 🔗 Code & benchmark: #VLM #3DVision #LLMs
14,895 görüntüleme • 1 yıl önce •via X (Twitter)
4 Yorum

Lennie Budgell ❇️1 yıl önce
This is some real great stuff I have been looking forward to seeing come into existence in such accuracy and types of usage. Finally. Thanks yall excited to get to playing around with the codr

PowerBeatsVR3 yıl önce
VR fitness app PowerBeatsVR is NOW LIVE on the official Meta Quest store! Get fit in VR without any expensive subscription:

Wenbo Hu1 yıl önce
Great work! I have a general question about why CUT3R is preferred over VGGT for spatial encoder?

Zhiwen(Aaron) Fan1 yıl önce
Great question. We’re aiming to equip VLMs with metric-scale geometric sensing.
