Загрузка видео...
Не удалось загрузить видео
We present VLM-3R: a Vision-Language Model capable of 3D spatial reasoning from monocular video, grounding visual cues, geometry, and camera motion. ✅ No depth sensor ✅ No pre-built 3D maps ✅ End-to-end spatial + temporal reasoning 🔗 Code & benchmark: #VLM #3DVision #LLMs
14,895 просмотров • 1 год назад •via X (Twitter)
Комментарии: 4

Lennie Budgell ❇️1 год назад
This is some real great stuff I have been looking forward to seeing come into existence in such accuracy and types of usage. Finally. Thanks yall excited to get to playing around with the codr

PowerBeatsVR3 лет назад
VR fitness app PowerBeatsVR is NOW LIVE on the official Meta Quest store! Get fit in VR without any expensive subscription:

Wenbo Hu1 год назад
Great work! I have a general question about why CUT3R is preferred over VGGT for spatial encoder?

Zhiwen(Aaron) Fan1 год назад
Great question. We’re aiming to equip VLMs with metric-scale geometric sensing.
