Video yükleniyor...
Video Yüklenemedi
If you want a vision encoder for dexterous manipulation, what should be the most important part to model? 🤔 Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control. But collecting... show more
26,068 görüntüleme • 3 ay önce •via X (Twitter)
5 Yorum

Great achievement! another step towards making pre-trained models generalize from egocentric videos. And a good example of how useful it is to have anthropomorphic hand for inference on the robot!

How easy is it to use? This easy:

This work would not have been possible without an incredible roster of co-authors and collaborators. First and foremost, a massive thank you to the leads who spearheaded this research and turned a massive two-year vision into reality: 🔸Yuvan Sharma, 🔸Dantong Niu @Dantong_Niu, and 🔸Anirudh Pai @apai253. Building robust robotics takes a village. A huge shoutout to the dedicated team who spent countless hours collecting data and pushing through rigorous evaluations. Flawless execution from: 🔸Zekai Wang, 🔸Zhuoyang Liu @liu_zhuoyang13, 🔸Baifeng Shi @baifeng_shi, 🔸Stefano Saravalle, 🔸Boning Shao, 🔸Ruijie Zheng, 🔸Jing Wang, 🔸Konstantinos Kallidromitis, 🔸Yusuke Kato, and 🔸Fabio Galasso! Finally, special thanks for the invaluable mentorship, support, and help shaping the narrative: 🔸Yuke Zhu @yukez, 🔸Danfei Xu @danfei_xu, 🔸Linxi Fan @DrJimFan, 🔸Trevor Darrell @trevordarrell, and 🔸Jitendra Malik @JitendraMalikCV. A truly monumental team effort from BAIR @berkeley_ai and NVIDIA @nvidia! 🚀

This is awesome. Looking forward to building with it

Action-aware vision feels like a key piece for robotics. Recognizing objects is one thing, but understanding what can be done with them is where embodied models need to go. How will these representations change as more robot data becomes available?
