Loading video...
Video Failed to Load
If you want a vision encoder for dexterous manipulation, what should be the most important part to model? ๐ค Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control. But collecting... show more
22,201 views โข 1 month ago โขvia X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
