Загрузка видео...
Не удалось загрузить видео
If you want a vision encoder for dexterous manipulation, what should be the most important part to model? 🤔 Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control. But collecting... show more
22,201 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 0
Нет доступных комментариев
Здесь появятся комментарии из оригинального поста
