
Roei Herzig
@roeiherzig • 1,890 subscribers
Postdoc @berkeley_ai. Researcher @IBMResearch. PhD @TelAvivUni. Working on Multimodal Robot Foundation Models and Structured Physical AI.
Videos

If you want a vision encoder for dexterous manipulation, what should be the most important part to model? 🤔 Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control. But collecting annotated robotic trajectories at scale is SUPER expensive and largely unrealistic. So, how do we bridge this gap? We introduce CAIP (Contrastive Action-Image Pre-training) ⬇️ 🔸 Action-centric upstream: we align visual observations with action chunks through a contrastive objective. 🔸 Human video as a proxy: we represent 3D human hand poses analogously to robotic end-effector actions, tapping into a massive source of human demonstrations. 🔸 Massive scale: pre-trained on over 32,000 hours of manipulation video, driving both sample efficiency and robust generalization. 🔸 Hardware proven: achieves a 76% average success rate on a real-world Dexmate Vega bimanual Dexmate manipulator with dual 22-DoF Sharpa Wave hands Sharpa . 🔸 State-of-the-art: significantly outperforms strong baselines like DINOv2, SigLIP, MVP, and Qwen3.5 ViT across complex tasks, even under unexpected lighting changes and visual distractors. 🌐 Project: 📝 Blog: 📄 Paper: 💻 Model:
Roei Herzig22,201 Aufrufe • vor 1 Monat

What happens when vision🤝 robotics meet? Happy to share our new work on Pretraining Robotic Foundational Models!🔥 ARM4R is an Autoregressive Robotic Model that leverages low-level 4D Representations learned from human video data to yield a better robotic model. Berkeley AI Research😊
Roei Herzig62,805 Aufrufe • vor 1 Jahr

Children learn to manipulate the world by playing with toys — can robots do the same? 🧸🤖 We show that robots trained on 250 "toys" made of 4 shape primitives (🔵,🔶,🧱,💍) can generalize grasping to real objects. Jitendra MALIK trevordarrell Shankar Sastry Berkeley AI Research😊
Roei Herzig27,488 Aufrufe • vor 9 Monaten

Can instruction tuning be used in robotics?🚨Excited to release LLARVA, our new Vision-Action Instruction-tuned LMM for robotics! Key aspects: • Perform vision-action instruction tuning • Align vision and action modalities • Pretrained on robotic instruction data Berkeley AI Research
Roei Herzig10,517 Aufrufe • vor 2 Jahren
Keine weiteren Inhalte verfügbar