Loading video...

Video Failed to Load

Go Home

Led by Google DeepMind, we present ALOHA 2 ๐Ÿค™: An Enhanced Low-Cost Hardware for Bimanual Teleoperation. ALOHA 2 ๐Ÿค™ significantly improves the durability of the original ALOHA ๐Ÿ–๏ธ, enabling fleet-scale data collection on more complex tasks. As usual, everything is open-sourced!

144,078 views โ€ข 2 years ago โ€ขvia X (Twitter)

10 Comments

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

Before diving into the hardware, we also release a *proper* ALOHA sim model with SysID, thanks to @kevin_zakka @the_real_btaba @ayzwah. Even if you donโ€™t have the hardware, there is now a way to perform complex tasks with ALOHA in Mujoco!

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

We start by improving the grippers: to make them grasp better and more robust. We use a low-friction rail design that transmits 2x more force to the gripper tips. We also change the grip tape layout to improve grasping of small objects. Led by @SpencerGoodric6 and Thinh Nguyen

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

We use the same rail design on the leader side. To further improve ergonomics, we replace the original servo with a lower gear ratio one that is easier to backdrive. This results in a 10x reduction in friction that the operator needs to overcome when opening grippers!

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

Next, we improve the gravity compensation of the leader arm. With a constant-force retractor and a spring-pulley system, the arm can "float" in most places. It is also much more durable than the original rubberbands!

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

Last but not least: we simplify the frame surrounding the workcell while maintaining the rigidity of the camera mounting points. This opens up the space for both human-robot collaborators and props for the robot to interact with.

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

To learn more, please visit our website: Paper: Tutorial: Designs: Sim:

Tony Z. Zhao's profile picture
Tony Z. Zhao2 years ago

Thanks to the core ALOHA 2 Team: @RandomRobotics @chelseabfinn @peteflorence @SpencerGoodric6 Thinh Nguyen @JonathanTompson @ayzwah @tonyzzhao and those who helped with hardware, software, data, simulation, and user studies: Jorge Aldaco, Robert Baruch, Jeff Bingham, Sanky Chan, Kenneth Draper, @debidatta, Wayne Gramlich, Torr Hage, @AlexHerzog00, Jonathan Hoech, Ian Storz, @the_real_btaba, @leilatakayama, Ted Wahrburg, Sichun Xu, Sergey Yaroshenko, and @kevin_zakka

Masato Kobayashi @ใ‚‹ใฃใจ๐Ÿบ's profile picture
Masato Kobayashi @ใ‚‹ใฃใจ๐Ÿบ2 years ago

@GoogleDeepMind @tonyzzhao This is an incredibly fantastic achievement! I'm excited! Please look forward to our report on interesting research about imitation learning as well! This is part of that research.

Brett Adcock's profile picture
Brett Adcock2 years ago

@GoogleDeepMind good work Tony / team

Keerthana Gopalakrishnan's profile picture
Keerthana Gopalakrishnan2 years ago

@GoogleDeepMind lmfao @peteflorence and @andyzeng_ ๐Ÿคฃ

Related Videos

This is how ALOHA's "teleoperation" system works - a fancy word for "remote control". Training robots will be more and more like playing games in the physical world. A human operates a "joystick++" to perform tasks and collect data, or intervene if there's any safety concern. There's actually a learning curve to master the controller, much like practicing gaming skills. Teleoperation can be done in many different ways. ALOHA is an impressive custom-built system with very low cost. Here're a few alternatives: (1) Motion Capture (MoCap): apply the MoCap systems used for Hollywood movies to capture the fine-grained motions of hand joints. There would be no "embodiment gap" if the robot hand has 5 fingers. For instance, a demonstrator can wear a CyberGlove ( and manipulate the objects. CyberGlove will capture the motion signals & haptic feedback in real-time, which can be re-targeted onto the humanoid. (2) Wearing gloves & markers can be clumsy. An alternative way to do MoCap is through computer vision. DexPilot from NVIDIA enables marker-less and glove-free data collection. The human operator simply uses their bare hands to perform the tasks. 4 Intel RealSense depth cameras and 2 NVIDIA Titan XP GPUs (yeah, 2019 work) translate the pixels to precise motion signals for robot learning. (3) VR Headset: turn the training room into a VR game and "role play" the robot. This has the advantage of scalable remote data collection - annotators from around the world can contribute without coming onsite. VR demonstration technique appeared in research projects like the iGibson home robot simulator, an initiative that I participated in at Stanford: Behind-the-scene video by Litian Liang

Jim Fan

124,588 views โ€ข 2 years ago

I was really impressed by the UMI gripper (Cheng Chi et al.), but a key limitation is that **force-related data wasnโ€™t captured**: humans feel haptic feedback through the mechanical springs, but the robot couldnโ€™t leverage that info, limiting the dataโ€™s value for fine-grained manipulation tasks. Led by my amazing students Yolanda Zhu and Binghao Huang, we designed a **portable visuo-tactile gripper** by integrating our dense, flexible tactile arrays with the UMI gripper to enable large-scale in-the-wild data collection. ๐Ÿ”— We demonstrate **cross-modal representation learning** and **downstream policy learning** on tasks requiring in-hand state estimation (e.g., test tube reorientation) and fine-grained force sensing (e.g., pipette fluid transfer). Key takeaways: - Our flexible tactile arrays store the rich haptic information humans perceive as dense tactile signals. - Portability and robustness are key for in-the-wild data collection; our portable gripper is compact, lightweight, and durable. - Touch provides precise, robust measurements of in-hand object pose, invariant to lighting and viewpoint. - Cross-modal pretraining on large-scale in-the-wild data significantly improves policy robustness and sample efficiency (as shown many times before โ€” and verified again here!). Also check out our previous investigations of dense, flexible tactile grids for understanding human-robot-environment interactions: - Dense tactile glove (Nature โ€™19): - 3D-ViTac (CoRL โ€™24):

Yunzhu Li

13,188 views โ€ข 1 year ago

Excited to announce GR00T N1, the worldโ€™s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to โ€œhallucinateโ€ new synthetic data that features accurate physics in pixels. Using Jensenโ€™s words, โ€œsystematically infinite dataโ€! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that โ€œrendersโ€ smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Letโ€™s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: ๐Ÿงต

Jim Fan

466,442 views โ€ข 1 year ago