Loading video...

Video Failed to Load

Go Home

Excited to introduce TWIST2, our next-generation humanoid data collection system. TWIST2 is portable (use anywhere, no MoCap), scalable (100+ demos in 15 mins), and holistic (unlock major whole-body human skills). Fully open-sourced:

101,142 views • 10 months ago •via X (Twitter)

33 Comments

Yanjie Ze's profile picture
Yanjie Ze10 months ago

2/n Why to build TWIST2? I still remember 1 year ago, painfully teleoperating a full-size humanoid for simple pick & place to collect a few demonstrations. Even getting 10 episodes was tough, and we had to lock the lower body to keep the robot stable. That is a well-known bottleneck of teleop: "Not scalable". This problem gets much harder in humanoids: the state and action space are much larger than tabletop arms, and dynamics are more complex. video from iDP3 (

Yanjie Ze's profile picture
Yanjie Ze10 months ago

3/n TWIST2 system overview We finally have such a teleop system that has all the features we need for general-purpose robot data collection: - Scalable - Portable - Holistic

Yanjie Ze's profile picture
Yanjie Ze10 months ago

4/n Portable You can just put on a VR headset (and 2 motion trackers) and start teleoperating anywhere. Rapid setup, no motion capture required. From walk in with the VR device to finish setup, it takes only 1 min.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

5/n Holistic We find that egocentric active vision is essential to achieve long-horizon complex tasks. Therefore, we build a 2-DoF low-cost neck, which can be easily attached/detached to Unitree G1 without need to remove its original head.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

6/n Holistic This is what it looks like in VR when doing teleoperation. The robot view is floating in the center of the VR view.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

7/n Teleop skills Our system achieves human-like whole-body skills in a unified manner. Case 1: folding 3 towels consecutively, a very long-horizon and dexterous whole-body manipulation task.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

8/n Teleop skills Case 2: walk around like humans

Yanjie Ze's profile picture
Yanjie Ze10 months ago

9/n Teleop skills Case 3: Pick an object from the ground

Yanjie Ze's profile picture
Yanjie Ze10 months ago

10/n Scalable data collection TWIST2 enables data collection at an unprecedented scale: ~50 mobile manipulation demos in 15 minutes

Yanjie Ze's profile picture
Yanjie Ze10 months ago

11/n Scalable data collection TWIST2 enables data collection at an unprecedented scale: ~100 manipulation demos in 15 minutes

Yanjie Ze's profile picture
Yanjie Ze10 months ago

12/n Visuomotor policy learning With TWIST2 data, we further build a hierarchical visuomotor policy learning framework, that controls the full body of a humanoid with imitation learning.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

13/n Visuomotor policy learning That means, our visuomotor policy can not only pick & place objects,

Yanjie Ze's profile picture
Yanjie Ze10 months ago

14/n Visuomotor policy learning ...but can also use feet to kick a T-shaped box to the target region, i.e., Humanoid Kick-T.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

15/n Visuomotor policy learning Our visuomotor policy predicts future whole-body joint positions. Across *Space & Time*.

Yanjie Ze's profile picture
Yanjie Ze10 months ago

16/n Humanoid data We also believe humanoid datasets should be universally shareable. So we built a platform to host and visualize humanoid data, and welcome contributions from the community:

Yanjie Ze's profile picture
Yanjie Ze10 months ago

n/n That's all. Hope you enjoy our work. This work was done while interning at Amazon FAR, with incredible collaborators @sihengzhao @KenWangWeizhuo, and guidance from amazing advisors @akanazawa @rocky_duan @pabbeel @GuanyaShi @jiajunwu_cs Karen. Website: arXiv: YouTube:

Cheng Chi's profile picture
Cheng Chi10 months ago

Glad to see humanoid teleop getting to useable state for manipulation! Congrats Yanjie

William Mason (feat Bear)'s profile picture
William Mason (feat Bear)10 months ago

@Scobleizer @zeyanjie, would love to chat about collaborating (@Contact_CI)

James's profile picture
James10 months ago

Wow this is nice! I didn't know Pico had motion trackers like that, those look very interesting! How is the tracking and SLAM compared to Quest 3?

RYC's profile picture
RYC10 months ago

I am curious why you choose ZED2 rather than RealSense D435i as the head camera?

Emerson Segura's profile picture
Emerson Segura10 months ago

@chichengcc Nice work!

Neil Nie's profile picture
Neil Nie10 months ago

This is so cool!! Congrats @ZeYanjie!

Durable's profile picture
Durable10 months ago

super cool!! can it do my laundry for me?? (^_~)

Rxct Ftvy's profile picture
Rxct Ftvy9 months ago

So futuristic, good way to train, good work.

Tian Fang's profile picture
Tian Fang10 months ago

@ZeYanjie I remember that you mentioned the human motions in videos can also be retarget to robots instantly in the RoboPaper talk, please remind me which papers are you referring to? Thank you,

Wang Fu's profile picture
Wang Fu10 months ago

牛逼!

Miguel's profile picture
Miguel9 months ago

The paper mentions a Discord server, is it still active? The link on the paper website is no longer working

luiz's profile picture
luiz9 months ago

Hello! I don’t have a Pico 4, so I tried deploying TWIST2 with the provided motion lib. Something still feels quite weird, the walking isn’t fluid (the video below is from the 001 example, which walks forward, turns, and walks back). Any hint about a misconfig I might have made?

MAX ONBOARDER ⭕'s profile picture
MAX ONBOARDER ⭕10 months ago

@Scobleizer This is soo good

JJ Walker's profile picture
JJ Walker10 months ago

TWIST2 is like the Swiss Army knife for robotics-portable, scalable, and open-source. Seriously game-changing stuff! Which whole-body skill do you want to see demoed next?

Lorenzo Pieri's profile picture
Lorenzo Pieri10 months ago

The neck looks great, thanks for sharing with the community!

Jakie PLA's profile picture
Jakie PLA10 months ago

TWIST2's open-source approach is brilliant! Portable, scalable AND holistic? This will supercharge humanoid tech development. Can't wait to try the 2-DoF neck for active perception!

Tobe Duru's profile picture
Tobe Duru10 months ago

game changer for data collection

Related Videos

In just one week, Binh and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

23,236 views • 7 months ago

NEWS: Humanoid robotics company Figure has released Helix 02, what they claim in their most capable humanoid model yet. "A single neural system that controls the full body directly from pixels, enabling dexterous, long horizon autonomy across an entire room: • Autonomous, long‑horizon loco-manipulation: Helix 02 unloads and reloads a dishwasher across a full-sized kitchen - a four-minute, end-to-end autonomous task that integrates walking, manipulation, and balance with no resets and no human intervention. We believe this is the longest horizon, most complex task completed autonomously by a humanoid robot to date. • All sensors in. All actuators out: Helix 02 connects every onboard sensor - vision, touch, and proprioception - directly to every actuator through a single unified visuomotor neural network. • Human-like whole body control from human data: All results are enabled by System 0, a learned whole‑body controller trained on over 1,000 hours of human motion data and sim‑to‑real reinforcement learning. System 0 replaces 109,504 lines of hand‑engineered C++ with a single neural prior for stable, natural motion. • New classes of dexterity: With Figure 03’s embedded tactile sensing and palm cameras, Helix 02 performs manipulation that was previously out of reach: extracting individual pills, dispensing precise syringe volumes, and singulating small, irregular objects from clutter despite self‑occlusion. Helix 02 is trained on over 1,000 hours of human motion data and integrates vision, touch, and proprioception."

Sawyer Merritt

624,910 views • 7 months ago