Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Excited to introduce TWIST2, our next-generation humanoid data collection system. TWIST2 is portable (use anywhere, no MoCap), scalable (100+ demos in 15 mins), and holistic (unlock major whole-body human skills). Fully open-sourced:

101,142 Aufrufe • vor 10 Monaten •via X (Twitter)

33 Kommentare

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

2/n Why to build TWIST2? I still remember 1 year ago, painfully teleoperating a full-size humanoid for simple pick & place to collect a few demonstrations. Even getting 10 episodes was tough, and we had to lock the lower body to keep the robot stable. That is a well-known bottleneck of teleop: "Not scalable". This problem gets much harder in humanoids: the state and action space are much larger than tabletop arms, and dynamics are more complex. video from iDP3 (

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

3/n TWIST2 system overview We finally have such a teleop system that has all the features we need for general-purpose robot data collection: - Scalable - Portable - Holistic

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

4/n Portable You can just put on a VR headset (and 2 motion trackers) and start teleoperating anywhere. Rapid setup, no motion capture required. From walk in with the VR device to finish setup, it takes only 1 min.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

5/n Holistic We find that egocentric active vision is essential to achieve long-horizon complex tasks. Therefore, we build a 2-DoF low-cost neck, which can be easily attached/detached to Unitree G1 without need to remove its original head.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

6/n Holistic This is what it looks like in VR when doing teleoperation. The robot view is floating in the center of the VR view.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

7/n Teleop skills Our system achieves human-like whole-body skills in a unified manner. Case 1: folding 3 towels consecutively, a very long-horizon and dexterous whole-body manipulation task.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

8/n Teleop skills Case 2: walk around like humans

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

9/n Teleop skills Case 3: Pick an object from the ground

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

10/n Scalable data collection TWIST2 enables data collection at an unprecedented scale: ~50 mobile manipulation demos in 15 minutes

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

11/n Scalable data collection TWIST2 enables data collection at an unprecedented scale: ~100 manipulation demos in 15 minutes

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

12/n Visuomotor policy learning With TWIST2 data, we further build a hierarchical visuomotor policy learning framework, that controls the full body of a humanoid with imitation learning.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

13/n Visuomotor policy learning That means, our visuomotor policy can not only pick & place objects,

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

14/n Visuomotor policy learning ...but can also use feet to kick a T-shaped box to the target region, i.e., Humanoid Kick-T.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

15/n Visuomotor policy learning Our visuomotor policy predicts future whole-body joint positions. Across *Space & Time*.

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

16/n Humanoid data We also believe humanoid datasets should be universally shareable. So we built a platform to host and visualize humanoid data, and welcome contributions from the community:

Profilbild von Yanjie Ze
Yanjie Zevor 10 Monaten

n/n That's all. Hope you enjoy our work. This work was done while interning at Amazon FAR, with incredible collaborators @sihengzhao @KenWangWeizhuo, and guidance from amazing advisors @akanazawa @rocky_duan @pabbeel @GuanyaShi @jiajunwu_cs Karen. Website: arXiv: YouTube:

Profilbild von Cheng Chi
Cheng Chivor 10 Monaten

Glad to see humanoid teleop getting to useable state for manipulation! Congrats Yanjie

Profilbild von William Mason (feat Bear)
William Mason (feat Bear)vor 10 Monaten

@Scobleizer @zeyanjie, would love to chat about collaborating (@Contact_CI)

Profilbild von James
Jamesvor 10 Monaten

Wow this is nice! I didn't know Pico had motion trackers like that, those look very interesting! How is the tracking and SLAM compared to Quest 3?

Profilbild von RYC
RYCvor 10 Monaten

I am curious why you choose ZED2 rather than RealSense D435i as the head camera?

Profilbild von Emerson Segura
Emerson Seguravor 10 Monaten

@chichengcc Nice work!

Profilbild von Neil Nie
Neil Nievor 10 Monaten

This is so cool!! Congrats @ZeYanjie!

Profilbild von Durable
Durablevor 10 Monaten

super cool!! can it do my laundry for me?? (^_~)

Profilbild von Rxct Ftvy
Rxct Ftvyvor 9 Monaten

So futuristic, good way to train, good work.

Profilbild von Tian Fang
Tian Fangvor 10 Monaten

@ZeYanjie I remember that you mentioned the human motions in videos can also be retarget to robots instantly in the RoboPaper talk, please remind me which papers are you referring to? Thank you,

Profilbild von Wang Fu
Wang Fuvor 10 Monaten

牛逼!

Profilbild von Miguel
Miguelvor 9 Monaten

The paper mentions a Discord server, is it still active? The link on the paper website is no longer working

Profilbild von luiz
luizvor 9 Monaten

Hello! I don’t have a Pico 4, so I tried deploying TWIST2 with the provided motion lib. Something still feels quite weird, the walking isn’t fluid (the video below is from the 001 example, which walks forward, turns, and walks back). Any hint about a misconfig I might have made?

Profilbild von MAX ONBOARDER ⭕
MAX ONBOARDER ⭕vor 10 Monaten

@Scobleizer This is soo good

Profilbild von JJ Walker
JJ Walkervor 10 Monaten

TWIST2 is like the Swiss Army knife for robotics-portable, scalable, and open-source. Seriously game-changing stuff! Which whole-body skill do you want to see demoed next?

Profilbild von Lorenzo Pieri
Lorenzo Pierivor 10 Monaten

The neck looks great, thanks for sharing with the community!

Profilbild von Jakie PLA
Jakie PLAvor 10 Monaten

TWIST2's open-source approach is brilliant! Portable, scalable AND holistic? This will supercharge humanoid tech development. Can't wait to try the 2-DoF neck for active perception!

Profilbild von Tobe Duru
Tobe Duruvor 10 Monaten

game changer for data collection

Ähnliche Videos

In just one week, Binh and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

23,236 Aufrufe • vor 7 Monaten

NEWS: Humanoid robotics company Figure has released Helix 02, what they claim in their most capable humanoid model yet. "A single neural system that controls the full body directly from pixels, enabling dexterous, long horizon autonomy across an entire room: • Autonomous, long‑horizon loco-manipulation: Helix 02 unloads and reloads a dishwasher across a full-sized kitchen - a four-minute, end-to-end autonomous task that integrates walking, manipulation, and balance with no resets and no human intervention. We believe this is the longest horizon, most complex task completed autonomously by a humanoid robot to date. • All sensors in. All actuators out: Helix 02 connects every onboard sensor - vision, touch, and proprioception - directly to every actuator through a single unified visuomotor neural network. • Human-like whole body control from human data: All results are enabled by System 0, a learned whole‑body controller trained on over 1,000 hours of human motion data and sim‑to‑real reinforcement learning. System 0 replaces 109,504 lines of hand‑engineered C++ with a single neural prior for stable, natural motion. • New classes of dexterity: With Figure 03’s embedded tactile sensing and palm cameras, Helix 02 performs manipulation that was previously out of reach: extracting individual pills, dispensing precise syringe volumes, and singulating small, irregular objects from clutter despite self‑occlusion. Helix 02 is trained on over 1,000 hours of human motion data and integrates vision, touch, and proprioception."

Sawyer Merritt

624,910 Aufrufe • vor 7 Monaten