Video yükleniyor...
Video Yüklenemedi
How do you give a humanoid the general motion capability? Not just single motions, but all motion? Introducing SONIC, our new work on supersizing motion tracking for natural humanoid control. We argue that motion tracking is the scalable foundation task for humanoids. So we "supersized" it: 9k+ GPU hours... show more
62,751 görüntüleme • 10 ay önce •via X (Twitter)
15 Yorum

Now, let's dive into how we achieved this: First, we gather lots of high-quality humanoid motion data, all 100m frames, 700 hours of it. We trained a universal humanoid motion tracker on it using 128 GPUs for a couple of days. (Yes, this is PHC-to-real if you follow my work :) The resulting tracker is great. High coverage on the motion dataset. However, tracking alone is not enough. It's not useful if we just show that we can track any motion from a large dataset. We need to track out-of-distribution motion, useful motion. How do we make motion tracking the scalable foundation for humanoid control?

We begin by designing a kinematic planner that can, like a gaming character, control humanoids to move on command based on gamepad control. The planner itself is a foundation-model level innovation that supports many. All of the footage below was taken in one take, over 10 minutes, controlled by a gamepad.. All of the footage below was taken in one take over 10 mins controlled by a gamepad.

Now that we have the planner (we say that a lot within our team these days), you can do some neat things like decoupling the upper body (VR three points) and the lower body (kinematic planner) for natural movement teleoperation

Of course, you will be able to teleoperate with whole-body movement, all upper and lower body, with motion provided by a VR headset

Adding GENMO ( to the mix, we also support teleoperating with videos. Mapping from video-captured human motion to humanoids.

With GENMO ( we can generate music-to-motion and text-to-motion, which are mapped directly to humanoid behaviors.

If we can teleoperate, we can collect data for Vision-Language-Actions models (VLAs)! Here we collect 300 trajectories, fine-tune the newest Gr00t N-series model ( and achieve a 95% success rate on this loco-manipluation task during our initial trials.

All of the above behaviors are using the same, single control policy that runs onboard the G1's orin. All models and training code will be released.

The dream team: Co-first authors @_ye_yuan @TingwuWang @80gg_overmind | Core contributors @eric_srchen Fernando, @ziang_cao , @jiefengli_jeff , David, @BenQingwei , @dennis | Our team @dngxngxng3 , @cyrushogg, Lina, Edy, Eugene, @TairanHe99 , @HaoruXue , @_wenlixiao , Zi | our leaders @jankautz, @Dr_YanChang, @UmarIqb , @DrJimFan , @yukez more videos and details at:

100% Game Changer. Congratulations!!!

Great 👍

So, in the end will you be selling Unitree bots with Sonic Control or are you hoping someone else will?

Insane work to the community! Thanks for sharing

Well done & nice name!

haha
