Loading video...

Video Failed to Load

Go Home

Humanoid motion tracking performance is greatly determined by retargeting quality! Introducing ๐—ข๐—บ๐—ป๐—ถ๐—ฅ๐—ฒ๐˜๐—ฎ๐—ฟ๐—ด๐—ฒ๐˜๐ŸŽฏ, generating high-quality interaction-preserving data from human motions for learning complex humanoid skills with ๐—บ๐—ถ๐—ป๐—ถ๐—บ๐—ฎ๐—น RL: - 5 rewards, - 4 DR terms, - Proprio. ONLY, - NO history/curriculum. Ready for agile, human-like ๐Ÿค–? (Best with ๐ŸŽง) ๐Ÿ”—...

824,889 views โ€ข 11 months ago โ€ขvia X (Twitter)

35 Comments

Zhen Wu's profile picture
Zhen Wu11 months ago

Existing retargeting often produces artifacts like foot-skating and penetration โŒ. To compensate, RL policies rely on complex ad-hoc reward terms, forcing a trade-off between accurate motion tracking and correcting errors like slipping or bad contacts. OmniRetarget fixes this at the source! โœ… Using an "interaction mesh," it generates clean, physically feasible trajectories by explicitly preserving the spatial and contact relationships between the agent, terrain, and objects. โœจ 2/9

Zhen Wu's profile picture
Zhen Wu11 months ago

The result of this high-quality data? We can train diverse skills like box carrying ๐Ÿ“ฆ, slope crawling ๐Ÿพ, and platform climbing ๐Ÿง— with a radically simplified RL process! All policies use just 5 reward terms, achieving successful zero-shot sim-to-real transfer! ๐ŸŽฏโžก๏ธ๐Ÿฆพ 3/9

Zhen Wu's profile picture
Zhen Wu11 months ago

What about scalability? OmniRetarget transforms a SINGLE human demo into diverse motion clips. We can systematically vary terrain height, object size, and initial poses. Best of all, these augmented skills transfer directly from sim to our real-world hardware! ๐Ÿค–โžก๏ธ๐Ÿฆพ 4/9

Zhen Wu's profile picture
Zhen Wu11 months ago

And it's not just for a specific robot! Our framework is highly general and adapts to different robot embodiments, including the @UnitreeRobotics H1 and the @boosterobotics T1. We can retarget complex object-carrying and platform-climbing skills across these different robots with minimal changes. 5/9

Zhen Wu's profile picture
Zhen Wu11 months ago

But how much better is our data? ๐Ÿค” Compared to widely-used baselines, our motions show far fewer physical artifactsโ€”virtually zero foot-skating and penetrationโ€”while better preserving contact. This allows us to use an open-sourced RL framework (BeyondMimic) without hyperparameters tuning, while baselines fail to achieve high success rates in this setting. 6/9

Zhen Wu's profile picture
Zhen Wu11 months ago

Our grand finale: A complex, long-horizon dynamic sequence, all driven by a proprioceptive-only policy (no vision/LIDAR)! In this task, the robot carries a chair to a platform, uses it as a step to climb up, then leaps off and performs a parkour-style roll to absorb the landing. This pushes the boundaries of agile, human-like loco-manipulation! 7/9

Zhen Wu's profile picture
Zhen Wu11 months ago

Standing on the shoulders of giants! Our work builds on amazing research in the community๐Ÿ’ก. We use the "interaction mesh" ๐Ÿ•ธ๏ธ [1], [2] to preserve spatial relationships and leverage the minimal RL formulation from works like BeyondMimic [3]. Our long-horizon sequence is a nod to the incredible Boston Dynamics Atlas demos ๐Ÿค– [4]! [1] E. S. L. Ho, T. Komura, and C.-L. Tai, โ€œSpatial relationship preserving character motion adaptation,โ€ ACM Transactions on Graphics, 2010. [2] S. Nakaoka and T. Komura, โ€œInteraction mesh based motion adaptation for biped humanoid robots,โ€ in Humanoids, 2012. [3] Q. Liao, T. E. Truong, X. Huang, G. Tevet, K. Sreenath, and C. K. Liu, โ€œBeyondmimic: From motion tracking to versatile humanoid control via guided diffusion,โ€ arXiv e-prints, pp. arXivโ€“2508, 2025. [4] Boston Dynamics, โ€œAtlas Gets a Grip,โ€ YouTube, available: https:// 8/9

Zhen Wu's profile picture
Zhen Wu11 months ago

We are open-sourcing over 4 hours of high-quality, retargeted trajectories! Website: ArXiv: Datasets: Huge shout out to the amazing team: @lujieyang98, @x_h_ucb, @akanazawa, @pabbeel, @carlo_sferrazza, @ckarenliu, @rocky_duan, and @GuanyaShi from Amazon FAR, MIT, UC Berkeley, Stanford, and CMU! This work was done during an internship at Amazon FAR (Frontier AI & Robotics). 9/9

PrismaX's profile picture
PrismaX11 months ago

climbed like a real human. Soon, we will see robot parkour competitions.

Brent Yi's profile picture
Brent Yi11 months ago

Beautiful results!!! And cliffhanger ๐Ÿ˜‚

Zhen Wu's profile picture
Zhen Wu11 months ago

Thanks Brent! Yeah there are more exciting parkour-style motions on the way ๐Ÿš—๐Ÿš—๐Ÿš—๐Ÿ˜

Milan Kovac's profile picture
Milan Kovac11 months ago

Very cool work

petr tyurin's profile picture
petr tyurin11 months ago

I donโ€™t know you personally yet, but you superstar better than 1000 Kardashians and Ronaldo ๐Ÿš€๐Ÿค–

Carlos DP ๐Ÿค–๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Carlos DP ๐Ÿค–๐Ÿ‡บ๐Ÿ‡ธ11 months ago

Incredible work! Really solid result

Laksh's profile picture
Laksh11 months ago

@Scobleizer This looks like a huge step forward for more natural humanoid motion ๐Ÿ‘

Xiatao Sun's profile picture
Xiatao Sun11 months ago

Impressive work! Fixing retargeting artifacts at the source rather than with complex reward engineering is the right approach. The long-horizon parkour sequence is stunning!

Jon Conley (e/acc)'s profile picture
Jon Conley (e/acc)11 months ago

This is truly impressive to see how generalizeable this is and also simplifies the skill transfer process to potentially hundreds of humanoid robot vendors. Wonโ€™t be surprised to see lots of robotics companies building upon this work in the future.

Hongsuk Benjamin Choi's profile picture
Hongsuk Benjamin Choi11 months ago

this is what the community needs:) and these videos are really impressive!

Rathor's profile picture
Rathor11 months ago

Time for #HumanoidGeneration

Watching learner's profile picture
Watching learner11 months ago

It is really good to see motion tracking improved like this. OmniRetarget proves how much the right starting point matters. That is also what AIOZ Movie Review Challenge does. It gives creators a fun and simple way to share movies and connect.

Guojian Wang's profile picture
Guojian Wang11 months ago

@grok What's the meaning of retargeting?

Mike Domainer's profile picture
Mike Domainer11 months ago

should fit

Dr Diogenesโ„ข ๐Ÿ”ญ ๐Ÿงฌ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Dr Diogenesโ„ข ๐Ÿ”ญ ๐Ÿงฌ๐Ÿ‡บ๐Ÿ‡ธ11 months ago

Thatโ€™s some really nice work! Congratulations!

Jude Onyenze's profile picture
Jude Onyenze11 months ago

@TairanHe99 If you had to choose which is more efficient, learning from Third-Person Human or Learning from Motion Capture

Alfred Cueva's profile picture
Alfred Cueva11 months ago

Interesting work! Specially since it doesnโ€™t need to undergo curriculum training. Could OmniRetarget be made adaptive to the downstream RL task or policy uncertainty, dynamically refining trajectories?

Louis Le Lay's profile picture
Louis Le Lay11 months ago

Wow interesting stuff and you say it's generalizable to other robots? Def interested ๐Ÿ˜

Chung Min Kim's profile picture
Chung Min Kim11 months ago

wow!!! ๐Ÿฅน๐Ÿซถ amazing job this is so cool!!

176 feet per second's profile picture
176 feet per second11 months ago

Pelvic and hip mobility is bloody amazing

VibeCodeTeddy's profile picture
VibeCodeTeddy11 months ago

Impressive work on solving those artifact issues! Streamlined RL with fewer reward terms is a game changer. What's next for scalability challenges?

schakleton's profile picture
schakleton11 months ago

Very impressive

GHOSTROSIN's profile picture
GHOSTROSIN11 months ago

@Tesla_Optimus

Maxin ๐Ÿ‡ฌ๐Ÿ‡ญ's profile picture
Maxin ๐Ÿ‡ฌ๐Ÿ‡ญ11 months ago

@Threadreaderapp unroll

Dirty Indy ๐ŸŸฅ๐ŸŸง๐ŸŸจ's profile picture
Dirty Indy ๐ŸŸฅ๐ŸŸง๐ŸŸจ2 months ago

Any real world applications such as construction or working a garbage sorting facility, parkour stuff is cool but I am seeking enlightenment elsewhere in reducing human toil

Big Legend (โ–,โ–)'s profile picture
Big Legend (โ–,โ–)11 months ago

That's great.

Yichao Zhong's profile picture
Yichao Zhong11 months ago

Congrats! When will you open-source the motion retargeting code๐Ÿฅน

Related Videos

NEWS: Humanoid robotics company Figure has released Helix 02, what they claim in their most capable humanoid model yet. "A single neural system that controls the full body directly from pixels, enabling dexterous, long horizon autonomy across an entire room: โ€ข Autonomous, longโ€‘horizon loco-manipulation: Helix 02 unloads and reloads a dishwasher across a full-sized kitchen - a four-minute, end-to-end autonomous task that integrates walking, manipulation, and balance with no resets and no human intervention. We believe this is the longest horizon, most complex task completed autonomously by a humanoid robot to date. โ€ข All sensors in. All actuators out: Helix 02 connects every onboard sensor - vision, touch, and proprioception - directly to every actuator through a single unified visuomotor neural network. โ€ข Human-like whole body control from human data: All results are enabled by System 0, a learned wholeโ€‘body controller trained on over 1,000 hours of human motion data and simโ€‘toโ€‘real reinforcement learning. System 0 replaces 109,504 lines of handโ€‘engineered C++ with a single neural prior for stable, natural motion. โ€ข New classes of dexterity: With Figure 03โ€™s embedded tactile sensing and palm cameras, Helix 02 performs manipulation that was previously out of reach: extracting individual pills, dispensing precise syringe volumes, and singulating small, irregular objects from clutter despite selfโ€‘occlusion. Helix 02 is trained on over 1,000 hours of human motion data and integrates vision, touch, and proprioception."

Sawyer Merritt

624,910 views โ€ข 7 months ago

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. ๐Ÿค– Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. Itโ€™s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. โžค Radical Openness: X-Humanoid is open-sourcing the full stackโ€”robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. โžค High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. โžค True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. โžค Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. Itโ€™s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHub๐Ÿค–

49,440 views โ€ข 7 months ago

I donโ€™t know if we live in a Matrix, but I know for sure that robots will spend most of their lives in simulation. Let machines train machines. Iโ€™m excited to introduce DexMimicGen, a massive-scale synthetic data generator that enables a humanoid robot to learn complex skills from only a handful of human demonstrations. Yes, as few as 5! DexMimicGen addresses the biggest pain point in robotics: where do we get data? Unlike with LLMs, where vast amounts of texts are readily available, you cannot simply download motor control signals from the internet. So researchers teleoperate the robots to collect motion data via XR headsets. They have to repeat the same skill over and over and over again, because neural nets are data hungry. This is a very slow and uncomfortable process. At NVIDIA, we believe the majority of high-quality tokens for robot foundation models will come from simulation. What DexMimicGen does is to trade GPU compute time for human time. It takes one motion trajectory from human, and multiplies into 1000s of new trajectories. A robot brain trained on this augmented dataset will generalize far better in the real world. Think of DexMimicGen as a learning signal amplifier. It maps a small dataset to a large (de facto infinite) dataset, using physics simulation in the loop. In this way, we free humans from babysitting the bots all day. The future of robot data is generative. The future of the entire robot learning pipeline will also be generative. ๐Ÿงต

Jim Fan

165,246 views โ€ข 1 year ago

New framework: Kick down your robot, it will get back up every time ๐Ÿฅ‹ Chinese startup RoboParty is a Beijing startup founded April 2025 by Huang Yi, originally shipping ROBOTO ORIGIN, the world's first full-stack open-source bipedal humanoid. They released UFO: Unsupervised Reinforcement Learning Framework for Humanoid Control. DEFINITIONS -> what differs is where the learning signal comes from: - SUPERVISED: humans supply the right answers (labels), the model imitates them. - UNSUPERVISED: no answer key, the model finds structure in raw data on its own. - REINFORCEMENT LEARNING: no answer key either, the model tries things and a reward scores each attempt. โ†’ UNSUPERVISED RL: trial and error where the agent invents its own rewards, instead of engineers hand-writing one per task. REPRESENTATION LEARNING: compress raw states into a useful internal map. TEMPORAL DISTANCE: distance on that map is "how many steps from A to B." CONTRASTIVE: trained by pulling together what's close in time, pushing apart what isn't. -> CONTRASTIVE TEMPORAL-DISTANCE REPRESENTATION LEARNING: the model builds an internal map of body states where distance means how many steps it takes to get from one to another. It is trained by contrast: states that occur close together in a movement get pulled together in the map, randomly paired states get pushed apart. UFO is an open-source training framework that teaches humanoid robots skills, like getting up, walking, goal-reaching, teleoperation, without reference motions -> no motion-capture or human-video demonstrations to imitate. Its core is TeCH, a contrastive temporal-distance representation-learning algorithm: the robot explores, builds pseudo-goals by temporal rolling, and learns goal-conditioned policies from a single unified progress reward. One framework trains five different robots (Unitree G1/H1, RoboParty RP0/RP1, AgiBot X2) with automatic config conversion in ~2โ€“3 hours per robot! The real novelty here "no demonstrations at all". No data-collection arms race,the dominant humanoid-locomotion recipe is tracking: imitate mocap/retargeted-human reference trajectories. The robot self-generates goals from its own exploration and learns from a progress reward, needing zero reference motion data. Everybody else is fighting over data acquisition, while this team just teleports out of the race entirely (inb4 "competition is for losers ๐Ÿ’€ ). This strategy reminds me of the DeepSeek playbook applied to robots: open-source the whole stack to become the global default and commoditize everyone else. RoboParty is giving away hardware and now control software (UFO) to be the Android of humanoids. Yet another reason for the US to ban Chinese open models perhaps ๐Ÿฅถ ? What I also really like about this approach is the cross-embodiment infrastructure, one framework trains Unitree G1/H1, RoboParty RP0/RP1, and AgiBot X2 with automatic configuration conversion. Just like Physical Intelligence, RoboParty seems to place itself as a neutral hardware agnostic middle man. Also woth mentioning: their ability ot perform stable skill injection, e.g. adding a cartwheel without forgetting how to walk. A common failure of RL humanoid policies is that teaching a new agile skill destabilizes the existing ones (catastrophic forgetting). UFO claims you can inject rare motions (cartwheel) without collapsing learned behavior. If it holds, that's a significant incremental/continual skill-learning! But again, I have to underline it: no arXiv, no external validation, no success-rate numbers. -> robotics badely needs an independent unbiased evaluator imho. Still, look at that cool demo: robot is getting kicked and pushed around (serious disturbance) during teleoperation (controlled the person at the back wearing the VR headset), and still managed to always get back up. This is some serious demonstration of stability and robustness!

Lรฉo

36,112 views โ€ข 1 month ago

Physics-based Motion Retargeting from Sparse Inputs paper page: Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers, providing very limited sensor data of the user's pose. Another challenge is that an avatar might have a different skeleton structure than a human and the mapping between them is unclear. In this work we address both of these challenges. We introduce a method to retarget motions in real-time from sparse human sensor data to characters of various morphologies. Our method uses reinforcement learning to train a policy to control characters in a physics simulator. We only require human motion capture data for training, without relying on artist-generated animations for each avatar. This allows us to use large motion capture datasets to train general policies that can track unseen users from real and sparse data in real-time. We demonstrate the feasibility of our approach on three characters with different skeleton structure: a dinosaur, a mouse-like creature and a human. We show that the avatar poses often match the user surprisingly well, despite having no sensor information of the lower body available. We discuss and ablate the important components in our framework, specifically the kinematic retargeting step, the imitation, contact and action reward as well as our asymmetric actor-critic observations. We further explore the robustness of our method in a variety of settings including unbalancing, dancing and sports motions.

AK

106,527 views โ€ข 3 years ago