正在加载视频...

视频加载失败

Humanoid motion tracking performance is greatly determined by retargeting quality! Introducing 𝗢𝗺𝗻𝗶𝗥𝗲𝘁𝗮𝗿𝗴𝗲𝘁🎯, generating high-quality interaction-preserving data from human motions for learning complex humanoid skills with 𝗺𝗶𝗻𝗶𝗺𝗮𝗹 RL: - 5 rewards, - 4 DR terms, - Proprio. ONLY, - NO history/curriculum. Ready for agile, human-like 🤖? (Best with 🎧) 🔗...

824,889 次观看 • 11 个月前 •via X (Twitter)

35 条评论

Zhen Wu 的头像
Zhen Wu11 个月前

Existing retargeting often produces artifacts like foot-skating and penetration ❌. To compensate, RL policies rely on complex ad-hoc reward terms, forcing a trade-off between accurate motion tracking and correcting errors like slipping or bad contacts. OmniRetarget fixes this at the source! ✅ Using an "interaction mesh," it generates clean, physically feasible trajectories by explicitly preserving the spatial and contact relationships between the agent, terrain, and objects. ✨ 2/9

Zhen Wu 的头像
Zhen Wu11 个月前

The result of this high-quality data? We can train diverse skills like box carrying 📦, slope crawling 🐾, and platform climbing 🧗 with a radically simplified RL process! All policies use just 5 reward terms, achieving successful zero-shot sim-to-real transfer! 🎯➡️🦾 3/9

Zhen Wu 的头像
Zhen Wu11 个月前

What about scalability? OmniRetarget transforms a SINGLE human demo into diverse motion clips. We can systematically vary terrain height, object size, and initial poses. Best of all, these augmented skills transfer directly from sim to our real-world hardware! 🤖➡️🦾 4/9

Zhen Wu 的头像
Zhen Wu11 个月前

And it's not just for a specific robot! Our framework is highly general and adapts to different robot embodiments, including the @UnitreeRobotics H1 and the @boosterobotics T1. We can retarget complex object-carrying and platform-climbing skills across these different robots with minimal changes. 5/9

Zhen Wu 的头像
Zhen Wu11 个月前

But how much better is our data? 🤔 Compared to widely-used baselines, our motions show far fewer physical artifacts—virtually zero foot-skating and penetration—while better preserving contact. This allows us to use an open-sourced RL framework (BeyondMimic) without hyperparameters tuning, while baselines fail to achieve high success rates in this setting. 6/9

Zhen Wu 的头像
Zhen Wu11 个月前

Our grand finale: A complex, long-horizon dynamic sequence, all driven by a proprioceptive-only policy (no vision/LIDAR)! In this task, the robot carries a chair to a platform, uses it as a step to climb up, then leaps off and performs a parkour-style roll to absorb the landing. This pushes the boundaries of agile, human-like loco-manipulation! 7/9

Zhen Wu 的头像
Zhen Wu11 个月前

Standing on the shoulders of giants! Our work builds on amazing research in the community💡. We use the "interaction mesh" 🕸️ [1], [2] to preserve spatial relationships and leverage the minimal RL formulation from works like BeyondMimic [3]. Our long-horizon sequence is a nod to the incredible Boston Dynamics Atlas demos 🤖 [4]! [1] E. S. L. Ho, T. Komura, and C.-L. Tai, “Spatial relationship preserving character motion adaptation,” ACM Transactions on Graphics, 2010. [2] S. Nakaoka and T. Komura, “Interaction mesh based motion adaptation for biped humanoid robots,” in Humanoids, 2012. [3] Q. Liao, T. E. Truong, X. Huang, G. Tevet, K. Sreenath, and C. K. Liu, “Beyondmimic: From motion tracking to versatile humanoid control via guided diffusion,” arXiv e-prints, pp. arXiv–2508, 2025. [4] Boston Dynamics, “Atlas Gets a Grip,” YouTube, available: https:// 8/9

Zhen Wu 的头像
Zhen Wu11 个月前

We are open-sourcing over 4 hours of high-quality, retargeted trajectories! Website: ArXiv: Datasets: Huge shout out to the amazing team: @lujieyang98, @x_h_ucb, @akanazawa, @pabbeel, @carlo_sferrazza, @ckarenliu, @rocky_duan, and @GuanyaShi from Amazon FAR, MIT, UC Berkeley, Stanford, and CMU! This work was done during an internship at Amazon FAR (Frontier AI & Robotics). 9/9

PrismaX 的头像
PrismaX11 个月前

climbed like a real human. Soon, we will see robot parkour competitions.

Brent Yi 的头像
Brent Yi11 个月前

Beautiful results!!! And cliffhanger 😂

Zhen Wu 的头像
Zhen Wu11 个月前

Thanks Brent! Yeah there are more exciting parkour-style motions on the way 🚗🚗🚗😝

Milan Kovac 的头像
Milan Kovac11 个月前

Very cool work

petr tyurin 的头像
petr tyurin11 个月前

I don’t know you personally yet, but you superstar better than 1000 Kardashians and Ronaldo 🚀🤖

Carlos DP 🤖🇺🇸 的头像
Carlos DP 🤖🇺🇸11 个月前

Incredible work! Really solid result

Laksh 的头像
Laksh11 个月前

@Scobleizer This looks like a huge step forward for more natural humanoid motion 👏

Xiatao Sun 的头像
Xiatao Sun11 个月前

Impressive work! Fixing retargeting artifacts at the source rather than with complex reward engineering is the right approach. The long-horizon parkour sequence is stunning!

Jon Conley (e/acc) 的头像
Jon Conley (e/acc)11 个月前

This is truly impressive to see how generalizeable this is and also simplifies the skill transfer process to potentially hundreds of humanoid robot vendors. Won’t be surprised to see lots of robotics companies building upon this work in the future.

Hongsuk Benjamin Choi 的头像
Hongsuk Benjamin Choi11 个月前

this is what the community needs:) and these videos are really impressive!

Rathor 的头像
Rathor11 个月前

Time for #HumanoidGeneration

Watching learner 的头像
Watching learner11 个月前

It is really good to see motion tracking improved like this. OmniRetarget proves how much the right starting point matters. That is also what AIOZ Movie Review Challenge does. It gives creators a fun and simple way to share movies and connect.

Guojian Wang 的头像
Guojian Wang11 个月前

@grok What's the meaning of retargeting?

Mike Domainer 的头像
Mike Domainer11 个月前

should fit

Dr Diogenes™ 🔭 🧬🇺🇸 的头像
Dr Diogenes™ 🔭 🧬🇺🇸11 个月前

That’s some really nice work! Congratulations!

Jude Onyenze 的头像
Jude Onyenze11 个月前

@TairanHe99 If you had to choose which is more efficient, learning from Third-Person Human or Learning from Motion Capture

Alfred Cueva 的头像
Alfred Cueva11 个月前

Interesting work! Specially since it doesn’t need to undergo curriculum training. Could OmniRetarget be made adaptive to the downstream RL task or policy uncertainty, dynamically refining trajectories?

Louis Le Lay 的头像
Louis Le Lay11 个月前

Wow interesting stuff and you say it's generalizable to other robots? Def interested 😁

Chung Min Kim 的头像
Chung Min Kim11 个月前

wow!!! 🥹🫶 amazing job this is so cool!!

176 feet per second 的头像
176 feet per second11 个月前

Pelvic and hip mobility is bloody amazing

VibeCodeTeddy 的头像
VibeCodeTeddy11 个月前

Impressive work on solving those artifact issues! Streamlined RL with fewer reward terms is a game changer. What's next for scalability challenges?

schakleton 的头像
schakleton11 个月前

Very impressive

GHOSTROSIN 的头像
GHOSTROSIN11 个月前

@Tesla_Optimus

Maxin 🇬🇭 的头像
Maxin 🇬🇭11 个月前

@Threadreaderapp unroll

Dirty Indy 🟥🟧🟨 的头像
Dirty Indy 🟥🟧🟨2 个月前

Any real world applications such as construction or working a garbage sorting facility, parkour stuff is cool but I am seeking enlightenment elsewhere in reducing human toil

Big Legend (❖,❖) 的头像
Big Legend (❖,❖)11 个月前

That's great.

Yichao Zhong 的头像
Yichao Zhong11 个月前

Congrats! When will you open-source the motion retargeting code🥹

相关视频

NEWS: Humanoid robotics company Figure has released Helix 02, what they claim in their most capable humanoid model yet. "A single neural system that controls the full body directly from pixels, enabling dexterous, long horizon autonomy across an entire room: • Autonomous, long‑horizon loco-manipulation: Helix 02 unloads and reloads a dishwasher across a full-sized kitchen - a four-minute, end-to-end autonomous task that integrates walking, manipulation, and balance with no resets and no human intervention. We believe this is the longest horizon, most complex task completed autonomously by a humanoid robot to date. • All sensors in. All actuators out: Helix 02 connects every onboard sensor - vision, touch, and proprioception - directly to every actuator through a single unified visuomotor neural network. • Human-like whole body control from human data: All results are enabled by System 0, a learned whole‑body controller trained on over 1,000 hours of human motion data and sim‑to‑real reinforcement learning. System 0 replaces 109,504 lines of hand‑engineered C++ with a single neural prior for stable, natural motion. • New classes of dexterity: With Figure 03’s embedded tactile sensing and palm cameras, Helix 02 performs manipulation that was previously out of reach: extracting individual pills, dispensing precise syringe volumes, and singulating small, irregular objects from clutter despite self‑occlusion. Helix 02 is trained on over 1,000 hours of human motion data and integrates vision, touch, and proprioception."

Sawyer Merritt

624,910 次观看 • 7 个月前

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. 🤖 Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. It’s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. ➤ Radical Openness: X-Humanoid is open-sourcing the full stack—robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. ➤ High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. ➤ True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. ➤ Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. It’s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHub🤖

49,440 次观看 • 7 个月前

I don’t know if we live in a Matrix, but I know for sure that robots will spend most of their lives in simulation. Let machines train machines. I’m excited to introduce DexMimicGen, a massive-scale synthetic data generator that enables a humanoid robot to learn complex skills from only a handful of human demonstrations. Yes, as few as 5! DexMimicGen addresses the biggest pain point in robotics: where do we get data? Unlike with LLMs, where vast amounts of texts are readily available, you cannot simply download motor control signals from the internet. So researchers teleoperate the robots to collect motion data via XR headsets. They have to repeat the same skill over and over and over again, because neural nets are data hungry. This is a very slow and uncomfortable process. At NVIDIA, we believe the majority of high-quality tokens for robot foundation models will come from simulation. What DexMimicGen does is to trade GPU compute time for human time. It takes one motion trajectory from human, and multiplies into 1000s of new trajectories. A robot brain trained on this augmented dataset will generalize far better in the real world. Think of DexMimicGen as a learning signal amplifier. It maps a small dataset to a large (de facto infinite) dataset, using physics simulation in the loop. In this way, we free humans from babysitting the bots all day. The future of robot data is generative. The future of the entire robot learning pipeline will also be generative. 🧵

Jim Fan

165,246 次观看 • 1 年前

New framework: Kick down your robot, it will get back up every time 🥋 Chinese startup RoboParty is a Beijing startup founded April 2025 by Huang Yi, originally shipping ROBOTO ORIGIN, the world's first full-stack open-source bipedal humanoid. They released UFO: Unsupervised Reinforcement Learning Framework for Humanoid Control. DEFINITIONS -> what differs is where the learning signal comes from: - SUPERVISED: humans supply the right answers (labels), the model imitates them. - UNSUPERVISED: no answer key, the model finds structure in raw data on its own. - REINFORCEMENT LEARNING: no answer key either, the model tries things and a reward scores each attempt. → UNSUPERVISED RL: trial and error where the agent invents its own rewards, instead of engineers hand-writing one per task. REPRESENTATION LEARNING: compress raw states into a useful internal map. TEMPORAL DISTANCE: distance on that map is "how many steps from A to B." CONTRASTIVE: trained by pulling together what's close in time, pushing apart what isn't. -> CONTRASTIVE TEMPORAL-DISTANCE REPRESENTATION LEARNING: the model builds an internal map of body states where distance means how many steps it takes to get from one to another. It is trained by contrast: states that occur close together in a movement get pulled together in the map, randomly paired states get pushed apart. UFO is an open-source training framework that teaches humanoid robots skills, like getting up, walking, goal-reaching, teleoperation, without reference motions -> no motion-capture or human-video demonstrations to imitate. Its core is TeCH, a contrastive temporal-distance representation-learning algorithm: the robot explores, builds pseudo-goals by temporal rolling, and learns goal-conditioned policies from a single unified progress reward. One framework trains five different robots (Unitree G1/H1, RoboParty RP0/RP1, AgiBot X2) with automatic config conversion in ~2–3 hours per robot! The real novelty here "no demonstrations at all". No data-collection arms race,the dominant humanoid-locomotion recipe is tracking: imitate mocap/retargeted-human reference trajectories. The robot self-generates goals from its own exploration and learns from a progress reward, needing zero reference motion data. Everybody else is fighting over data acquisition, while this team just teleports out of the race entirely (inb4 "competition is for losers 💀 ). This strategy reminds me of the DeepSeek playbook applied to robots: open-source the whole stack to become the global default and commoditize everyone else. RoboParty is giving away hardware and now control software (UFO) to be the Android of humanoids. Yet another reason for the US to ban Chinese open models perhaps 🥶 ? What I also really like about this approach is the cross-embodiment infrastructure, one framework trains Unitree G1/H1, RoboParty RP0/RP1, and AgiBot X2 with automatic configuration conversion. Just like Physical Intelligence, RoboParty seems to place itself as a neutral hardware agnostic middle man. Also woth mentioning: their ability ot perform stable skill injection, e.g. adding a cartwheel without forgetting how to walk. A common failure of RL humanoid policies is that teaching a new agile skill destabilizes the existing ones (catastrophic forgetting). UFO claims you can inject rare motions (cartwheel) without collapsing learned behavior. If it holds, that's a significant incremental/continual skill-learning! But again, I have to underline it: no arXiv, no external validation, no success-rate numbers. -> robotics badely needs an independent unbiased evaluator imho. Still, look at that cool demo: robot is getting kicked and pushed around (serious disturbance) during teleoperation (controlled the person at the back wearing the VR headset), and still managed to always get back up. This is some serious demonstration of stability and robustness!

Léo

36,112 次观看 • 1 个月前

Physics-based Motion Retargeting from Sparse Inputs paper page: Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers, providing very limited sensor data of the user's pose. Another challenge is that an avatar might have a different skeleton structure than a human and the mapping between them is unclear. In this work we address both of these challenges. We introduce a method to retarget motions in real-time from sparse human sensor data to characters of various morphologies. Our method uses reinforcement learning to train a policy to control characters in a physics simulator. We only require human motion capture data for training, without relying on artist-generated animations for each avatar. This allows us to use large motion capture datasets to train general policies that can track unseen users from real and sparse data in real-time. We demonstrate the feasibility of our approach on three characters with different skeleton structure: a dinosaur, a mouse-like creature and a human. We show that the avatar poses often match the user surprisingly well, despite having no sensor information of the lower body available. We discuss and ablate the important components in our framework, specifically the kinematic retargeting step, the imitation, contact and action reward as well as our asymmetric actor-critic observations. We further explore the robustness of our method in a variety of settings including unbalancing, dancing and sports motions.

AK

106,527 次观看 • 3 年前