Загрузка видео...

Не удалось загрузить видео

На главную

Humanoid motion tracking performance is greatly determined by retargeting quality! Introducing 𝗢𝗺𝗻𝗶𝗥𝗲𝘁𝗮𝗿𝗴𝗲𝘁🎯, generating high-quality interaction-preserving data from human motions for learning complex humanoid skills with 𝗺𝗶𝗻𝗶𝗺𝗮𝗹 RL: - 5 rewards, - 4 DR terms, - Proprio. ONLY, - NO history/curriculum. Ready for agile, human-like 🤖? (Best with 🎧) 🔗...

824,889 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 35

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

Existing retargeting often produces artifacts like foot-skating and penetration ❌. To compensate, RL policies rely on complex ad-hoc reward terms, forcing a trade-off between accurate motion tracking and correcting errors like slipping or bad contacts. OmniRetarget fixes this at the source! ✅ Using an "interaction mesh," it generates clean, physically feasible trajectories by explicitly preserving the spatial and contact relationships between the agent, terrain, and objects. ✨ 2/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

The result of this high-quality data? We can train diverse skills like box carrying 📦, slope crawling 🐾, and platform climbing 🧗 with a radically simplified RL process! All policies use just 5 reward terms, achieving successful zero-shot sim-to-real transfer! 🎯➡️🦾 3/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

What about scalability? OmniRetarget transforms a SINGLE human demo into diverse motion clips. We can systematically vary terrain height, object size, and initial poses. Best of all, these augmented skills transfer directly from sim to our real-world hardware! 🤖➡️🦾 4/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

And it's not just for a specific robot! Our framework is highly general and adapts to different robot embodiments, including the @UnitreeRobotics H1 and the @boosterobotics T1. We can retarget complex object-carrying and platform-climbing skills across these different robots with minimal changes. 5/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

But how much better is our data? 🤔 Compared to widely-used baselines, our motions show far fewer physical artifacts—virtually zero foot-skating and penetration—while better preserving contact. This allows us to use an open-sourced RL framework (BeyondMimic) without hyperparameters tuning, while baselines fail to achieve high success rates in this setting. 6/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

Our grand finale: A complex, long-horizon dynamic sequence, all driven by a proprioceptive-only policy (no vision/LIDAR)! In this task, the robot carries a chair to a platform, uses it as a step to climb up, then leaps off and performs a parkour-style roll to absorb the landing. This pushes the boundaries of agile, human-like loco-manipulation! 7/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

Standing on the shoulders of giants! Our work builds on amazing research in the community💡. We use the "interaction mesh" 🕸️ [1], [2] to preserve spatial relationships and leverage the minimal RL formulation from works like BeyondMimic [3]. Our long-horizon sequence is a nod to the incredible Boston Dynamics Atlas demos 🤖 [4]! [1] E. S. L. Ho, T. Komura, and C.-L. Tai, “Spatial relationship preserving character motion adaptation,” ACM Transactions on Graphics, 2010. [2] S. Nakaoka and T. Komura, “Interaction mesh based motion adaptation for biped humanoid robots,” in Humanoids, 2012. [3] Q. Liao, T. E. Truong, X. Huang, G. Tevet, K. Sreenath, and C. K. Liu, “Beyondmimic: From motion tracking to versatile humanoid control via guided diffusion,” arXiv e-prints, pp. arXiv–2508, 2025. [4] Boston Dynamics, “Atlas Gets a Grip,” YouTube, available: https:// 8/9

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

We are open-sourcing over 4 hours of high-quality, retargeted trajectories! Website: ArXiv: Datasets: Huge shout out to the amazing team: @lujieyang98, @x_h_ucb, @akanazawa, @pabbeel, @carlo_sferrazza, @ckarenliu, @rocky_duan, and @GuanyaShi from Amazon FAR, MIT, UC Berkeley, Stanford, and CMU! This work was done during an internship at Amazon FAR (Frontier AI & Robotics). 9/9

Фото профиля PrismaX
PrismaX11 месяцев назад

climbed like a real human. Soon, we will see robot parkour competitions.

Фото профиля Brent Yi
Brent Yi11 месяцев назад

Beautiful results!!! And cliffhanger 😂

Фото профиля Zhen Wu
Zhen Wu11 месяцев назад

Thanks Brent! Yeah there are more exciting parkour-style motions on the way 🚗🚗🚗😝

Фото профиля Milan Kovac
Milan Kovac11 месяцев назад

Very cool work

Фото профиля petr tyurin
petr tyurin11 месяцев назад

I don’t know you personally yet, but you superstar better than 1000 Kardashians and Ronaldo 🚀🤖

Фото профиля Carlos DP 🤖🇺🇸
Carlos DP 🤖🇺🇸11 месяцев назад

Incredible work! Really solid result

Фото профиля Laksh
Laksh11 месяцев назад

@Scobleizer This looks like a huge step forward for more natural humanoid motion 👏

Фото профиля Xiatao Sun
Xiatao Sun11 месяцев назад

Impressive work! Fixing retargeting artifacts at the source rather than with complex reward engineering is the right approach. The long-horizon parkour sequence is stunning!

Фото профиля Jon Conley (e/acc)
Jon Conley (e/acc)11 месяцев назад

This is truly impressive to see how generalizeable this is and also simplifies the skill transfer process to potentially hundreds of humanoid robot vendors. Won’t be surprised to see lots of robotics companies building upon this work in the future.

Фото профиля Hongsuk Benjamin Choi
Hongsuk Benjamin Choi11 месяцев назад

this is what the community needs:) and these videos are really impressive!

Фото профиля Rathor
Rathor11 месяцев назад

Time for #HumanoidGeneration

Фото профиля Watching learner
Watching learner11 месяцев назад

It is really good to see motion tracking improved like this. OmniRetarget proves how much the right starting point matters. That is also what AIOZ Movie Review Challenge does. It gives creators a fun and simple way to share movies and connect.

Фото профиля Guojian Wang
Guojian Wang11 месяцев назад

@grok What's the meaning of retargeting?

Фото профиля Mike Domainer
Mike Domainer11 месяцев назад

should fit

Фото профиля Dr Diogenes™ 🔭 🧬🇺🇸
Dr Diogenes™ 🔭 🧬🇺🇸11 месяцев назад

That’s some really nice work! Congratulations!

Фото профиля Jude Onyenze
Jude Onyenze11 месяцев назад

@TairanHe99 If you had to choose which is more efficient, learning from Third-Person Human or Learning from Motion Capture

Фото профиля Alfred Cueva
Alfred Cueva11 месяцев назад

Interesting work! Specially since it doesn’t need to undergo curriculum training. Could OmniRetarget be made adaptive to the downstream RL task or policy uncertainty, dynamically refining trajectories?

Фото профиля Louis Le Lay
Louis Le Lay11 месяцев назад

Wow interesting stuff and you say it's generalizable to other robots? Def interested 😁

Фото профиля Chung Min Kim
Chung Min Kim11 месяцев назад

wow!!! 🥹🫶 amazing job this is so cool!!

Фото профиля 176 feet per second
176 feet per second11 месяцев назад

Pelvic and hip mobility is bloody amazing

Фото профиля VibeCodeTeddy
VibeCodeTeddy11 месяцев назад

Impressive work on solving those artifact issues! Streamlined RL with fewer reward terms is a game changer. What's next for scalability challenges?

Фото профиля schakleton
schakleton11 месяцев назад

Very impressive

Фото профиля GHOSTROSIN
GHOSTROSIN11 месяцев назад

@Tesla_Optimus

Фото профиля Maxin 🇬🇭
Maxin 🇬🇭11 месяцев назад

@Threadreaderapp unroll

Фото профиля Dirty Indy 🟥🟧🟨
Dirty Indy 🟥🟧🟨2 месяцев назад

Any real world applications such as construction or working a garbage sorting facility, parkour stuff is cool but I am seeking enlightenment elsewhere in reducing human toil

Фото профиля Big Legend (❖,❖)
Big Legend (❖,❖)11 месяцев назад

That's great.

Фото профиля Yichao Zhong
Yichao Zhong11 месяцев назад

Congrats! When will you open-source the motion retargeting code🥹

Похожие видео

NEWS: Humanoid robotics company Figure has released Helix 02, what they claim in their most capable humanoid model yet. "A single neural system that controls the full body directly from pixels, enabling dexterous, long horizon autonomy across an entire room: • Autonomous, long‑horizon loco-manipulation: Helix 02 unloads and reloads a dishwasher across a full-sized kitchen - a four-minute, end-to-end autonomous task that integrates walking, manipulation, and balance with no resets and no human intervention. We believe this is the longest horizon, most complex task completed autonomously by a humanoid robot to date. • All sensors in. All actuators out: Helix 02 connects every onboard sensor - vision, touch, and proprioception - directly to every actuator through a single unified visuomotor neural network. • Human-like whole body control from human data: All results are enabled by System 0, a learned whole‑body controller trained on over 1,000 hours of human motion data and sim‑to‑real reinforcement learning. System 0 replaces 109,504 lines of hand‑engineered C++ with a single neural prior for stable, natural motion. • New classes of dexterity: With Figure 03’s embedded tactile sensing and palm cameras, Helix 02 performs manipulation that was previously out of reach: extracting individual pills, dispensing precise syringe volumes, and singulating small, irregular objects from clutter despite self‑occlusion. Helix 02 is trained on over 1,000 hours of human motion data and integrates vision, touch, and proprioception."

Sawyer Merritt

624,910 просмотров • 7 месяцев назад

X-Humanoid just officially dropped Embodied Tien Kung 3.0, A universal platform designed to be way more open and developer-friendly. 🤖 Built on their Wise Kaiwu AI platform, this next-gen humanoid is all about slashing development costs. It’s a fully interoperable ecosystem that supports everything from tactile interaction to high-dynamic motion control at a full humanoid scale. ➤ Radical Openness: X-Humanoid is open-sourcing the full stack—robot body, motion control, VLM/VLA models, and the RoboMIND dataset. It fully supports ROS2, MQTT, and TCP/IP, so developers can customize use cases without re-engineering the basics. ➤ High-Performance Hardware: With high-torque integrated joints, Tien Kung 3.0 can clear 1-meter (3.3ft) obstacles and handle dexterous moves like kneeling and bending. It hits millimeter-level precision, making it a solid fit for industrial-grade tasks. ➤ True Autonomy: The bot runs a continuous perception-decision-execution loop. It uses world models to break down complex language commands and VLA models for real-time obstacle avoidance and navigation. ➤ Scalable Collaboration: The platform moves beyond single-unit tasks to support multi-robot collaboration with autonomous scheduling. It’s built to move embodied AI from the lab straight into real-world commercial and industrial environments. Source: X-Humanoid #Humanoid #OpenSource #Robotics #EmbodiedAI #PhysicalAI #Automation #XHumanoid #TienKung #WiseKaiwu

RoboHub🤖

49,440 просмотров • 7 месяцев назад

I don’t know if we live in a Matrix, but I know for sure that robots will spend most of their lives in simulation. Let machines train machines. I’m excited to introduce DexMimicGen, a massive-scale synthetic data generator that enables a humanoid robot to learn complex skills from only a handful of human demonstrations. Yes, as few as 5! DexMimicGen addresses the biggest pain point in robotics: where do we get data? Unlike with LLMs, where vast amounts of texts are readily available, you cannot simply download motor control signals from the internet. So researchers teleoperate the robots to collect motion data via XR headsets. They have to repeat the same skill over and over and over again, because neural nets are data hungry. This is a very slow and uncomfortable process. At NVIDIA, we believe the majority of high-quality tokens for robot foundation models will come from simulation. What DexMimicGen does is to trade GPU compute time for human time. It takes one motion trajectory from human, and multiplies into 1000s of new trajectories. A robot brain trained on this augmented dataset will generalize far better in the real world. Think of DexMimicGen as a learning signal amplifier. It maps a small dataset to a large (de facto infinite) dataset, using physics simulation in the loop. In this way, we free humans from babysitting the bots all day. The future of robot data is generative. The future of the entire robot learning pipeline will also be generative. 🧵

Jim Fan

165,246 просмотров • 1 год назад

New framework: Kick down your robot, it will get back up every time 🥋 Chinese startup RoboParty is a Beijing startup founded April 2025 by Huang Yi, originally shipping ROBOTO ORIGIN, the world's first full-stack open-source bipedal humanoid. They released UFO: Unsupervised Reinforcement Learning Framework for Humanoid Control. DEFINITIONS -> what differs is where the learning signal comes from: - SUPERVISED: humans supply the right answers (labels), the model imitates them. - UNSUPERVISED: no answer key, the model finds structure in raw data on its own. - REINFORCEMENT LEARNING: no answer key either, the model tries things and a reward scores each attempt. → UNSUPERVISED RL: trial and error where the agent invents its own rewards, instead of engineers hand-writing one per task. REPRESENTATION LEARNING: compress raw states into a useful internal map. TEMPORAL DISTANCE: distance on that map is "how many steps from A to B." CONTRASTIVE: trained by pulling together what's close in time, pushing apart what isn't. -> CONTRASTIVE TEMPORAL-DISTANCE REPRESENTATION LEARNING: the model builds an internal map of body states where distance means how many steps it takes to get from one to another. It is trained by contrast: states that occur close together in a movement get pulled together in the map, randomly paired states get pushed apart. UFO is an open-source training framework that teaches humanoid robots skills, like getting up, walking, goal-reaching, teleoperation, without reference motions -> no motion-capture or human-video demonstrations to imitate. Its core is TeCH, a contrastive temporal-distance representation-learning algorithm: the robot explores, builds pseudo-goals by temporal rolling, and learns goal-conditioned policies from a single unified progress reward. One framework trains five different robots (Unitree G1/H1, RoboParty RP0/RP1, AgiBot X2) with automatic config conversion in ~2–3 hours per robot! The real novelty here "no demonstrations at all". No data-collection arms race,the dominant humanoid-locomotion recipe is tracking: imitate mocap/retargeted-human reference trajectories. The robot self-generates goals from its own exploration and learns from a progress reward, needing zero reference motion data. Everybody else is fighting over data acquisition, while this team just teleports out of the race entirely (inb4 "competition is for losers 💀 ). This strategy reminds me of the DeepSeek playbook applied to robots: open-source the whole stack to become the global default and commoditize everyone else. RoboParty is giving away hardware and now control software (UFO) to be the Android of humanoids. Yet another reason for the US to ban Chinese open models perhaps 🥶 ? What I also really like about this approach is the cross-embodiment infrastructure, one framework trains Unitree G1/H1, RoboParty RP0/RP1, and AgiBot X2 with automatic configuration conversion. Just like Physical Intelligence, RoboParty seems to place itself as a neutral hardware agnostic middle man. Also woth mentioning: their ability ot perform stable skill injection, e.g. adding a cartwheel without forgetting how to walk. A common failure of RL humanoid policies is that teaching a new agile skill destabilizes the existing ones (catastrophic forgetting). UFO claims you can inject rare motions (cartwheel) without collapsing learned behavior. If it holds, that's a significant incremental/continual skill-learning! But again, I have to underline it: no arXiv, no external validation, no success-rate numbers. -> robotics badely needs an independent unbiased evaluator imho. Still, look at that cool demo: robot is getting kicked and pushed around (serious disturbance) during teleoperation (controlled the person at the back wearing the VR headset), and still managed to always get back up. This is some serious demonstration of stability and robustness!

Léo

36,112 просмотров • 1 месяц назад

Physics-based Motion Retargeting from Sparse Inputs paper page: Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers, providing very limited sensor data of the user's pose. Another challenge is that an avatar might have a different skeleton structure than a human and the mapping between them is unclear. In this work we address both of these challenges. We introduce a method to retarget motions in real-time from sparse human sensor data to characters of various morphologies. Our method uses reinforcement learning to train a policy to control characters in a physics simulator. We only require human motion capture data for training, without relying on artist-generated animations for each avatar. This allows us to use large motion capture datasets to train general policies that can track unseen users from real and sparse data in real-time. We demonstrate the feasibility of our approach on three characters with different skeleton structure: a dinosaur, a mouse-like creature and a human. We show that the avatar poses often match the user surprisingly well, despite having no sensor information of the lower body available. We discuss and ablate the important components in our framework, specifically the kinematic retargeting step, the imitation, contact and action reward as well as our asymmetric actor-critic observations. We further explore the robustness of our method in a variety of settings including unbalancing, dancing and sports motions.

AK

106,527 просмотров • 3 лет назад