Loading video...

Video Failed to Load

Go Home

#CVPR2024 Highlight🌟 SimXR: Real-Time Simulated Avatar from Head-Mounted Sensors Controlling simulated humanoids and estimate full-body pose using SLAM images and headset pose in real time. 🌐: 📜: 🧑🏻‍💻: Input: four SLAM views from Quest 2 headset and 6DoF headset pose. Output: full-body humanoid control signals. Leveraging the laws of...

36,413 views • 2 years ago •via X (Twitter)

11 Comments

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

How to learn a photon to actuations full-body control policy? Distill from PHC! First, we create large-scale synthetic data with paired full-body motion and egocentric images. Then, train a control policy to consume these input streams and query PHC for pseudo-gt torques as supervision.

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

While we only train using synthetic data, SimXR generalizes to real-world data captures. Here is some synthetic data results: we randomize background and lighting at a per-frame level.

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

This pipeline is also applicable to AR headsets! Here is some results trained and tested on the Aria Digital Twin (ADT) dataset (two views).

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

All code and data will be released at Repo is a work in progress at the moment 🤫

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

Thanks to everyone on the team! @jinkuncao @Me_Rawal @awinkler_ Jing Huang @kkitani @xuweipeng000 Work done at Meta Reality Labs, Pittsburgh in collaboration with CMU.

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

Thoughts: learning a universal humanoid motion imitator that has scaled to a large motion dataset like AMASS really pays off. For tasks like animation, pose estimation, and teleoperation, either in simulation only or real-world (cue a strong imitator can already be used for many tasks (using kinematic pose as the intermediate representation). It is also the gateway for acquiring motor skills (cue where you can learn compact motion representations that will be universally applicable to downstream tasks. PULSE has been tremendously useful in my projects, which I will share soon 😁

Chen Tessler's profile picture
Chen Tessler2 years ago

Nice! I need to hook up my quest to Isaac. This looks so much fun! 🤩

FutureMapper ᯅ's profile picture
FutureMapper ᯅ2 years ago

How are you accessing the camera?

PowerBeatsVR's profile picture
PowerBeatsVR3 years ago

Get ready for a full-body VR workout that’s fun, fast, and intuitive — Play PowerBeatsVR (Now on Meta Quest) 🔥

haareblond's profile picture
haareblond2 years ago

wow cool this might make upper or full body tracking possible for some diy projects. very impressive. :o

Zhengyi “Zen” Luo's profile picture
Zhengyi “Zen” Luo2 years ago

Meta’s IOBT will add support for that very soon 😇

Related Videos

Physics-based Motion Retargeting from Sparse Inputs paper page: Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers, providing very limited sensor data of the user's pose. Another challenge is that an avatar might have a different skeleton structure than a human and the mapping between them is unclear. In this work we address both of these challenges. We introduce a method to retarget motions in real-time from sparse human sensor data to characters of various morphologies. Our method uses reinforcement learning to train a policy to control characters in a physics simulator. We only require human motion capture data for training, without relying on artist-generated animations for each avatar. This allows us to use large motion capture datasets to train general policies that can track unseen users from real and sparse data in real-time. We demonstrate the feasibility of our approach on three characters with different skeleton structure: a dinosaur, a mouse-like creature and a human. We show that the avatar poses often match the user surprisingly well, despite having no sensor information of the lower body available. We discuss and ablate the important components in our framework, specifically the kinematic retargeting step, the imitation, contact and action reward as well as our asymmetric actor-critic observations. We further explore the robustness of our method in a variety of settings including unbalancing, dancing and sports motions.

AK

106,527 views • 3 years ago