Loading video...

Video Failed to Load

Go Home

Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun...

96,195 views โ€ข 3 months ago โ€ขvia X (Twitter)

15 Comments

Mahi Shafiullah ๐Ÿ ๐Ÿค–'s profile picture
Mahi Shafiullah ๐Ÿ ๐Ÿค–3 months ago

The idea is simple: throw in any reasonable human video and some compute, and out comes robot actions doing the same thing the human is doing. But to make it work, we had to tackle the breadth of human data, hand-object interaction, and dexterous retargeting. 2/n

Mahi Shafiullah ๐Ÿ ๐Ÿค–'s profile picture
Mahi Shafiullah ๐Ÿ ๐Ÿค–3 months ago

Most internet-grade โ€œhuman dataโ€ has terrible quality: often theyโ€™re blurry, with no meaningful interaction happening, and hand/objects are constantly going out of frame. Want to know if your human data is robot ready? Pass them through our Do as I Do. For example: we start with 2000 100DOH clips that are already filtered for hand/object interactions, and find that only 107 of them actually have the necessary info. Do as I Do can reconstruct 83 of those. 3/n

Mahi Shafiullah ๐Ÿ ๐Ÿค–'s profile picture
Mahi Shafiullah ๐Ÿ ๐Ÿค–3 months ago

For hand-object tracking from monocular RGB, we used many of the nice perks of modern CV, and made stuff on our own when that wasnโ€™t enough. For example, did you know you can get SAM3D to do 4D object tracking? Under the hood, itโ€™s a diffusion decoder; so tracking โ‰ˆ biasing the distribution with current object priors! The technical details matter, so check out the paper. 4/n

Mahi Shafiullah ๐Ÿ ๐Ÿค–'s profile picture
Mahi Shafiullah ๐Ÿ ๐Ÿค–3 months ago

Sometimes there is genuine ambiguity in the data, and vision canโ€™t resolve it. What can you do in this case? You use physics! The physics-grounding step that transfers motion to robot actions also validates that any motion we propose is achievable by a robot. Thus, even when the object wants to float away, the physics staples it to the hand. 5/n

Mahi Shafiullah ๐Ÿ ๐Ÿค–'s profile picture
Mahi Shafiullah ๐Ÿ ๐Ÿค–3 months ago

Lots of gems like this in the full work, led by @bhawna_paliwal_, @HarithejaE, & @willjhliang. Also special thanks to @pabbeel & @JitendraMalikCV for advice, and @kyutai_labs for the compute support. Try our code today: ๐Ÿ“: ๐ŸŒ: โš™๏ธ:

Chenhao Li's profile picture
Chenhao Li3 months ago

Very cool man

Michael Cho - Rbt/Acc's profile picture
Michael Cho - Rbt/Acc3 months ago

Great stuff!! I got lots of qns (like how u do the physics grounding)...would love to have u guys come on robopapers to share more ๐Ÿ™

Sharpa's profile picture
Sharpa3 months ago

incredible achievement! congrats to the team!

R2rule's profile picture
R2rule3 months ago

"Do what I do" is a foolish approach... It's full of unnecessary and useless data which comes from ring and pinky fingers. People use only 3 fingers in 99.9% of their daily tasks. โ€œMove object as i moveโ€ is the reliable model to teach AI..

SEAR's profile picture
SEAR3 months ago

Lab data is too fake Just pay people to cook dinner in smart glasses for the real chaotic data

SEAR's profile picture
SEAR3 months ago

Lab data is too fake Paying people to cook in smart glasses gives robots the real chaos they need

Paolo AI's profile picture
Paolo AI3 months ago

I have some hands on experience with a similar approach and what we should do on the IK for Mujoco is to use additional constraints to favour natural human positions and avoid singularity points. Example:

EgoScale's profile picture
EgoScale3 months ago

This is the right kind of messy. Human videos wonโ€™t look like clean teleop data, but they contain the stuff robots actually struggle with: contact, recovery, small corrections, changing context. Making that usable is the unlock.

Liam "OD"'s profile picture
Liam "OD"3 months ago

Exactly. Teleoperation is boutique. I suspect open-source robotics scales on uncurated internet video.

Roger Bellver's profile picture
Roger Bellver3 months ago

Very interesting! And thank you for the contribution to the public through a paper. I find the potential to use AI-generated videos to train physical AI robot incredibly interesting. Another example of AI training AI, how great.

Related Videos

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's ๐Ÿ”ฅ Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. โ†’ 1,000 hours pre-training โ†’ 20% normalised task performance โ†’ 10,000 hours โ†’ 28% โ†’ 100,000 hours โ†’ 45% โ†’ 1,000,000 hours โ†’ 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ โ™ป๏ธ Join the weekly robotics newsletter, and never miss any news โ†’

Lukas Ziegler

23,681 views โ€ข 1 month ago

Exciting updates on Project GR00T! We discover a systematic way to scale up robot data, tackling the most painful pain point in robotics. The idea is simple: human collects demonstration on a real robot, and we multiply that data 1000x or more in simulation. Letโ€™s break it down: 1. We use Apple Vision Pro (yes!!) to give the human operator first person control of the humanoid. Vision Pro parses human hand pose and retargets the motion to the robot hand, all in real time. From the humanโ€™s point of view, they are immersed in another body like the Avatar. Teleoperation is slow and time-consuming, but we can afford to collect a small amount of data. 2. We use RoboCasa, a generative simulation framework, to multiply the demonstration data by varying the visual appearance and layout of the environment. In Jensenโ€™s keynote video below, the humanoid is now placing the cup in hundreds of kitchens with a huge diversity of textures, furniture, and object placement. We only have 1 physical kitchen at the GEAR Lab in NVIDIA HQ, but we can conjure up infinite ones in simulation. 3. Finally, we apply MimicGen, a technique to multiply the above data even more by varying the *motion* of the robot. MimicGen generates vast number of new action trajectories based on the original human data, and filters out failed ones (e.g. those that drop the cup) to form a much larger dataset. To sum up, given 1 human trajectory with Vision Pro -> RoboCasa produces N (varying visuals) -> MimicGen further augments to NxM (varying motions). This is the way to trade compute for expensive human data by GPU-accelerated simulation. A while ago, I mentioned that teleoperation is fundamentally not scalable, because we are always limited by 24 hrs/robot/day in the world of atoms. Our new GR00T synthetic data pipeline breaks this barrier in the world of bits. Scaling has been so much fun for LLMs, and it's finally our turn to have fun in robotics! We are building tools to enable everyone in the ecosystem to scale up with us. Links in thread:

Jim Fan

364,670 views โ€ข 2 years ago