Loading video...

Video Failed to Load

Go Home

🤖What if a robot could perform a new task just from a natural language command, with zero demonstrations? Our new work, NovaFlow, makes it possible! We use pre-trained video generative model to create a video of the task, then translate it into a plan for real-world robot execution. 1/6...

105,915 views • 11 months ago •via X (Twitter)

34 Comments

Hongyu Li's profile picture
Hongyu Li11 months ago

We start by giving a video model a text prompt like "hang the mug." It generates a short clip of a human performing the task, acting as a visual instruction manual. This video provides the commonsense physical knowledge needed for the robot to understand the motion. 🎥 2/6

Hongyu Li's profile picture
Hongyu Li11 months ago

The generated video is 2D, so how does the robot understand 3D space? We lift it! 📏 A monocular video depth estimation model predicts depth for every frame. Then, we calibrate it with the robot's real-world depth camera for precise spatial awareness. 3/6

Hongyu Li's profile picture
Hongyu Li11 months ago

🤔 But how do you bridge the gap between that video and physical control? We introduce "actionable 3D object flow," a novel representation distilled from the video that guides the robot’s motion for complex rigid, articulated, and even deformable objects. 🌟 4/6

Hongyu Li's profile picture
Hongyu Li11 months ago

NovaFlow shows incredible generalization. From rigid blocks to deformable ropes, on both a tabletop arm and a mobile Spot! It adapts to new objects and embodiments without any retraining. It achieves SOTA zero-shot results, even outperforming policies trained on 10-30 demos. 🏆 5/6

Hongyu Li's profile picture
Hongyu Li11 months ago

By leveraging the commonsense knowledge in video models, NovaFlow helps solve the robot data bottleneck. This is a leap towards more general-purpose robots. 🚀 Dive deeper into our work: 📜 Paper: 🌐 Project Website: 👀 P.S. You can check out the 3D generated videos and planned robot trajectories in a live viewer on our website! 6/6

Hongyu Li's profile picture
Hongyu Li11 months ago

This is a collaboration with @ling_feng_sun, @YafeiHuCMU, Duy Ta, Jenny Barry, and @jiahuifu_carol.

Nils's profile picture
Nils11 months ago

Beautiful.

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks, Nils!

Xiatao Sun's profile picture
Xiatao Sun11 months ago

Using video generation models to create task plans from language is clever! The actionable 3D object flow representation seems like the key to making this work across embodiments.

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks, Xiatao. I believe video models capture the most general form of dynamics, although they can’t be directly understood by robots. Actionable flow enables robots to comprehend the generated videos.

Zhiyang (Frank) Dou's profile picture
Zhiyang (Frank) Dou11 months ago

Nice work!

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks, Zhiyang!

Bozheng Li's profile picture
Bozheng Li11 months ago

great work!

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks, Bozheng!

Balveer Singh's profile picture
Balveer Singh11 months ago

Wow, this is such a cool approach.

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks🙏

Ted Xiao's profile picture
Ted Xiao11 months ago

Nice work!

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks, Ted! 🙏

Sam Green's profile picture
Sam Green11 months ago

Laying the groundwork for Soma

Atharva Sawant's profile picture
Atharva Sawant11 months ago

So cool

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks

Cryptlesh's profile picture
Cryptlesh11 months ago

Wow. It’s like imagining whole take on head and then implementing it like human break does. Kind of.

Hongyu Li's profile picture
Hongyu Li11 months ago

Exactly, a video model serves as the imagination capabilities of human beings.

Bercan's profile picture
Bercan11 months ago

Great work. Can you try on these examples in mcap.

Hongyu Li's profile picture
Hongyu Li11 months ago

Thanks for the great suggestions! Egocentric data is more challenging, and we’re still working on it!

Bercan's profile picture
Bercan11 months ago

Happy to work together we are also working on real2sim

Stanley Wei's profile picture
Stanley Wei11 months ago

That will save a lot of time. Hope you post some cool demos.

BlockChan's profile picture
BlockChan9 months ago

holy shit i personaly owned name for a while and have no idea what to build

towerofshadow's profile picture
towerofshadow11 months ago

is it more effective to train robots using human actions that are tracked using actionable flow, or to teleoperate robots into direct movement data?

Hongyu Li's profile picture
Hongyu Li11 months ago

NovaFlow is object-centric and embodiment-agnostic. So it should work the same for both human and robot data. In practice, we find video generative models work better with generating human data.

Danny Ng's profile picture
Danny Ng11 months ago

I wonder how long it takes to perform one task?

Hongyu Li's profile picture
Hongyu Li11 months ago

It depends on the choice of video model. For Veo, it takes around 2-3min for the entire pipeline.

Urs Gehrig's profile picture
Urs Gehrig11 months ago

Zero shot Robotic AI manipulation.

Mariano Phielipp's profile picture
Mariano Phielipp6 months ago

Interesting work. How you see the work of Florence, Manuelli, Redrake in the context of this work Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation Florence,Manuelli, Russ Tedrake?

Related Videos