Загрузка видео...

Не удалось загрузить видео

На главную

🤖What if a robot could perform a new task just from a natural language command, with zero demonstrations? Our new work, NovaFlow, makes it possible! We use pre-trained video generative model to create a video of the task, then translate it into a plan for real-world robot execution. 1/6...

105,915 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 34

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

We start by giving a video model a text prompt like "hang the mug." It generates a short clip of a human performing the task, acting as a visual instruction manual. This video provides the commonsense physical knowledge needed for the robot to understand the motion. 🎥 2/6

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

The generated video is 2D, so how does the robot understand 3D space? We lift it! 📏 A monocular video depth estimation model predicts depth for every frame. Then, we calibrate it with the robot's real-world depth camera for precise spatial awareness. 3/6

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

🤔 But how do you bridge the gap between that video and physical control? We introduce "actionable 3D object flow," a novel representation distilled from the video that guides the robot’s motion for complex rigid, articulated, and even deformable objects. 🌟 4/6

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

NovaFlow shows incredible generalization. From rigid blocks to deformable ropes, on both a tabletop arm and a mobile Spot! It adapts to new objects and embodiments without any retraining. It achieves SOTA zero-shot results, even outperforming policies trained on 10-30 demos. 🏆 5/6

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

By leveraging the commonsense knowledge in video models, NovaFlow helps solve the robot data bottleneck. This is a leap towards more general-purpose robots. 🚀 Dive deeper into our work: 📜 Paper: 🌐 Project Website: 👀 P.S. You can check out the 3D generated videos and planned robot trajectories in a live viewer on our website! 6/6

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

This is a collaboration with @ling_feng_sun, @YafeiHuCMU, Duy Ta, Jenny Barry, and @jiahuifu_carol.

Фото профиля Nils
Nils11 месяцев назад

Beautiful.

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks, Nils!

Фото профиля Xiatao Sun
Xiatao Sun11 месяцев назад

Using video generation models to create task plans from language is clever! The actionable 3D object flow representation seems like the key to making this work across embodiments.

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks, Xiatao. I believe video models capture the most general form of dynamics, although they can’t be directly understood by robots. Actionable flow enables robots to comprehend the generated videos.

Фото профиля Zhiyang (Frank) Dou
Zhiyang (Frank) Dou11 месяцев назад

Nice work!

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks, Zhiyang!

Фото профиля Bozheng Li
Bozheng Li11 месяцев назад

great work!

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks, Bozheng!

Фото профиля Balveer Singh
Balveer Singh11 месяцев назад

Wow, this is such a cool approach.

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks🙏

Фото профиля Ted Xiao
Ted Xiao11 месяцев назад

Nice work!

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks, Ted! 🙏

Фото профиля Sam Green
Sam Green11 месяцев назад

Laying the groundwork for Soma

Фото профиля Atharva Sawant
Atharva Sawant11 месяцев назад

So cool

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks

Фото профиля Cryptlesh
Cryptlesh11 месяцев назад

Wow. It’s like imagining whole take on head and then implementing it like human break does. Kind of.

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Exactly, a video model serves as the imagination capabilities of human beings.

Фото профиля Bercan
Bercan11 месяцев назад

Great work. Can you try on these examples in mcap.

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

Thanks for the great suggestions! Egocentric data is more challenging, and we’re still working on it!

Фото профиля Bercan
Bercan11 месяцев назад

Happy to work together we are also working on real2sim

Фото профиля Stanley Wei
Stanley Wei11 месяцев назад

That will save a lot of time. Hope you post some cool demos.

Фото профиля BlockChan
BlockChan9 месяцев назад

holy shit i personaly owned name for a while and have no idea what to build

Фото профиля towerofshadow
towerofshadow11 месяцев назад

is it more effective to train robots using human actions that are tracked using actionable flow, or to teleoperate robots into direct movement data?

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

NovaFlow is object-centric and embodiment-agnostic. So it should work the same for both human and robot data. In practice, we find video generative models work better with generating human data.

Фото профиля Danny Ng
Danny Ng11 месяцев назад

I wonder how long it takes to perform one task?

Фото профиля Hongyu Li
Hongyu Li11 месяцев назад

It depends on the choice of video model. For Veo, it takes around 2-3min for the entire pipeline.

Фото профиля Urs Gehrig
Urs Gehrig11 месяцев назад

Zero shot Robotic AI manipulation.

Фото профиля Mariano Phielipp
Mariano Phielipp6 месяцев назад

Interesting work. How you see the work of Florence, Manuelli, Redrake in the context of this work Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation Florence,Manuelli, Russ Tedrake?

Похожие видео