Загрузка видео...
Не удалось загрузить видео
Introducing Video Language Planning! By planning across the space of generated videos/language, we can synthesize long-horizon video plans and solve much longer horizon tasks than existing baseline (such as RT-2 and PALM-E). (1/5)
90,272 просмотров • 3 лет назад •via X (Twitter)
Комментарии: 10

Video Language Planning (VLP) combines the strengths of VLMs and text-to-video models to jointly synthesize detailed video plans of actions to execute using a tree-search procedure. Video plans are then converted to actions using a goal-conditioned policy. (2/5)

Planning substantially improves the performance of VLP. Below, we illustrate generated video plans to construct a line by either: (1) using a single text-to-video model, (2) using a VLM + text-to-video model without search, (3) using a VLM+text-to-video with search. (3/5)

By planning and refining a sequence of actions, on long-horizon tasks, VLP substantially outperforms existing approaches such as RT-2 or PALM-E which autoregressively predict actions and often get stuck at OOD states The importance of model based planning @ylecun! (4/5)

Work done with amazing collaborators @mengjiao_yang , @peteflorence , @xf1280, @ayzwah, @brian_ichter, Pierre Sermanet, @TianheYu, @pabbeel, Josh Tenenbaum, Leslie Kaelbling, @andyzeng_, Jonathan Thompson (5/5)

can i interpret it as a combination or rollout policy, critic function and learned dynamic model, by these things we can do muzero style planning.

yup, where the rollouts are surprisingly super consistent even after many number of steps

@Scobleizer Incredible, this is a mind-blowing innovation.

Awesome work! 🦾🤖✨

Wow that's great

Beautiful! You guys are on fire!
