Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Video Language Planning! By planning across the space of generated videos/language, we can synthesize long-horizon video plans and solve much longer horizon tasks than existing baseline (such as RT-2 and PALM-E). (1/5)

90,272 просмотров • 3 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля Yilun Du
Yilun Du3 лет назад

Video Language Planning (VLP) combines the strengths of VLMs and text-to-video models to jointly synthesize detailed video plans of actions to execute using a tree-search procedure. Video plans are then converted to actions using a goal-conditioned policy. (2/5)

Фото профиля Yilun Du
Yilun Du3 лет назад

Planning substantially improves the performance of VLP. Below, we illustrate generated video plans to construct a line by either: (1) using a single text-to-video model, (2) using a VLM + text-to-video model without search, (3) using a VLM+text-to-video with search. (3/5)

Фото профиля Yilun Du
Yilun Du3 лет назад

By planning and refining a sequence of actions, on long-horizon tasks, VLP substantially outperforms existing approaches such as RT-2 or PALM-E which autoregressively predict actions and often get stuck at OOD states The importance of model based planning @ylecun! (4/5)

Фото профиля Yilun Du
Yilun Du3 лет назад

Work done with amazing collaborators @mengjiao_yang , @peteflorence , @xf1280, @ayzwah, @brian_ichter, Pierre Sermanet, @TianheYu, @pabbeel, Josh Tenenbaum, Leslie Kaelbling, @andyzeng_, Jonathan Thompson (5/5)

Фото профиля Xidong Feng
Xidong Feng3 лет назад

can i interpret it as a combination or rollout policy, critic function and learned dynamic model, by these things we can do muzero style planning.

Фото профиля Yilun Du
Yilun Du3 лет назад

yup, where the rollouts are surprisingly super consistent even after many number of steps

Фото профиля Rufus
Rufus3 лет назад

@Scobleizer Incredible, this is a mind-blowing innovation.

Фото профиля William Lamkin
William Lamkin3 лет назад

Awesome work! 🦾🤖✨

Фото профиля Jaisurya
Jaisurya3 лет назад

Wow that's great

Фото профиля lucas
lucas3 лет назад

Beautiful! You guys are on fire!

Похожие видео

JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models paper page: Achieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle certain long-horizon tasks in an open world. However, they still struggle when the number of open-world tasks could potentially be infinite and lack the capability to progressively enhance task completion as game time progresses. We introduce JARVIS-1, an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, all within the popular yet challenging open-world Minecraft universe. Specifically, we develop JARVIS-1 on top of pre-trained multimodal language models, which map visual observations and textual instructions to plans. The plans will be ultimately dispatched to the goal-conditioned controllers. We outfit JARVIS-1 with a multimodal memory, which facilitates planning using both pre-trained knowledge and its actual game survival experiences. In our experiments, JARVIS-1 exhibits nearly perfect performances across over 200 varying tasks from the Minecraft Universe Benchmark, ranging from entry to intermediate levels. JARVIS-1 has achieved a completion rate of 12.5% in the long-horizon diamond pickaxe task. This represents a significant increase up to 5 times compared to previous records. Furthermore, we show that JARVIS-1 is able to self-improve following a life-long learning paradigm thanks to multimodal memory, sparking a more general intelligence and improved autonomy.

AK

141,440 просмотров • 2 лет назад

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 просмотров • 3 лет назад