Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Video Language Planning! By planning across the space of generated videos/language, we can synthesize long-horizon video plans and solve much longer horizon tasks than existing baseline (such as RT-2 and PALM-E). (1/5)

90,272 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Yilun Du
Yilun Duvor 3 Jahren

Video Language Planning (VLP) combines the strengths of VLMs and text-to-video models to jointly synthesize detailed video plans of actions to execute using a tree-search procedure. Video plans are then converted to actions using a goal-conditioned policy. (2/5)

Profilbild von Yilun Du
Yilun Duvor 3 Jahren

Planning substantially improves the performance of VLP. Below, we illustrate generated video plans to construct a line by either: (1) using a single text-to-video model, (2) using a VLM + text-to-video model without search, (3) using a VLM+text-to-video with search. (3/5)

Profilbild von Yilun Du
Yilun Duvor 3 Jahren

By planning and refining a sequence of actions, on long-horizon tasks, VLP substantially outperforms existing approaches such as RT-2 or PALM-E which autoregressively predict actions and often get stuck at OOD states The importance of model based planning @ylecun! (4/5)

Profilbild von Yilun Du
Yilun Duvor 3 Jahren

Work done with amazing collaborators @mengjiao_yang , @peteflorence , @xf1280, @ayzwah, @brian_ichter, Pierre Sermanet, @TianheYu, @pabbeel, Josh Tenenbaum, Leslie Kaelbling, @andyzeng_, Jonathan Thompson (5/5)

Profilbild von Xidong Feng
Xidong Fengvor 3 Jahren

can i interpret it as a combination or rollout policy, critic function and learned dynamic model, by these things we can do muzero style planning.

Profilbild von Yilun Du
Yilun Duvor 3 Jahren

yup, where the rollouts are surprisingly super consistent even after many number of steps

Profilbild von Rufus
Rufusvor 3 Jahren

@Scobleizer Incredible, this is a mind-blowing innovation.

Profilbild von William Lamkin
William Lamkinvor 3 Jahren

Awesome work! 🦾🤖✨

Profilbild von Jaisurya
Jaisuryavor 3 Jahren

Wow that's great

Profilbild von lucas
lucasvor 3 Jahren

Beautiful! You guys are on fire!

Ähnliche Videos

JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models paper page: Achieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle certain long-horizon tasks in an open world. However, they still struggle when the number of open-world tasks could potentially be infinite and lack the capability to progressively enhance task completion as game time progresses. We introduce JARVIS-1, an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, all within the popular yet challenging open-world Minecraft universe. Specifically, we develop JARVIS-1 on top of pre-trained multimodal language models, which map visual observations and textual instructions to plans. The plans will be ultimately dispatched to the goal-conditioned controllers. We outfit JARVIS-1 with a multimodal memory, which facilitates planning using both pre-trained knowledge and its actual game survival experiences. In our experiments, JARVIS-1 exhibits nearly perfect performances across over 200 varying tasks from the Minecraft Universe Benchmark, ranging from entry to intermediate levels. JARVIS-1 has achieved a completion rate of 12.5% in the long-horizon diamond pickaxe task. This represents a significant increase up to 5 times compared to previous records. Furthermore, we show that JARVIS-1 is able to self-improve following a life-long learning paradigm thanks to multimodal memory, sparking a more general intelligence and improved autonomy.

AK

141,440 Aufrufe • vor 2 Jahren

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 Aufrufe • vor 3 Jahren