Loading video...

Video Failed to Load

Go Home

Punchline: World models == VQA (about the future)! Planning with world models can be powerful for robotics/control. But most world models are video generators trained to predict everything, including irrelevant pixels and distractions. We ask - what if a world model only predicted the semantic information necessary for decision-making?...

61,208 views • 8 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕 ✈️ RSS

1,077,988 views • 1 month ago

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,708 views • 3 years ago

In 2026, AI world models will take the spotlight in storytelling - powering new types of interactive experiences & digital economies not seen before World models are progressing rapidly - Marble World Labs and Genie 3 Google DeepMind already generate 3D environments from text prompts, allowing users to explore them as if they were video games As creators adopt these tools, new storytelling formats will emerge. One genre I'm excited about is "generative Minecraft" - where players co-create virtual worlds together by vibe coding with world models. Game mechanics could be programmable with natural language - ex. "create a paintbrush that changes the color of anything I touch to pink" World models will also likely give rise to not just a single game, but an entire new category of generative world experiences - you could have a horror experience where you’re hiding from generated monsters, or a D&D experience where you’re roaming an infinite fantasy world with friends And with a common base model for the underlying worlds, these experiences could be inter-connected in a multiverse we could only dream about previously in science fiction A key affordance here is the role of consumers as co-creators - you can wander the multiverse as a tourist, or break out your pickaxe and become a creator anytime. This in turn would give rise to new digital economies - with creators making a living building and selling interoperable assets, serving as a guide for new players, etc The opportunity is enormous - a new category of generative worlds would not only create a new storytelling medium unlike any we’ve seen before, but also be rich training grounds for agents, robotics, and AGI If you’re excited about building new interactive experiences or the virtual economy stack with world models - we’d love to hear from you!

Jon Lai

29,871 views • 7 months ago