Video wird geladen...
Video konnte nicht geladen werden
For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and depth data any robot's sensors would observe on any trajectory. Robots can now be trained and tested in far more spaces. Until now, scanning spaces like these required expensive equipment and... show more
147,955 Aufrufe • vor 19 Tagen •via X (Twitter)
10 Kommentare

Atlas understands both the spatial structure of the world and how it evolves over time. This lets us turn a handful of ordinary cameras into a "bullet time" multiview capture studio, capturing dynamic moments frozen in time.

Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models. Passing more images gives Atlas more context: the more it sees, the less it imagines.

By positioning multiple input views within the model's spatial context, Atlas allows you to generate image and video frames with precise control. Here, we hand-designed a camera trajectory to generate a 1 minute video at 1440p resolution from seven reference images.

We pre-trained Atlas from scratch to take multimodal inputs, including camera movement, and turn it into 3D grounded views. Atlas puts you in control: direct the views, reconstruct real spaces with your inputs, and build explorable worlds. Read more:

Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

All of this is powered by one unified architecture: a multimodal autoregressive diffusion transformer, pretrained from scratch. This foundation blends the best of modern LLMs and video models, benefiting from the architectural, algorithmic, and systems advances from both areas.

Atlas is our scalable foundation model that can perceive, generate, reason, and interact with virtual and physical worlds. We're opening up early access in the coming weeks. Learn more and sign up here:

How accurately can this reconstruct the real world from a live video feed? Also is it running on anything in near realtime? Seems like image inputs could be used to generate a spatial map a robot can use for navigation.

@grok is this like a substitute to Lidar sensor ?

this may matter more than a prettier simulator: every real space becomes a test site. can Atlas persist object-state changes across actions, or is it mainly generating sensor observations? without persistent state, the robot learns a visual world, not a causal one.
