Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and depth data any robot's sensors would observe on any trajectory. Robots can now be trained and tested in far more spaces. Until now, scanning spaces like these required expensive equipment and...

147,955 görüntüleme • 19 gün önce •via X (Twitter)

10 Yorum

World Labs profil fotoğrafı
World Labs19 gün önce

Atlas understands both the spatial structure of the world and how it evolves over time. This lets us turn a handful of ordinary cameras into a "bullet time" multiview capture studio, capturing dynamic moments frozen in time.

World Labs profil fotoğrafı
World Labs19 gün önce

Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models. Passing more images gives Atlas more context: the more it sees, the less it imagines.

World Labs profil fotoğrafı
World Labs19 gün önce

By positioning multiple input views within the model's spatial context, Atlas allows you to generate image and video frames with precise control. Here, we hand-designed a camera trajectory to generate a 1 minute video at 1440p resolution from seven reference images.

World Labs profil fotoğrafı
World Labs19 gün önce

We pre-trained Atlas from scratch to take multimodal inputs, including camera movement, and turn it into 3D grounded views. Atlas puts you in control: direct the views, reconstruct real spaces with your inputs, and build explorable worlds. Read more:

World Labs profil fotoğrafı
World Labs19 gün önce

Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

World Labs profil fotoğrafı
World Labs19 gün önce

All of this is powered by one unified architecture: a multimodal autoregressive diffusion transformer, pretrained from scratch. This foundation blends the best of modern LLMs and video models, benefiting from the architectural, algorithmic, and systems advances from both areas.

World Labs profil fotoğrafı
World Labs19 gün önce

Atlas is our scalable foundation model that can perceive, generate, reason, and interact with virtual and physical worlds. We're opening up early access in the coming weeks. Learn more and sign up here:

Sammy profil fotoğrafı
Sammy19 gün önce

How accurately can this reconstruct the real world from a live video feed? Also is it running on anything in near realtime? Seems like image inputs could be used to generate a spatial map a robot can use for navigation.

Ezatullah Matin wakily profil fotoğrafı
Ezatullah Matin wakily18 gün önce

@grok is this like a substitute to Lidar sensor ?

Mob · INT AIOS profil fotoğrafı
Mob · INT AIOS18 gün önce

this may matter more than a prettier simulator: every real space becomes a test site. can Atlas persist object-state changes across actions, or is it mainly generating sensor observations? without persistent state, the robot learns a visual world, not a causal one.

Benzer Videolar

Force-sensing fingers! 🧤 Stanford researchers just released UMI-FT, a handheld data collection platform that puts compact six-axis force/torque sensors on each finger, enabling finger-level wrench measurements alongside RGB, depth, and pose data. Many manipulation tasks require careful force modulation: too little force and the task fails, too much and you cause damage. But commercial force/torque sensors are expensive, bulky, and fragile, which has limited large-scale force-aware policy learning. UMI-FT changes the economics. The platform uses an iPhone for RGB vision, ultrawide RGB, depth, and pose via ARKit, with each finger sensorized using a CoinFT sensor to capture per-finger wrench information during manipulation. This multimodal data trains an adaptive compliance policy that predicts position targets, grasp force, and stiffness for execution on standard compliance controllers. The learned policy runs slowest and generates reference targets, while model-based compliance and force controllers provide delicate 6D compliance control and real-time force modulation. They tested on three contact-rich, force-sensitive tasks: whiteboard wiping (locate eraser, grasp, wipe until clean), skewering zucchini (grasp slice firmly, push onto stick until punctured), and lightbulb insertion (grasp bulb, align bayonet pin with socket slit, insert while overcoming spring force, rotate to light up). The results are clear. Policies without compliance struggle to modulate contact force and trigger safety faults from excessive force. Policies without force sensing fail to grasp unseen objects or resist reaction forces, causing slippage. Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

12,868 görüntüleme • 7 ay önce