Загрузка видео...

Не удалось загрузить видео

На главную

For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and depth data any robot's sensors would observe on any trajectory. Robots can now be trained and tested in far more spaces. Until now, scanning spaces like these required expensive equipment and...

147,955 просмотров • 18 дней назад •via X (Twitter)

Комментарии: 10

Фото профиля World Labs
World Labs18 дней назад

Atlas understands both the spatial structure of the world and how it evolves over time. This lets us turn a handful of ordinary cameras into a "bullet time" multiview capture studio, capturing dynamic moments frozen in time.

Фото профиля World Labs
World Labs18 дней назад

Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models. Passing more images gives Atlas more context: the more it sees, the less it imagines.

Фото профиля World Labs
World Labs18 дней назад

By positioning multiple input views within the model's spatial context, Atlas allows you to generate image and video frames with precise control. Here, we hand-designed a camera trajectory to generate a 1 minute video at 1440p resolution from seven reference images.

Фото профиля World Labs
World Labs18 дней назад

We pre-trained Atlas from scratch to take multimodal inputs, including camera movement, and turn it into 3D grounded views. Atlas puts you in control: direct the views, reconstruct real spaces with your inputs, and build explorable worlds. Read more:

Фото профиля World Labs
World Labs18 дней назад

Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

Фото профиля World Labs
World Labs18 дней назад

All of this is powered by one unified architecture: a multimodal autoregressive diffusion transformer, pretrained from scratch. This foundation blends the best of modern LLMs and video models, benefiting from the architectural, algorithmic, and systems advances from both areas.

Фото профиля World Labs
World Labs18 дней назад

Atlas is our scalable foundation model that can perceive, generate, reason, and interact with virtual and physical worlds. We're opening up early access in the coming weeks. Learn more and sign up here:

Фото профиля Sammy
Sammy18 дней назад

How accurately can this reconstruct the real world from a live video feed? Also is it running on anything in near realtime? Seems like image inputs could be used to generate a spatial map a robot can use for navigation.

Фото профиля Ezatullah Matin wakily
Ezatullah Matin wakily18 дней назад

@grok is this like a substitute to Lidar sensor ?

Фото профиля Mob · INT AIOS
Mob · INT AIOS18 дней назад

this may matter more than a prettier simulator: every real space becomes a test site. can Atlas persist object-state changes across actions, or is it mainly generating sensor observations? without persistent state, the robot learns a visual world, not a causal one.

Похожие видео

Force-sensing fingers! 🧤 Stanford researchers just released UMI-FT, a handheld data collection platform that puts compact six-axis force/torque sensors on each finger, enabling finger-level wrench measurements alongside RGB, depth, and pose data. Many manipulation tasks require careful force modulation: too little force and the task fails, too much and you cause damage. But commercial force/torque sensors are expensive, bulky, and fragile, which has limited large-scale force-aware policy learning. UMI-FT changes the economics. The platform uses an iPhone for RGB vision, ultrawide RGB, depth, and pose via ARKit, with each finger sensorized using a CoinFT sensor to capture per-finger wrench information during manipulation. This multimodal data trains an adaptive compliance policy that predicts position targets, grasp force, and stiffness for execution on standard compliance controllers. The learned policy runs slowest and generates reference targets, while model-based compliance and force controllers provide delicate 6D compliance control and real-time force modulation. They tested on three contact-rich, force-sensitive tasks: whiteboard wiping (locate eraser, grasp, wipe until clean), skewering zucchini (grasp slice firmly, push onto stick until punctured), and lightbulb insertion (grasp bulb, align bayonet pin with socket slit, insert while overcoming spring force, rotate to light up). The results are clear. Policies without compliance struggle to modulate contact force and trigger safety faults from excessive force. Policies without force sensing fail to grasp unseen objects or resist reaction forces, causing slippage. Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

12,868 просмотров • 7 месяцев назад