Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

When we say Atlas has pixel-perfect camera control, we mean it. Atlas precisely follows input camera parameters, including non-planar projections such as the Brown-Conrady distortion model and the Kannala-Brandt fish-eye model. Panini Bhamidipati invented a novel method of camera conditioning and it works beautifully. 🧵 [1/N]

92,777 Aufrufe • vor 15 Tagen •via X (Twitter)

24 Kommentare

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

Pinhole cameras are a given -- Atlas can perform the iconic dolly zoom effortlessly.

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

Non-planar projects are also no problem. Here's a fish-eye projection with the Kannala-Brandt model.

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

And pincushion, barrel, and tangential distortion under the Brown-Conrady model.

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

Read more about Atlas: Request early access:

Profilbild von Panini Bhamidipati
Panini Bhamidipativor 15 Tagen

Thank you! It was a very fun problem to work on with this incredible team.

Profilbild von Bilawal Sidhu
Bilawal Sidhuvor 15 Tagen

@ychngji6 @BhamidipatiPan1 Interesting! So can yall reconstruct a scene given a sparse set of say 8mm fish eye images taken on a full frame sensor?

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

@ychngji6 @BhamidipatiPan1 8mm might stretch the current model a bit, but in principle yes.

Profilbild von Bilawal Sidhu
Bilawal Sidhuvor 15 Tagen

@ychngji6 @BhamidipatiPan1 Very cool

Profilbild von saietta
saiettavor 15 Tagen

following the camera parameters is the easy half, the harder one is whether the reconstructed geometry stays put once you move past a clean single-subject shot. cluttered scenes with occlusion and reflective surfaces are where world models usually drift, curious if that holds at this precision

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

@BhamidipatiPan1 Does this qualify?

Profilbild von saietta
saiettavor 14 Tagen

@BhamidipatiPan1 Good stress test, plenty of clutter and occlusion in there. Curious if it still holds on the reflective tool surfaces, that's usually where the drift shows up first, not the diffuse stuff.

Profilbild von Xi WANG
Xi WANGvor 15 Tagen

@BhamidipatiPan1 Super nice job! Thanks for sharing! Have you consider adding more characteristics like bokeh / aperture and shutter speed (motion blur), which are also very physical and optical.

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

@BhamidipatiPan1 Not yet, but good idea!

Profilbild von Dre
Drevor 10 Tagen

@BhamidipatiPan1 Very cool. On uncalibrated input, ballpark how tight are the recovered intrinsics and distortion vs a ChArUco calibration? Sub-pixel reprojection, or more like a few px?

Profilbild von Preyforge
Preyforgevor 15 Tagen

@BhamidipatiPan1 I trust camera control after I check the generated sequence, not the still; continuity has to survive the path users actually hit.

Profilbild von BarkinBoss
BarkinBossvor 14 Tagen

@BhamidipatiPan1 Atlas pulling off pixel‑perfect camera control is wild engineering. When you can model distortion this precisely, you’re not just simulating reality, you’re bending it. Thoughts?

Profilbild von saietta
saiettavor 14 Tagen

@BhamidipatiPan1 Good stress test, plenty of clutter and occlusion in there. Curious if it still holds on the reflective tool surfaces, that's usually where the drift shows up first, not the diffuse stuff.

Profilbild von Derek Austin
Derek Austinvor 15 Tagen

@BhamidipatiPan1 What units are used in Atlas? One thing I hate in blender is figuring out how large to make something compared to my scene

Profilbild von Keunhong Park
Keunhong Parkvor 15 Tagen

@BhamidipatiPan1 Whatever units you want. It's grounded to your inputs

Profilbild von Derek Austin
Derek Austinvor 15 Tagen

@BhamidipatiPan1 I am loving all the demo twitter videos - anyway I could get access?

Profilbild von Alex Nichol
Alex Nicholvor 15 Tagen

@BhamidipatiPan1 Ray embeddings?

Profilbild von Ryan Edkins
Ryan Edkinsvor 15 Tagen

@BhamidipatiPan1 Most camera-conditioned models fake it with a pinhole assumption, so bothering to nail actual fisheye distortion is a genuine technical achievement.

Profilbild von Sagar
Sagarvor 15 Tagen

@BhamidipatiPan1 Camera control is what makes it usable for real shots. Where does it still drift, long moves or fast ones?

Profilbild von Si(super intelligence)
Si(super intelligence)vor 15 Tagen

@BhamidipatiPan1 👍👍

Ähnliche Videos

Mistral AI Releases Robostral Navigate: An 8B Model Enabling Robots to Navigate Complex Environments Hitting 76.6% on R2R-CE With One RGB Camera. No LiDAR. No depth sensor. No multi-camera rig. Here's how it works. 👇 1. Pointing, not metric commands The model predicts the pixel coordinates of the next target in the camera view, plus the arrival orientation. Working in pixel space keeps it robust to camera intrinsics and world scale. When the target leaves the frame, it falls back to local displacements ("2m forward, 1.5m left, turn 25°"). 2. Grounding-first No open-source VLM base. It starts from Mistral's grounding model (pointing, counting, localization). Navigation emerges once the model knows where things are. → ~400,000 trajectories across 6,000 simulated scenes 3. Prefix-caching for training A tree-based attention mask packs a full episode into one sequence — all time steps in a single forward pass. → 22× fewer training tokens; months of training done in days 4. Online RL on top After supervised training, CISPO adds trial-and-error learning to fight distribution shift from behavior cloning. → +3.2% success rate from RL alone 5. The numbers (R2R-CE, Matterport3D) → 76.6% success on validation unseen → +9.7 pts over best single-camera approach → +4.5 pts over best depth/multi-camera system The key takeaway: state-of-the-art continuous VLN without a sensor stack — grounding-init, pixel-space actions, prefix-cached SFT, and online RL, on one RGB camera. Full analysis: Technical details: Mistral AI Mistral AI for Developers

Marktechpost AI

39,955 Aufrufe • vor 2 Monaten