Video wird geladen...
Video konnte nicht geladen werden
our new model Atlas also happens to be a capable text-to-image generator, providing it with a strong foundation of “world knowledge”... this generalizes across both single and multi view domains, meaning you can step into even highly stylized scenes with precise 3D control
103,214 Aufrufe • vor 10 Tagen •via X (Twitter)
36 Kommentare

Question purely out of curiosity: can you recover an explicit 3D representation (say, a textured mesh) from Atlas, such that renderings of the explicit representation closely reproduce the video (say, for example “hard surface” models and relatively simple materials)? This kind of question often gets side stepped (e.g., “meshes are old news”), but for me it’s a scientific question rather than an application question: what information is the model predicting, and which variables are still underdetermined?

Incredible, congrats Ben!

Can you let me play GTA VI in Atlas 😌🙏🏼

yes

Hahaha thanks 🤣

Incredible work. If I have a variety of viewpoints of the same scene, can Atlas ingest all that ground truth and use it to synthesise the missing perspectives? Would this work with synchronised video frames too?

exactly!

Beautiful style adherence! The fabric on the octopus is amazing

very cool, can you show it doing a couple 360s? would like to see if it keeps persistence

this one 🫠 so good

nice work

Sick

The move from image generation to controllable 3D worlds is especially exciting. Curious how far this kind of world knowledge can generalize across unseen environments and viewpoints.

I was waiting for this. Marble etc were awesome but I can't wait to experiment. Working on history projects (London UK, really top people) and we have such deep ideas but needed something like Atlas to get us there. Storytelling has a new tool kit and you guys rocking it. EPIC

Very impressive. I wonder if one day, AI would read a novel and produce a cinematic version of it.

wow that 2D japanese scene though

Takes one back to NeRF days

Claiming Atlas has 'world knowledge' from 2D text-to-image training is a massive stretch. How does it handle actual physical consistency and occlusion when you try to navigate those stylized scenes?

That world knowledge foundation generalizing across views is clever.

wf

Great 👍🏻

@venturetwins How do we get started using this?

very cool!

World knowledge plus precise 3D control is a powerful combination

you did more for open research with the original nerf code release than most labs manage in a decade. really hoping atlas follows the same path

So in the future, video editing will become world editing? like world editor in RTS game?

nice!

The cinematography and real world simulation is unmatched!

😍😍😍😍

Precise 3D control across single and multi-view domains solves the biggest headache in AI home transformation workflows. When passing raw balcony photos into Atlas, does the underlying world knowledge lock structural dimensions (like railings and glass planes) when pulling dynamic camera orbits?

Precise 3D control from a side feature, wow

Yep were in a simulation

The 3D consistency here is really impressive. Feels like a meaningful step toward truly controllable world models

Extremely cool, congrats to the team

Does that work with anime ?

when we revisit the same place twice are the details identical or it's being generated differently? How long does the same details persist? I think I can find some fun use cases with this new tool

