Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

our new model Atlas also happens to be a capable text-to-image generator, providing it with a strong foundation of “world knowledge”... this generalizes across both single and multi view domains, meaning you can step into even highly stylized scenes with precise 3D control

103,214 Aufrufe • vor 10 Tagen •via X (Twitter)

36 Kommentare

Profilbild von Keenan Crane
Keenan Cranevor 10 Tagen

Question purely out of curiosity: can you recover an explicit 3D representation (say, a textured mesh) from Atlas, such that renderings of the explicit representation closely reproduce the video (say, for example “hard surface” models and relatively simple materials)? This kind of question often gets side stepped (e.g., “meshes are old news”), but for me it’s a scientific question rather than an application question: what information is the model predicting, and which variables are still underdetermined?

Profilbild von Thariq
Thariqvor 10 Tagen

Incredible, congrats Ben!

Profilbild von Mazeyar
Mazeyarvor 10 Tagen

Can you let me play GTA VI in Atlas 😌🙏🏼

Profilbild von Ben Mildenhall
Ben Mildenhallvor 10 Tagen

yes

Profilbild von Mazeyar
Mazeyarvor 10 Tagen

Hahaha thanks 🤣

Profilbild von Mathew Tizard
Mathew Tizardvor 10 Tagen

Incredible work. If I have a variety of viewpoints of the same scene, can Atlas ingest all that ground truth and use it to synthesise the missing perspectives? Would this work with synchronised video frames too?

Profilbild von Ben Mildenhall
Ben Mildenhallvor 9 Tagen

exactly!

Profilbild von POM
POMvor 10 Tagen

Beautiful style adherence! The fabric on the octopus is amazing

Profilbild von bone
bonevor 10 Tagen

very cool, can you show it doing a couple 360s? would like to see if it keeps persistence

Profilbild von Ian Curtis
Ian Curtisvor 10 Tagen

this one 🫠 so good

Profilbild von Nicholas Bardy
Nicholas Bardyvor 10 Tagen

nice work

Profilbild von Troy Kirwin
Troy Kirwinvor 10 Tagen

Sick

Profilbild von XGEN Labs
XGEN Labsvor 4 Tagen

The move from image generation to controllable 3D worlds is especially exciting. Curious how far this kind of world knowledge can generalize across unseen environments and viewpoints.

Profilbild von Doug Thompson
Doug Thompsonvor 10 Tagen

I was waiting for this. Marble etc were awesome but I can't wait to experiment. Working on history projects (London UK, really top people) and we have such deep ideas but needed something like Atlas to get us there. Storytelling has a new tool kit and you guys rocking it. EPIC

Profilbild von Yücel İnanoğulları
Yücel İnanoğullarıvor 9 Tagen

Very impressive. I wonder if one day, AI would read a novel and produce a cinematic version of it.

Profilbild von Nicholas H
Nicholas Hvor 10 Tagen

wow that 2D japanese scene though

Profilbild von ritwik
ritwikvor 10 Tagen

Takes one back to NeRF days

Profilbild von Amara Wolfe
Amara Wolfevor 10 Tagen

Claiming Atlas has 'world knowledge' from 2D text-to-image training is a massive stretch. How does it handle actual physical consistency and occlusion when you try to navigate those stylized scenes?

Profilbild von Peachy J
Peachy Jvor 10 Tagen

That world knowledge foundation generalizing across views is clever.

Profilbild von COSMOS
COSMOSvor 9 Tagen

wf

Profilbild von man ki Hui
man ki Huivor 9 Tagen

Great 👍🏻

Profilbild von Jason Farris
Jason Farrisvor 10 Tagen

@venturetwins How do we get started using this?

Profilbild von Pablo Antonio
Pablo Antoniovor 10 Tagen

very cool!

Profilbild von THC Humor 💹🧲
THC Humor 💹🧲vor 9 Tagen

World knowledge plus precise 3D control is a powerful combination

Profilbild von Nathan Quantum
Nathan Quantumvor 9 Tagen

you did more for open research with the original nerf code release than most labs manage in a decade. really hoping atlas follows the same path

Profilbild von Shawn Neal
Shawn Nealvor 9 Tagen

So in the future, video editing will become world editing? like world editor in RTS game?

Profilbild von Neek
Neekvor 10 Tagen

nice!

Profilbild von Manaen
Manaenvor 9 Tagen

The cinematography and real world simulation is unmatched!

Profilbild von 김란지
김란지vor 9 Tagen

😍😍😍😍

Profilbild von Alex AI Home
Alex AI Homevor 9 Tagen

Precise 3D control across single and multi-view domains solves the biggest headache in AI home transformation workflows. When passing raw balcony photos into Atlas, does the underlying world knowledge lock structural dimensions (like railings and glass planes) when pulling dynamic camera orbits?

Profilbild von AI Mastery Guide
AI Mastery Guidevor 10 Tagen

Precise 3D control from a side feature, wow

Profilbild von Logerhaus
Logerhausvor 10 Tagen

Yep were in a simulation

Profilbild von kaen_
kaen_vor 9 Tagen

The 3D consistency here is really impressive. Feels like a meaningful step toward truly controllable world models

Profilbild von Callum Marley
Callum Marleyvor 10 Tagen

Extremely cool, congrats to the team

Profilbild von Owx
Owxvor 9 Tagen

Does that work with anime ?

Profilbild von Ben
Benvor 9 Tagen

when we revisit the same place twice are the details identical or it's being generated differently? How long does the same details persist? I think I can find some fun use cases with this new tool

Ähnliche Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 Aufrufe • vor 1 Jahr