正在加载视频...
视频加载失败
📢GaussianGPT: autoregressive 3D Gaussian scene generation. We introduce a GPT-style model that directly generates 3D Gaussian scenes, token by token, in a series of small, discrete decision steps. Generation, completion, and large-scale outpainting in a single pipeline. Unlike diffusion-based approaches, GaussianGPT explicitly models the scene distribution at every step,... show more
152,626 次观看 • 6 个月前 •via X (Twitter)
31 条评论

I'm not sure "token by token" is the right framing here. Gaussians aren't really discrete units the way words are, so forcing that paradigm seems like it's fighting the representation more than leveraging it.

Very cool work! We followed a quite similar pipeline a couple of year ago to generate scenes from a single-exemplar: In our case, latent features were DINO and Gaussians were decoded with a patch-based synthesis step (inspired by PatchMatch).

this is really interesting! does this mean it builds a scene "meter by meter" (tokens as splats done over distance)? if yes there is tons of usecases for this - eg worldmodels for robotics, interactive games, VR, or simply generation of scenes, etc etc all combined with really fast rendering longer term end-state here could be, a model that does "chain of thought" directly in latent space and allows a worldmodel to reason not in pixels but full environments at absurd speed lmk if understanding this is BS and also kudos @nicolasvluetzow @barbara_roessle @katha_schmid

Yes, pretty much, although tokens are smaller granularity than meters. We also see the same use cases :)

yes - was more meant as figure of speech obvious genuinely - this is the most exciting thing i have seen a long time in that space true 3d worldmodels - not 2d hacks

Incredible

Pretty coooool :-) You think this approach could work with SDF primitives too?

This is sick! 🔥No input photos, no COLMAP, no Meshroom, no classic Structure-from-Motion at generation time.

very cool!

that's sick I was trying to do AR a while back i feel like this would bang

@nicolasvluetzow Is this open source?

We are moving from capturing reality to predicting it. The data moat just got deeper.

how about object by object, LLM recalls the scene that gets generated in its space, whether in gauss splats or text clearly sequential representation here is shown, much like image generation in chatgpt, is it volumetric generation using LLM, but with using gaussian?

this is awesome!

Awesome work! Can’t wait to integrate this into our StudioTwin platform and bring it to the hands of UE5 filmmakers. What’s the performance like for a full scene generation?

dream come true type shi

HOly Ish 😮

@janusch_patas interesting

impressive!

Will you release all trained models? Would tremendously help with adjacent research

so we're basically teaching GPT to think in 3D now? next thing you know it'll be generating entire metaverses token by token while i'm still trying to get my gaussian splats to not look like abstract art gone wrong

love it

Cool

Are you saying this generates a point cloud that doesn’t break at certain camera angles?

does this create splats from images, text or as of now just edit them?

When will be codea? If there will be ability to choose image gen models?

Very cool!

cc @SirWrender

🚀

oMG

any credible application of this could 100x compute demand
