Video wird geladen...
Video konnte nicht geladen werden
Novel view synthesis has long been a core challenge in 3D vision. But how much 3D inductive bias is truly needed? —Surprisingly, very little! Introducing "LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias"—a fully transformer-based approach that enables scalable, generalizable, and fully data-driven novel view synthesis,... show more
114,947 Aufrufe • vor 1 Jahr •via X (Twitter)
26 Kommentare

(2/6) We introduce two architectures: (1) an encoder-decoder LVSM, and (2) a decoder-only LVSM. Both models bypass the 3D inductive biases used in previous methods—from 3D representations (e.g., NeRF, 3DGS) to network designs (e.g., epipolar projections, plane sweeps)—addressing novel view synthesis with a fully data-driven approach.

(3/6) Our methods can work for both objects and scenes from sparse inputs (2-4 views), and our best models outperform previous state-of-the-art methods by 1.5 to 3.5 dB PSNR.

(4/6) We observe that our LVSM also works with a single input view for many cases, despite only being trained with multi-view inputs. This observation shows the capability of LVSM to understand the 3D world, e.g. understanding depth, rather than just performing pixel-level view interpolation.

(5/6) Notably, our models surpass all previous methods even with reduced computational resources (1-2 GPUs).

(6/6) Thanks for my wonderful collaborators's support and help! @hanwenjiang1 @HaoTan5 @KaiZhang9546 @Sai__Bi @tianyuanzhang99 @fujun_luan @Jimantha @zexiangxu

very surprised with the extremely good lpips results, congrats on making a huge move in fast 3d reconstruction!

Thank you, Junlin!

This is super cool, I've been waiting for this to be invented! For a few years now, I've been taking certain photos with a couple of additional shots from different angles, knowing that some day I would be able to make 3D views! The world is catching up to me! Cool!

Very cool!

Thank you, Sida!

Pretty cool - congrats!!!

Thank you!

I can’t wait to play in your Holodeck.

This is super cool!

It is really a great work! I have some question about it. 1. How much GPU memory is required for inference? 2. What is the inference speed? Is it faster than stable video diffusion?

So powerful! 🔥

What do you need? Just a few pictures or pictures with accurate camera transformations?

what about dust3r?

Great work! How is the visual quality so good without using diffusion models? Especially when you zoom into details.

Cool work !

🔥🔥🚀

Do you think 3D Inductive Biases are unnecessary (in the context of NVS) Or do you think we don’t have enough 3D data / havent found a good representation yet?

wow!! really cool.. how sensitive is the method to having noisy poses?

Congrats! I like the capacity of the model to have good results with Single Image even if the model is trained for Multiple. And any info about the inference speed and code release date?

Wen app?

@theworldlabs Since many of us are waiting for so long and most of us are not really going to get access anytime soon you guys might as well look into this approach too
