Loading video...

Video Failed to Load

Go Home

Presenting MonoNeRF at #ICML2023 We train a generalizable NeRF from: ✅Large-scale monocular videos instead of one scene ✅No GT camera poses.📷🚫 Without per-scene optimization, the model can do view synthesis, depth estimation, camera pose estimation.

36,578 views • 3 years ago •via X (Twitter)

8 Comments

Xiaolong Wang's profile picture
Xiaolong Wang3 years ago

This work is extending from our previous work on Video Autoencoder, but a NeRF version. We firmly believe in the potential of learning 3D geometry from videos without the constraint of camera poses. This is the way to scale up! 2/n

Xiaolong Wang's profile picture
Xiaolong Wang3 years ago

Even trained without camera poses, it can still be used for camera pose estimation. 3/n

Xiaolong Wang's profile picture
Xiaolong Wang3 years ago

This is a joint effort with my student Yang Fu (@yangfu21) and old friend Ishan Misra (@imisra_). Looking forward to catching up in ICML. arxiv: 4/n

Yue Wang's profile picture
Yue Wang3 years ago

@JiaweiYang118

69420's profile picture
694203 years ago

Can it render in real time?

Jiatao Gu's profile picture
Jiatao Gu3 years ago

Amazing work!! Will you release the code & pretrained models soon?

Xiaolong Wang's profile picture
Xiaolong Wang3 years ago

Yes. Very soon I think! @yangfu21

Yuliang Zou's profile picture
Yuliang Zou3 years ago

Nice work! Not sure if I miss something, I did not find how to set d_i adaptively for each image and how to convert depth encoder features to this multiple nerf representation.

Related Videos