正在加载视频...

视频加载失败

Presenting MonoNeRF at #ICML2023 We train a generalizable NeRF from: ✅Large-scale monocular videos instead of one scene ✅No GT camera poses.📷🚫 Without per-scene optimization, the model can do view synthesis, depth estimation, camera pose estimation.

36,578 次观看 • 3 年前 •via X (Twitter)

8 条评论

Xiaolong Wang 的头像
Xiaolong Wang3 年前

This work is extending from our previous work on Video Autoencoder, but a NeRF version. We firmly believe in the potential of learning 3D geometry from videos without the constraint of camera poses. This is the way to scale up! 2/n

Xiaolong Wang 的头像
Xiaolong Wang3 年前

Even trained without camera poses, it can still be used for camera pose estimation. 3/n

Xiaolong Wang 的头像
Xiaolong Wang3 年前

This is a joint effort with my student Yang Fu (@yangfu21) and old friend Ishan Misra (@imisra_). Looking forward to catching up in ICML. arxiv: 4/n

Yue Wang 的头像
Yue Wang3 年前

@JiaweiYang118

69420 的头像
694203 年前

Can it render in real time?

Jiatao Gu 的头像
Jiatao Gu3 年前

Amazing work!! Will you release the code & pretrained models soon?

Xiaolong Wang 的头像
Xiaolong Wang3 年前

Yes. Very soon I think! @yangfu21

Yuliang Zou 的头像
Yuliang Zou3 年前

Nice work! Not sure if I miss something, I did not find how to set d_i adaptively for each image and how to convert depth encoder features to this multiple nerf representation.

相关视频