Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I'm excited to share our new work Align3R that estimates camera poses and consistent depth maps from a monocular video of a dynamic scene. Project page: Code: Paper:

56,547 Aufrufe • vor 1 Jahr •via X (Twitter)

9 Kommentare

Profilbild von Yuan Liu
Yuan Liuvor 1 Jahr

Our work is motivated by aligning estimated single-view depth maps with DUSt3R. To achieve this, we fine-tuned DUSt3R on 3D scenes with additional monocular depth maps as inputs.

Profilbild von Yuan Liu
Yuan Liuvor 1 Jahr

There is another excellent work MonST3R ( that also finetunes DUSt3R on dynamic scenes. We have borrowed the idea of flow losses from MonST3R. We thank the authors for releasing their excellent work.

Profilbild von Hector
Hectorvor 1 Jahr

amazing work

Profilbild von Adarsh Baghel
Adarsh Baghelvor 1 Jahr

isn't there something to fill in the white gaps?

Profilbild von Yuan Liu
Yuan Liuvor 1 Jahr

It is possible to do this with some diffusion models like CAT4D ( This problem is not well-studied yet.

Profilbild von Jiaqi Gu
Jiaqi Guvor 1 Jahr

Wondering, if the camera's movement only includes rotation but not translation. what are the results will be.

Profilbild von Yuan Liu
Yuan Liuvor 1 Jahr

Thank you! It can work even if the camera is static so I assume that it can handle rotated cameras without camera motion because the predicted 3D point maps provide some cues to solve pure rotations.

Profilbild von Ai Peasant
Ai Peasantvor 1 Jahr

So I need to upload a every frame? Not sure how it works.

Profilbild von Yuan Liu
Yuan Liuvor 1 Jahr

Yes. The demo only infers a few frames due to the limited GPU resources. Using the GitHub codes will be more convenient if we want to predict a whole video.

Ähnliche Videos