Loading video...
Video Failed to Load
Christian Rupprecht explains their interpretability research in 3D computer vision, testing if (and where in the model) multi-view transformers like VGGT, DepthAnything 3, and DUSt3R use point/patch correspondences to make sense of 3D scene geometry.
74,761 views • 6 months ago •via X (Twitter)
9 Comments

Full talk:

Paper:

Q: Are these truly geometric correspondences or just the semantic correspondences we get from the DINO encoder, or even from diffusion models like Stable Diffusion? A: Early layers are good at semantic correspondences but bad at geometric ones, and vice versa for later layers.

Related:

Can we extract these correspondences from VGGT then run classic bundle adjustment?

Possibly, but why would you want to do that? I'd expect them to be much less precise than, e.g., RoMA 2, not least because they're only patch-level (not pixel-level) correspondences.

Im thinking they're less precise too, but maybe more robust

Interesting!

That's a very good analysis to verify if models are truly able to learn epipolar geometry for monocular VSLAM. Every matching/depth estimation model that I have tested, produce few epipolar correspondences until explicitly trained to follow geometry.
