Загрузка видео...

Не удалось загрузить видео

На главную

Monocular depth estimation is “impossible” because one image can’t measure depth geometrically. Our iDisc #CVPR2023 can group pixels w/o supervision and learn depth inductive bias on groups. We get LiDAR-like (but denser) depth from single images! More:

52,267 просмотров • 3 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля AntiFisherYu001
AntiFisherYu0012 лет назад

You are a murderer

Фото профиля Adrian
Adrian3 лет назад

@ValueAnalyst1 @a_meta4 check it out

Фото профиля MinotaurOnLucy
MinotaurOnLucy3 лет назад

As an outsider, I have the following questions: What is the current state-of-the-art performance for out of distribution cases? If it is not good, will we see a foundational model that achieves good out of distribution performance in the near term, or would it be more mid term?

Фото профиля julien michot
julien michot3 лет назад

Why "impossible"? Here you get the most plausible depths.

Фото профиля Nikolaos Sarafianos
Nikolaos Sarafianos3 лет назад

This is great work! Did you test it on humans at various distances from the camera by any chance?

Фото профиля Fisher Yu
Fisher Yu3 лет назад

We test the method on street scenes that have people. However, it is indeed interesting to see whether we can use the method to estimate depth in human-centric scenes.

Фото профиля ζ Pedram ζ
ζ Pedram ζ3 лет назад

Neat idea, github?

Фото профиля Fisher Yu
Fisher Yu3 лет назад

github link: The full code will be released before CVPR 2023.

Фото профиля test bot
test bot3 лет назад

@Scobleizer 69

Фото профиля 𝘃𝗿𝗹𝗹𝗿𝘃
𝘃𝗿𝗹𝗹𝗿𝘃3 лет назад

@Scobleizer Great work!

Похожие видео

In collaboration with Intel, our Depth Fusion showcases the power of our LDM3D diffusion model in generating 360° views from text prompts provided by the user. The LDM3D diffusion model generates a 2D RGB image and its corresponding relative depth map providing a complete RGBD representation corresponding to the text prompt. The LDM 3D model is a specialized version of the stable diffusion V 1.4 model that has been modified to fit both image and depth map data.The model was then fine tuned on a subset of the Laion400M data set - large scale image caption data set. The depth maps used to fine tune our model were generated by the DPTBeiT large 512 depth estimation model that provides highly accurate relative depth estimates for each pixel. We take the generated 2D RGB image and depth map and use them to compute a 360° projection using touchdesigner. Touchdesigner is a versatile platform that allows for the creation of immersive and interactive multimedia experiences. Our application harnesses the power of touchdesigner to bring the generated 360° views to life, providing users with a unique and engaging way to experience their text prompts, whether it’s a description of a tranquil forest, a noisy cityscape or a futuristic sci fi world. Our depth fusion can bring these concepts to life in a vivid and immersive detail. - Scottie Fox, VP Engineering Blockade Labs ScottieFox #AI #VR #3D #gamedev #stablediffusion

Blockade Labs

11,439 просмотров • 3 лет назад