Загрузка видео...

Не удалось загрузить видео

На главную

Depth Anything V2 This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much finer and more

114,537 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля AK
AK2 лет назад

robust depth predictions through three key practices: 1) replacing all labeled real images with synthetic images, 2) scaling up the capacity of our teacher model, and 3) teaching student models via the bridge of large-scale pseudo-labeled real images. Compared with the

Фото профиля AK
AK2 лет назад

latest models built on Stable Diffusion, our models are significantly more efficient (more than 10x faster) and more accurate. We offer models of different scales (ranging from 25M to 1.3B params) to support extensive scenarios. Benefiting from their strong generalization

Фото профиля AK
AK2 лет назад

capability, we fine-tune them with metric depth labels to obtain our metric depth models. In addition to our models, considering the limited diversity and frequent noise in current test

Фото профиля AK
AK2 лет назад

sets, we construct a versatile evaluation benchmark with precise annotations and diverse scenes to facilitate future research.

Фото профиля AK
AK2 лет назад

paper page:

Фото профиля AK
AK2 лет назад

daily papers:

Фото профиля Tremeschin 🔱
Tremeschin 🔱2 лет назад

Still waiting for the transformers pipeline instead of a git clone huggingface repo install in my project 😅 Results are awesome for DAv2 !!

Фото профиля Amir Laylaz
Amir Laylaz2 лет назад

@oleg__chomp 👀👀

Фото профиля JohnYue122333
JohnYue1223332 лет назад

great

Фото профиля Sivan R. Hokayma
Sivan R. Hokayma2 лет назад

Holy shit this is awesome. Are the rgb values directly proportional to the distance to the camera or are they relative to other elements within the scene? i.e. will anything 1 ft away from the camera always be the same shade of red across different scenes?

Похожие видео

In collaboration with Intel, our Depth Fusion showcases the power of our LDM3D diffusion model in generating 360° views from text prompts provided by the user. The LDM3D diffusion model generates a 2D RGB image and its corresponding relative depth map providing a complete RGBD representation corresponding to the text prompt. The LDM 3D model is a specialized version of the stable diffusion V 1.4 model that has been modified to fit both image and depth map data.The model was then fine tuned on a subset of the Laion400M data set - large scale image caption data set. The depth maps used to fine tune our model were generated by the DPTBeiT large 512 depth estimation model that provides highly accurate relative depth estimates for each pixel. We take the generated 2D RGB image and depth map and use them to compute a 360° projection using touchdesigner. Touchdesigner is a versatile platform that allows for the creation of immersive and interactive multimedia experiences. Our application harnesses the power of touchdesigner to bring the generated 360° views to life, providing users with a unique and engaging way to experience their text prompts, whether it’s a description of a tranquil forest, a noisy cityscape or a futuristic sci fi world. Our depth fusion can bring these concepts to life in a vivid and immersive detail. - Scottie Fox, VP Engineering Blockade Labs ScottieFox #AI #VR #3D #gamedev #stablediffusion

Blockade Labs

11,439 просмотров • 3 лет назад