正在加载视频...

视频加载失败

Want to use Depth Anything, but need metric depth rather than relative depth? Thrilled to introduce Prompt Depth Anything, a new paradigm for accurate metric depth estimation with up to 4K resolution. 👉Key Message: Depth foundation models like DA have already internalized rich geometric knowledge of the 3D world...

67,750 次观看 • 1 年前 •via X (Twitter)

9 条评论

Bingyi Kang 的头像
Bingyi Kang1 年前

This project would not have been possible without the wonderful collaboration of @pengsida, Jingxiao Chen, @songyoupeng, @JiamingSuen, @ericliuof97, Hujun Bao, Jiashi Feng and @XiaoweiZhou5.

Bingyi Kang 的头像
Bingyi Kang1 年前

Check out for some interactive demos there:

Bingyi Kang 的头像
Bingyi Kang1 年前

We thank @huggingface for the generous support in serving our demo. @_akhaliq @pcuenq

Patryk Zoltowski 的头像
Patryk Zoltowski1 年前

Any reason why cannot use depth from front Truedepth camera? They are also available in ARKit (with head tracking config) and from my experience higher quality and bigger resolution (640x480) than those from Lidar. However only provided at 30fps.

Basid 的头像
Basid1 年前

exactly what I needed, good work

Redmond Hosting 的头像
Redmond Hosting1 年前

What about using this to upscale depth from a multi-view depth model like Depth Crafter?

Bingyi Kang 的头像
Bingyi Kang1 年前

The same idea can be applied to this setting. But the models released are only for iPhone's Lidar.

Xiaoyu Xiang 的头像
Xiaoyu Xiang1 年前

Wow, a strong baseline to compare in our upcoming paper!

hantmango 的头像
hantmango1 年前

Don't have Iphone....

相关视频

Alibaba presents MIMO Controllable Character Video Synthesis with Spatial Decomposed Modeling Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.

AK

149,079 次观看 • 2 年前