ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

๐Ÿ“ข๐Ÿ“ข๐Ÿ“ข ๐€๐‚๐Ÿ‘๐ƒ: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers TL;DR: for 3D camera control in generative video, it really helps knowing *which* part of your model you should mess with Internship by Sherwin Bahmani at Snap

23,040 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰ โ€ขvia X (Twitter)

5 ๆก่ฏ„่ฎบ

Andrea Tagliasacchi ๐Ÿ‡จ๐Ÿ‡ฆ ็š„ๅคดๅƒ
Andrea Tagliasacchi ๐Ÿ‡จ๐Ÿ‡ฆ1 ๅนดๅ‰

TL;DR (expanded): 1) "when" in the diffusion process you condition for camera matters (i.e. noise scheduler) 2) "how" in the diffusion process you condition for camera maters (i.e. architecture) 3) "what data" you give to your diffusion model to condition camera matters

Andrea Tagliasacchi ๐Ÿ‡จ๐Ÿ‡ฆ ็š„ๅคดๅƒ
Andrea Tagliasacchi ๐Ÿ‡จ๐Ÿ‡ฆ1 ๅนดๅ‰

Why, you ask? 1) camera motion is low-frequency... early denoising iterations deal with low-frequency content 2) early DiT blocks are enough to fine-tune for camera control... more and you lose quality 3) model needs to know what a static view of the dynamic world looks like

Andrea Tagliasacchi ๐Ÿ‡จ๐Ÿ‡ฆ ็š„ๅคดๅƒ
Andrea Tagliasacchi ๐Ÿ‡จ๐Ÿ‡ฆ1 ๅนดๅ‰

A shout to the collaborators @isskoro @guocheng_qian A. Siarohin @willimenapace @SergeyTulyakov at Snap and @DaveLindell at UofT.

Samarth Sinha ็š„ๅคดๅƒ
Samarth Sinha1 ๅนดๅ‰

@sherwinbahmani Congrats @sherwinbahmani !!

Abdullah Hamdi ็š„ๅคดๅƒ
Abdullah Hamdi1 ๅนดๅ‰

@sherwinbahmani Congrats to the team ! Amazing work

็›ธๅ…ณ่ง†้ข‘

๐Ÿ“ข๐Ÿ“ข ๐๐ž๐ซ๐œ๐‡๐ž๐š๐: ๐๐ž๐ซ๐œ๐ž๐ฉ๐ญ๐ฎ๐š๐ฅ ๐‡๐ž๐š๐ ๐Œ๐จ๐๐ž๐ฅ ๐Ÿ๐จ๐ซ ๐’๐ข๐ง๐ ๐ฅ๐ž-๐ˆ๐ฆ๐š๐ ๐ž ๐Ÿ‘๐ƒ ๐‡๐ž๐š๐ ๐‘๐ž๐œ๐จ๐ง๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ข๐จ๐ง & ๐„๐๐ข๐ญ๐ข๐ง๐ ๐Ÿ“ข๐Ÿ“ข PercHead reconstructs realistic 3D heads from a single image and enables disentangled 3D editing via geometric controls and style inputs from images or text. At its core is a generalized 3D head decoder trained with perceptual supervision from DINOv2 and SAM 2.1. We find that our new perceptual loss formulation improves reconstruction fidelity compared to commonly-used methods such as LPIPS. Our trained reconstruction model is able to generate 3D-consistent heads from a single input image. Even with challenging side-view inputs, the model robustly infers missing regions for a coherent, high-fidelity output. In addition, our architecture seamlessly adapts to downstream tasks: by swapping the encoder, we can transform the model into a disentangled 3D editing pipeline. In this scenario, we can control geometry through - potentially hand-drawn - segmentation maps, and condition style via image or text prompt. We also provide an interactive GUI to enable the exploration of our editing pipeline. ๐ŸŒ ๐Ÿ“ฝ๏ธ Great work by Antonio Oroz and Tobias Kirschstein

Matthias Niessner

18,855 ๆฌก่ง‚็œ‹ โ€ข 8 ไธชๆœˆๅ‰