Video yükleniyor...
Video Yüklenemedi
Introducing Ctrl-VI, a video sampling method allowing for a flexible set of user controls—ranging from coarse but easy-to-specify text prompts to precise camera/object trajectories. (1/n)
43,208 görüntüleme • 11 ay önce •via X (Twitter)
7 Yorum

To support control flexibility while maintaining output diversity, we take a variational inference approach to sample from a target distribution that accounts for all input constraints, leveraging a collection of pre-trained diffusion/flow video generation models. (2/n)

High-dimensionality in video data leads to difficulties in convergence with classical sampling/steering methods and weight degeneracy with importance sampling. We break down the sampling problem into easier subproblems of KL divergence minimization, … (3/n)

… and introduce conditional variables that serve as 3D-aware contexts (inferred via pre-trained novel view synthesis models), which reduce modes in the solution space such that particles are less likely to be trapped at local optima. (4/n)

The framework enables other applications such as video object insertion. More results on the website: Work done with my awesome collaborators: @HaoyiDuan (co-lead), @du_yilun, and @jiajunwu_cs. (5/n)

@du_yilun @jiajunwu_cs More generally speaking, this work is an example treating pre-trained models, in the form of p(x) or p(x|y) with fixed-modality conditional y, as proxies for black-box data distributions. As new online or task-specific data come in, … (6/n)

@du_yilun @jiajunwu_cs (cont’) there is a large design space to continuously learn model routing/composition to obtain p(x|y’) to incorporate extra conditioning signals y’ or adapt to test-time need. I’d love to chat more if you are excited about this direction! (n/n)

Congrats! Very impressive work.
