正在加载视频...

视频加载失败

VideoJAM is our new framework for improved motion generation from AI at Meta We show that video generators struggle with motion because the training objective favors appearance over dynamics. VideoJAM directly adresses this **without any extra data or scaling** 👇🧵

168,926 次观看 • 1 年前 •via X (Twitter)

11 条评论

Hila Chefer 的头像
Hila Chefer1 年前

Why do video generators struggle with motion? We found that the pixel-based loss barely changes when video frames are shuffled—showing it is nearly **invariant to temporal incoherence**. This leads models to ignore motion and prioritize appearance

Hila Chefer 的头像
Hila Chefer1 年前

Our Solution: VideoJAM VideoJAM instills an explicit motion prior by modifying the objective: the model predicts both appearance and motion from a **single learned representation.** This forces the model to capture both visuals and dynamics, improving motion understanding.

Hila Chefer 的头像
Hila Chefer1 年前

Inner-Guidance: Improving Motion at Inference At inference, we introduce **Inner-Guidance**—a method that leverages the **model’s own motion predictions** as a dynamic guidance signal, steering the generation toward coherent, realistic motion.

Hila Chefer 的头像
Hila Chefer1 年前

🎬 Results VideoJAM fine-tunes a pretrained video generator (DiT) on just 3M samples from its own training set—yet achieves remarkable motion coherence. It even outperforms highly competitive proprietary models like Sora and Kling in motion quality

Hila Chefer 的头像
Hila Chefer1 年前

This work was done during my internship at @AIatMeta 🎉 Huge thanks to my amazing collaborators @urielsinger @amit_zhr @YKirstain @adam_polyak90 Yaniv Taigman @liorwolf and @ShellySheynin Check out the project page for many more results and details:

Hila Chefer 的头像
Hila Chefer1 年前

Now on Huggingface daily papers 🤗 And arxiv 🥳

Akool Inc 的头像
Akool Inc1 年前

Need AI avatars or voiceovers? HeyGen and AKOOL both offer powerful AI video tools, but AKOOL delivers superior results with more precise body movements and lip sync. See the difference!

Alex Nasa 的头像
Alex Nasa1 年前

@AIatMeta A motion prediction auxiliary model make so much sense, did you run it on this benchmark by any chance?

Hila Chefer 的头像
Hila Chefer1 年前

@AIatMeta Interesting, thanks so much for the reference! I actually think physics is an interesting issue that is far (far) from being solved. Our representation is motion-based but I think it’ll be so cool if we could add an actual physics-driven representation to the mix 🙏🙏

Everett World 的头像
Everett World1 年前

@AIatMeta That's super impressive - can't wait to generate with it!

Hila Chefer 的头像
Hila Chefer1 年前

@AIatMeta Thanks Everett! 🫶

相关视频

Google presents Still-Moving Customized Video Generation without Customized Video Data Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a text-to-image (T2I) model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth or StyleDrop). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.

AK

40,485 次观看 • 2 年前