Загрузка видео...

Не удалось загрузить видео

На главную

Variable-length compressive tokenization is a promising direction. Not just for efficiency, but for actually learning powerful representations. FlexTok scratched the surface of this direction with images. But real-world data has additional structure that can be tapped, such as the temporal structure of videos. VideoFlexTok develops this concept for video,...

17,775 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 1

Фото профиля Sivan Doveh
Sivan Doveh4 месяцев назад

Great paper Amir! Super interesting

Похожие видео

Terence Tao, UCLA professor and the most decorated mathematician alive: "Funds pay $750K to combine weak signals into one real edge. I proved the thing that makes it work and makes it dangerous: in any long enough sequence, hidden structure is unavoidable. it always accumulates. the whole job is telling the real structure from the noise that only looks like it." this free lecture is the most decorated mathematician alive on the exact problem sitting underneath every factor model, and it costs nothing. at the board it's simple. Tao's lifelong theme is the line between structure and randomness. The Erdős discrepancy problem asks a deceptively simple thing: can you write an endless string of plus-ones and minus-ones that stays perfectly balanced forever? Tao proved you cannot. No matter how cleverly you try, imbalance, hidden structure, is forced to accumulate as the sequence grows. There is no such thing as a long stream of pure, structureless noise. That's the whole idea, minus the jargon. Which is exactly why a multi-factor model can work, and exactly why it can kill you. Stack enough weak signals and real structure will appear, because at scale structure is unavoidable. But so will fake structure, patterns that exist only because the data is long enough to force them. Same point as the post above: finding structure is guaranteed. Knowing which structure is an edge is the rare and expensive part. the mathematics is free and public. what nobody can sell you is the judgment to tell the structure the market will pay you for from the structure that exists only because you looked hard enough. That judgment is the alpha, and it takes years to build.

Rossst.03

213,931 просмотров • 1 месяц назад

Google presents Still-Moving Customized Video Generation without Customized Video Data Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a text-to-image (T2I) model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth or StyleDrop). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.

AK

40,489 просмотров • 2 лет назад