Загрузка видео...

Не удалось загрузить видео

На главную

Video might be the next intelligence substrate. Strikingly, video models are beginning to exhibit the same emergent reasoning behaviors first observed in LLMs—multi-path search, self-correction, and layer specialization. We demystify video reasoning and show it doesn’t happen frame-by-frame, but along diffusion steps. 🔗 📄 So, what's next? ;)

71,036 просмотров • 6 месяцев назад •via X (Twitter)

Комментарии: 15

Фото профиля Xun Huang
Xun Huang6 месяцев назад

It is well known that x0 prediction in early timestep represents the weighted average of all possible outcomes, but I feel it’s misleading to call it “reasoning”. How is “reasoning” defined here?

Фото профиля Zhongang Cai
Zhongang Cai6 месяцев назад

Great point — we don’t consider the presence of multiple possible outcomes at the start to be reasoning. Rather, we view reasoning as the subsequent selection process: early steps maintain multiple candidates, and later steps progressively refine and prune them toward a consistent solution. In that sense, it is analogous to latent reasoning in LLMs—except that here the process is more directly observable through the diffusion trajectory.

Фото профиля mrkelly
mrkelly6 месяцев назад

Video reasoning feels like magic until you see the compute bill.

Фото профиля Zhongang Cai
Zhongang Cai6 месяцев назад

Don’t worry, one year ago I thought video models could never be accelerated to run in real time. Now guess what ;)

Фото профиля 0xmusashi
0xmusashi6 месяцев назад

@kellypeilinchan

Фото профиля Zhongang Cai
Zhongang Cai6 месяцев назад

@kellypeilinchan This is already happening ;)

Фото профиля Thomas
Thomas6 месяцев назад

Interesting

Фото профиля John 🔮🌳 🌎☀️
John 🔮🌳 🌎☀️6 месяцев назад

Who knew “compression” could turn into AGI haha

Фото профиля Critter
Critter6 месяцев назад

A picture is worth a thousand words

Фото профиля Zhongang Cai
Zhongang Cai6 месяцев назад

True, and I wonder if words are necessary at all in many visual reasoning scenarios

Фото профиля Mohamed
Mohamed6 месяцев назад

@CritterObserves @caizhongang for video QA (even for very long videos), we proved here that they can be done end to end on visuals losslessly (without intermediate captioning at all) + logarithmic compute:

Фото профиля Zhongang Cai
Zhongang Cai6 месяцев назад

@CritterObserves Interesting work!

Фото профиля Yuanzhi
Yuanzhi5 месяцев назад

This seems to be like a natural result of bidirectional attention and diffusion; would like to know the observation on AR video gen model, e.g., diffusion forcing?

Фото профиля Zhongang Cai
Zhongang Cai5 месяцев назад

Great idea!

Фото профиля Wang Ruisi
Wang Ruisi6 месяцев назад

Our GitHub repository is now live for discussions and upcoming tool releases. Follow us for the latest updates! GitHub:

Похожие видео

AI has transformed how video is created. We think the next wave is about understanding it. Over the past few years, we've seen remarkable advances in video generation, editing, avatars, and creative tooling. An increasingly important problem is teaching machines to search, analyze, reason over, and extract insight from video - across massive libraries and live streams alike. We're calling this video intelligence, and we're actively looking to back founders building here. We're most excited about companies pushing on the core capabilities: - Video-native models - multimodal embeddings, temporal reasoning, and retrieval built specifically for video rather than adapted from image or text - Real-time and large-scale pipelines - infrastructure for processing, indexing, and querying video at the speed and scale enterprises actually need - Agentic and reasoning layers - systems that don't just retrieve clips but answer questions, surface anomalies, and take action on what they see The models and infrastructure to make this real are appearing to be crossing a capability threshold right now. Multimodal foundation models are maturing, storage costs have collapsed, and enterprises are sitting on years of unstructured video with no way to use it. That infrastructure unlocks a wide range of applications including media and sports workflows, security and physical operations, enterprise knowledge management, advertising analytics, robotics, and consumer products, where video has historically been dark data. If you're building in video intelligence at the model layer, the platform layer, or in a vertical application, we'd love to talk!

Jason Cui

36,205 просмотров • 4 месяцев назад