Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Video might be the next intelligence substrate. Strikingly, video models are beginning to exhibit the same emergent reasoning behaviors first observed in LLMs—multi-path search, self-correction, and layer specialization. We demystify video reasoning and show it doesn’t happen frame-by-frame, but along diffusion steps. 🔗 📄 So, what's next? ;)

71,036 Aufrufe • vor 6 Monaten •via X (Twitter)

15 Kommentare

Profilbild von Xun Huang
Xun Huangvor 6 Monaten

It is well known that x0 prediction in early timestep represents the weighted average of all possible outcomes, but I feel it’s misleading to call it “reasoning”. How is “reasoning” defined here?

Profilbild von Zhongang Cai
Zhongang Caivor 6 Monaten

Great point — we don’t consider the presence of multiple possible outcomes at the start to be reasoning. Rather, we view reasoning as the subsequent selection process: early steps maintain multiple candidates, and later steps progressively refine and prune them toward a consistent solution. In that sense, it is analogous to latent reasoning in LLMs—except that here the process is more directly observable through the diffusion trajectory.

Profilbild von mrkelly
mrkellyvor 6 Monaten

Video reasoning feels like magic until you see the compute bill.

Profilbild von Zhongang Cai
Zhongang Caivor 6 Monaten

Don’t worry, one year ago I thought video models could never be accelerated to run in real time. Now guess what ;)

Profilbild von 0xmusashi
0xmusashivor 6 Monaten

@kellypeilinchan

Profilbild von Zhongang Cai
Zhongang Caivor 6 Monaten

@kellypeilinchan This is already happening ;)

Profilbild von Thomas
Thomasvor 6 Monaten

Interesting

Profilbild von John 🔮🌳 🌎☀️
John 🔮🌳 🌎☀️vor 6 Monaten

Who knew “compression” could turn into AGI haha

Profilbild von Critter
Crittervor 6 Monaten

A picture is worth a thousand words

Profilbild von Zhongang Cai
Zhongang Caivor 6 Monaten

True, and I wonder if words are necessary at all in many visual reasoning scenarios

Profilbild von Mohamed
Mohamedvor 6 Monaten

@CritterObserves @caizhongang for video QA (even for very long videos), we proved here that they can be done end to end on visuals losslessly (without intermediate captioning at all) + logarithmic compute:

Profilbild von Zhongang Cai
Zhongang Caivor 6 Monaten

@CritterObserves Interesting work!

Profilbild von Yuanzhi
Yuanzhivor 5 Monaten

This seems to be like a natural result of bidirectional attention and diffusion; would like to know the observation on AR video gen model, e.g., diffusion forcing?

Profilbild von Zhongang Cai
Zhongang Caivor 5 Monaten

Great idea!

Profilbild von Wang Ruisi
Wang Ruisivor 6 Monaten

Our GitHub repository is now live for discussions and upcoming tool releases. Follow us for the latest updates! GitHub:

Ähnliche Videos

AI has transformed how video is created. We think the next wave is about understanding it. Over the past few years, we've seen remarkable advances in video generation, editing, avatars, and creative tooling. An increasingly important problem is teaching machines to search, analyze, reason over, and extract insight from video - across massive libraries and live streams alike. We're calling this video intelligence, and we're actively looking to back founders building here. We're most excited about companies pushing on the core capabilities: - Video-native models - multimodal embeddings, temporal reasoning, and retrieval built specifically for video rather than adapted from image or text - Real-time and large-scale pipelines - infrastructure for processing, indexing, and querying video at the speed and scale enterprises actually need - Agentic and reasoning layers - systems that don't just retrieve clips but answer questions, surface anomalies, and take action on what they see The models and infrastructure to make this real are appearing to be crossing a capability threshold right now. Multimodal foundation models are maturing, storage costs have collapsed, and enterprises are sitting on years of unstructured video with no way to use it. That infrastructure unlocks a wide range of applications including media and sports workflows, security and physical operations, enterprise knowledge management, advertising analytics, robotics, and consumer products, where video has historically been dark data. If you're building in video intelligence at the model layer, the platform layer, or in a vertical application, we'd love to talk!

Jason Cui

36,205 Aufrufe • vor 4 Monaten