
Zhongang Cai
@caizhongang • 1,381 subscribers
Staff Research Scientist, SenseTime Research. Multimodal, Video Reasoning, Spatial Intelligence, Virtual Humans. Ph.D., S-Lab/MMLab@NTU.
Videos

💡Videos and images should not be treated merely as inputs to understand or outputs to render. They can be first-class substrates for problem solving beyond language. 🎉Introducing VBVR-Pro: an infrastructure suite for visual intelligence: trainable, verifiable, RL-optimizable, and experimentally controllable. 🔗Homepage:
Zhongang Cai29,919 Aufrufe • vor 22 Tagen

Video might be the next intelligence substrate. Strikingly, video models are beginning to exhibit the same emergent reasoning behaviors first observed in LLMs—multi-path search, self-correction, and layer specialization. We demystify video reasoning and show it doesn’t happen frame-by-frame, but along diffusion steps. 🔗 📄 So, what's next? ;)
Zhongang Cai71,036 Aufrufe • vor 6 Monaten
Keine weiteren Inhalte verfügbar