Loading video...

Video Failed to Load

Go Home

Introduce CoT-VLA โ€“ Visual Chain-of-Thought reasoning for Robot Foundation Models! ๐Ÿค– By leveraging next-frame prediction as visual chain-of-thought reasoning, CoT-VLA uses future prediction to guide action generation and unlock large-scale video data for training. #CVPR2025

48,381 views โ€ข 1 year ago โ€ขvia X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

๐—œ'๐˜ƒ๐—ฒ ๐—ต๐—ฒ๐—ฎ๐—ฟ๐—ฑ ๐˜๐—ต๐—ถ๐˜€ ๐—ฎ ๐—น๐—ผ๐˜ ๐—ฟ๐—ฒ๐—ฐ๐—ฒ๐—ป๐˜๐—น๐˜†: "๐—ช๐—ฒ ๐˜๐—ฟ๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฑ ๐—ผ๐˜‚๐—ฟ ๐—ฟ๐—ผ๐—ฏ๐—ผ๐˜ ๐—ผ๐—ป ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜ ๐—ฎ๐—ป๐—ฑ ๐—ถ๐˜ ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฒ๐—ฑ ๐˜๐—ผ ๐—ฎ ๐—ป๐—ผ๐˜ƒ๐—ฒ๐—น ๐—ผ๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜ - ๐˜๐—ต๐—ฒ๐˜€๐—ฒ ๐—ป๐—ฒ๐˜„ ๐—ฉ๐—Ÿ๐—” ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€ ๐—ฎ๐—ฟ๐—ฒ ๐—ฐ๐—ฟ๐—ฎ๐˜‡๐˜†!" Let's talk about what's actually happening in that "A" (Action) part of your VLA model. The Vision and Language components? They're incredible. Pre-trained on internet-scale data, they understand objects, spatial relationships, and task instructions better than ever. But the Action component? That's still learned from scratch on your specific robot demonstrations. ๐—›๐—ฒ๐—ฟ๐—ฒ'๐˜€ ๐˜๐—ต๐—ฒ ๐—ฟ๐—ฒ๐—ฎ๐—น๐—ถ๐˜๐˜†: Your VLA model has internet-scale understanding of what a screwdriver looks like and what "tighten the screw" means. But the actual motor pattern for "rotating wrist while applying downward pressure"? That comes from your 500 robot demos. ๐—ช๐—ต๐—ฎ๐˜ ๐˜๐—ต๐—ถ๐˜€ ๐—บ๐—ฒ๐—ฎ๐—ป๐˜€ ๐—ณ๐—ผ๐—ฟ "๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป": โ€ข ๐—ฉ๐—ถ๐˜€๐—ถ๐—ผ๐—ป ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Recognises novel objects instantly (thanks to pre-training) โ€ข ๐—Ÿ๐—ฎ๐—ป๐—ด๐˜‚๐—ฎ๐—ด๐—ฒ ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Understands new task instructions (thanks to pre-training) โ€ข ๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Still limited to motor patterns seen during robot training Ask that same robot to "unscrew the bottle cap" and it fails because: โ€ข Vision: Recognises bottle and cap โ€ข Language: Understands "unscrew" โ€ข Action: Never learned the "twist while pulling" motor pattern ๐—ง๐—ต๐—ฒ ๐—ต๐—ฎ๐—ฟ๐—ฑ ๐˜๐—ฟ๐˜‚๐˜๐—ต ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ ๐—ฉ๐—Ÿ๐—” ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€: The "VL" gives you incredible zero-shot understanding. The "A" still requires task-specific demonstrations. We've cracked the perception and reasoning problem. We haven't cracked the motor generalisation problem.

Stephen James

51,356 views โ€ข 11 months ago

๐Ÿ“ข Our lab has been exploring 3D world models for years โ€” and weโ€™re thrilled to share **PhysTwin**: a milestone that reconstructs object appearance, geometry, and dynamics from just a few seconds of interaction! Led by the amazing Hanxiao Jiang ๐Ÿ‘‰ PhysTwin combines **Gaussian splatting** with **inverse dynamics optimization** based on simple **spring-mass** systems. โš™๏ธ The result? Real-time, action-conditioned 3D video prediction under novel interactions (i.e., 3D world models). ๐Ÿ”‘ A few key takeaways: 1. Having the right structure (e.g., particles/masses) helps navigate the trade-off between sample efficiency, generalization, and broad applicability. 2. Visual foundation models (VFMs) have matured to the point where they can provide rich supervision for world modeling (e.g., tracking, shape completion). 3. Beyond VFMs, many crucial components have come together in recent years: Gaussian splats for rendering, NVIDIA Warp for high-performance simulation, and scene/asset generation from a wide range of labs and companies. The future of 3D world models is looking bright! โœจ 4. The resulting digital twin supports a wide range of downstream applicationsโ€”especially in data generation and policy evaluation, thanks to its realistic rendering and simulation capabilities. ๐ŸŽฅ All code and data to reproduce the results, along with interactive demos, are available on the website. Check the following visualizations of: (1) observations, (2) reconstructed state/actions, (3) interactive digital twins, and (4) the overlays between real-world robot teleoperation and our modelโ€™s open-loop predictions.

Yunzhu Li

25,279 views โ€ข 1 year ago