Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introduce CoT-VLA โ€“ Visual Chain-of-Thought reasoning for Robot Foundation Models! ๐Ÿค– By leveraging next-frame prediction as visual chain-of-thought reasoning, CoT-VLA uses future prediction to guide action generation and unlock large-scale video data for training. #CVPR2025

48,963 Aufrufe โ€ข vor 1 Jahr โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

๐—œ'๐˜ƒ๐—ฒ ๐—ต๐—ฒ๐—ฎ๐—ฟ๐—ฑ ๐˜๐—ต๐—ถ๐˜€ ๐—ฎ ๐—น๐—ผ๐˜ ๐—ฟ๐—ฒ๐—ฐ๐—ฒ๐—ป๐˜๐—น๐˜†: "๐—ช๐—ฒ ๐˜๐—ฟ๐—ฎ๐—ถ๐—ป๐—ฒ๐—ฑ ๐—ผ๐˜‚๐—ฟ ๐—ฟ๐—ผ๐—ฏ๐—ผ๐˜ ๐—ผ๐—ป ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜ ๐—ฎ๐—ป๐—ฑ ๐—ถ๐˜ ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฒ๐—ฑ ๐˜๐—ผ ๐—ฎ ๐—ป๐—ผ๐˜ƒ๐—ฒ๐—น ๐—ผ๐—ฏ๐—ท๐—ฒ๐—ฐ๐˜ - ๐˜๐—ต๐—ฒ๐˜€๐—ฒ ๐—ป๐—ฒ๐˜„ ๐—ฉ๐—Ÿ๐—” ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€ ๐—ฎ๐—ฟ๐—ฒ ๐—ฐ๐—ฟ๐—ฎ๐˜‡๐˜†!" Let's talk about what's actually happening in that "A" (Action) part of your VLA model. The Vision and Language components? They're incredible. Pre-trained on internet-scale data, they understand objects, spatial relationships, and task instructions better than ever. But the Action component? That's still learned from scratch on your specific robot demonstrations. ๐—›๐—ฒ๐—ฟ๐—ฒ'๐˜€ ๐˜๐—ต๐—ฒ ๐—ฟ๐—ฒ๐—ฎ๐—น๐—ถ๐˜๐˜†: Your VLA model has internet-scale understanding of what a screwdriver looks like and what "tighten the screw" means. But the actual motor pattern for "rotating wrist while applying downward pressure"? That comes from your 500 robot demos. ๐—ช๐—ต๐—ฎ๐˜ ๐˜๐—ต๐—ถ๐˜€ ๐—บ๐—ฒ๐—ฎ๐—ป๐˜€ ๐—ณ๐—ผ๐—ฟ "๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป": โ€ข ๐—ฉ๐—ถ๐˜€๐—ถ๐—ผ๐—ป ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Recognises novel objects instantly (thanks to pre-training) โ€ข ๐—Ÿ๐—ฎ๐—ป๐—ด๐˜‚๐—ฎ๐—ด๐—ฒ ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Understands new task instructions (thanks to pre-training) โ€ข ๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ด๐—ฒ๐—ป๐—ฒ๐—ฟ๐—ฎ๐—น๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป: Still limited to motor patterns seen during robot training Ask that same robot to "unscrew the bottle cap" and it fails because: โ€ข Vision: Recognises bottle and cap โ€ข Language: Understands "unscrew" โ€ข Action: Never learned the "twist while pulling" motor pattern ๐—ง๐—ต๐—ฒ ๐—ต๐—ฎ๐—ฟ๐—ฑ ๐˜๐—ฟ๐˜‚๐˜๐—ต ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ ๐—ฉ๐—Ÿ๐—” ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€: The "VL" gives you incredible zero-shot understanding. The "A" still requires task-specific demonstrations. We've cracked the perception and reasoning problem. We haven't cracked the motor generalisation problem.

Stephen James

51,386 Aufrufe โ€ข vor 1 Jahr

๐Ÿš€ A better, faster co-folding-based binding affinity model. Predicting how tightly a drug candidate binds to its target is critical in drug discovery. It also requires massive computational resources. State-of-the-art models can take 20 seconds to a minute per prediction, impractical for the demands of large scale early-stage programs . ๐Ÿ’  Today, Recursionโ€™s Valence Labs is releasing Nesso-1: the fastest open-source co-folding-based binding affinity model available. At 1 second per prediction, itโ€™s roughly 20x faster than our previous collaboration on Boltz-2 while matching or surpassing its accuracy across public and internal benchmarks. By leveraging NVIDIA Healthcare cuEquivariance, weโ€™ve been able to further accelerate both training and inference by an additional 2-3x. We look forward to continuing to improve Nesso-1 in collaboration with NVIDIA. Weights and code are fully open-sourced. The core architectural ideas behind Nesso-1 build on the insight that coarse-grained co-folding representations can match full-atom models for affinity prediction at a fraction of the cost. Nesso-1 is the first open implementation of this approach with no proprietary dependencies, trained entirely on public data, built to be reproducible and extensible. Weโ€™re already using Nesso-1 internally in active drug discovery programs. Fast, reliable affinity prediction at scale is foundational to the kind of autonomous design loops that define our vision for Autonomous Precision Design and Nesso-1 is a meaningful step toward that. ๐Ÿ‘‰ Report: ๐Ÿ‘‰ Github: ๐Ÿ‘‰ HF:

Recursion

156,802 Aufrufe โ€ข vor 2 Monaten