Loading video...

Video Failed to Load

Go Home

Vision-language-action models (VLAs) need to REASON, but more importantly, they need to know WHEN to reason (or not)! Thrilled to introduce OneTwoVLA, a single, unified model that combines acting (System One) ⚡ and reasoning (System Two) 🤔, and can adaptively switch between these modes. Get ready for some serious...

22,501 views • 1 year ago •via X (Twitter)

10 Comments

Yang Gao's profile picture
Yang Gao1 year ago

💡 How does it work? OneTwoVLA only engages in reasoning at key steps – like completing a subtask or detecting an error. This reasoning includes scene descriptions, task plans, historical summaries, and more. At other times, OneTwoVLA generates actions based on its most recent reasoning. This cool capability comes from our specially designed adaptive inference framework and curated robot data with embodied reasoning.

Yang Gao's profile picture
Yang Gao1 year ago

🏞️ We've also developed a scalable pipeline for synthesizing embodied reasoning-centric vision-language data. This data is used for co-training with robot data, significantly enhancing the model's reasoning and generalization abilities!

Yang Gao's profile picture
Yang Gao1 year ago

OneTwoVLA packs diverse capabilities into a single model. First off, it excels at handling long-horizon manipulation tasks. Check out the video below where OneTwoVLA successfully tackles a super challenging task: hotpot cooking! 🔥🍲

Yang Gao's profile picture
Yang Gao1 year ago

Through co-training with vision-language data, OneTwoVLA can understand and complete task instructions it has never seen in its robot data (e.g., "Give me an icy cola." 🥤).

Yang Gao's profile picture
Yang Gao1 year ago

🔧 Error detection on the fly! OneTwoVLA can spot errors in real-time, rapidly reason about recovery strategies, and then generate corrective actions.

Yang Gao's profile picture
Yang Gao1 year ago

🤝 OneTwoVLA can also engage with humans naturally – seamlessly handling interventions and proactively seeking clarification when faced with ambiguities. It's a team player!

Yang Gao's profile picture
Yang Gao1 year ago

Finally, we've found that OneTwoVLA exhibits strong visual grounding capabilities. It understands spatial relationships, object attributes, and semantic features, even generalizing to objects unseen in its robot training data (e.g., GoPro 📷, Sprite 🥤, Starbucks Coffee ☕)!

Yang Gao's profile picture
Yang Gao1 year ago

Project website: Paper: Code: Data: Amazing work done with @lfqirrrrr, @RuiqianNai, @yingdong_hu99, @YouJiacheng, Junming Zhao

عبد العزيز الرفاعي's profile picture
عبد العزيز الرفاعي1 year ago

@Presidentlin

Aisha's profile picture
Aisha1 year ago

Do you think you can build it on your own to peel and chop onions? (prepare the ingredients to cook with?)

Related Videos