Загрузка видео...
Не удалось загрузить видео
Language following is a tough problem for VLAs: while these models can follow complex language, in practice getting datasets that enable language following is hard. We developed a method to counterfactually and automatically label data to improve language following! 🧵👇
44,274 просмотров • 1 год назад •via X (Twitter)
Комментарии: 22

The main idea in CAST is to train a policy that responds to simple atomic commands ("go left" or "go right"), and then artificially generate pairs of counterfactual instructions (e.g., "drive along the glass wall") and corresponding atomic command ("go left"). The atomic command leads to the action (which is easy to get from the atomic policy), and it is relabeled with the long-form instruction, thus providing a new dataset with many more alternative, counterfactual commands and corresponding artificial actions. Training on this data then gives us a VLA with better language following!

This method, CAST, ends up significantly improving language following as compared to just labeling the original data. To find out more, check out: website: paper: A really fun project led by @CatGlossop w/ @verityw_ Arjun Bhorkar @shahdhruv_

@berkeley_ai this is the same idea as Multitask Preplay. Interestingly, we show it predicts human behavior and improves AI generalization to new environments where tasks co-occur

wild stuff fr

feels like you just found a cheat code for scaling instruction diversity

I sincerely hope to have the opportunity to communicate and learn from you. You may contact me through: WeChat: 15986759218 Email: [email protected] Thank you very much!

feels like the missing link between scripted bots and true instruction-following agents

synthetic labels as leverage, fam. model scaling trick without raw data pain

real moves

feels like synthetic data finally hit its stride

that’s actually a slick way to boost data diversity without the insane labeling grind

Addressing the challenge of language following in Virtual Language Assistants (VLAs) is crucial for enhancing their functionality. The difficulty lies in acquiring datasets that are sufficiently nuanced to train these models effectively. Your innovative approach to counterfactually and automatically label data represents a significant advancement. By simulating alternative scenarios and systematically tagging data, you can create richer training sets that improve the model's comprehension and responsiveness. This method not only enhances the accuracy of language understanding but also reduces the reliance on manually curated datasets, accelerating the development of more intelligent VLAs.

synthetic semantics goin crazy rn

Feels like the kind of trick that makes small data feel infinite, love it

massive leap for scaling instruction following

wild how fake tasks end up teaching real world moves bro

Automating labels to teach VLAs the art of listening sounds like you’re training the next generation of conversational wizards!

Operating with synthetic commands boosts versatility

Sounds like the dataset grind finally got a real shortcut

crazy efficient cook right there

Dear Professor Levine, I am Wang Zhigang, Business Director of Tianji Robotics. Our company mainly specializes in 7-axis full-joint force-controlled humanoid dual arms. By applying force control and impedance algorithms, these arms enable safer interaction with humans.

Counterfactual auto-labeling is the leap VLAs needed to break the dataset bottleneck.
