Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Imitation learning is great, but needs us to have (near) optimal data. We throw away most other data (failures, evaluation data, suboptimal data, undirected play data), even though this data can be really useful and way cheaper! In our new work - RISE, we show a simple way to...

20,636 görüntüleme • 10 ay önce •via X (Twitter)

16 Yorum

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

To start - we know that BC can be fragile. Particularly, we know that BC struggles when placed in OOD configurations with sparse data coverage. Now we can of course collect more data - but this requires expensive high-quality demonstrations. But non-expert data can be plentiful - play data, failed attempts during teleoperation, or evaluation rollouts of existing policies. But it’s not clear how to actually use this data with BC. (2/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Our idea in Robust Imitation by Stitching from Experts (RISE) is simple - while you can’t just imitate non-optimal data, you can use it to learn how to *get back to expert states* using offline RL. It’s a particularly simple offline RL problem - no reward models needed; label expert data with a reward of 1, and non-expert data with 0, allowing dynamic programming to stitch together non-expert data with expert data. Importantly, it’s dead simple - throw in expert data with reward 1, non-expert data with reward 0 into the buffer and do offline RL with your favorite offline RL method. (3/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Simple enough right? - not so different from what @sidgreddy and @svlevine proposed in SQIL. Now, the issue is that plain old offline RL approaches often fail to stitch together disjoint trajectories under relatively sparse data coverage (especially from vision). So while there is a promise of using suboptimal data for recovery, it doesn’t work quite as well we’d want it to! (4/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Looking into it a bit deeper, Q functions are able to stitch reasonably well, but we find that the marginal action distribution captured by the policy is overly conservative, with little action coverage. This lack of coverage in the policy distribution makes it hard to actually make use of the stitching captured by the Q function. Simply put - the algorithm can't find the right actions to take because the policy distribution doesn't cover it! (5/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

The fix is easy! - allow the policy to “borrow” actions from nearby states. In doing so, policies can widen their action distributions in controlled ways, significantly improving their ability to stitch suboptimal data! This can be implemented in simple ways - e.g enforcing Lipschitz smoothness on the policy, or through simple data augmentation. Think of it like “fuzzing” nearby states to make stitching easier even in low data coverage settings. But the impact on policy recovery and success is huge! (6/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Let’s see how well this works! First, using RISE allows your imitation learning policies to be a *lot* more robust, using non-optimal data to recover from OOD scenarios. On the flipside, just imitation non-optimal data or standard offline RL is far less effective (as we see quantitatively) (7/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Secondly, we find that RISE can make use of failed or suboptimal demonstrations, significantly improving learned policies over standard imitation learning. Moreover this works across different tasks, including with deformables! (8/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Interestingly, this also extends to policy evaluations. We can iteratively keep reincorporating policy evalution rollouts into RISE to keep improving. Evaluation data is not wasted! (9/10)

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

What should you take away: 1) Use all your data, just throw it in the buffer with a 0. Don’t waste any of your failed data! 2) even suboptimal data can teach you how to recover back to experts 3) stitching can be hard in low coverage settings, some smoothness assumptions can really help! Fun diving deep into the guts of offline RL and learning a bunch about stitching behavior and recovery. Led by @kevin_huang8, Rosario Scalise, Cleah Winston along with Ayush Agrawal, @yunchuzh, @RohanBaijal, Markus Grotz, Byron Boots, @Ben_Burchfiel, @MashaItkina, Paarth Shah Curious to hear thoughts/feedback! Paper: Website: (10/10)

Edward Hu profil fotoğrafı
Edward Hu10 ay önce

really like using "auxiliary" data in general, since BC is quite restrictive on data needs. does the nearest-neighbor trick work for images? i wonder if u can generalize this to images by training a BC policy on some coarsened version of state (like a diffusion policy ;) )

Abhishek Gupta profil fotoğrafı
Abhishek Gupta10 ay önce

Yes works on images! We didn’t actually do nearest neighbors explicitly, but instead we did data augmentation with DINO features and we enforced Lipschitz smoothness with spectral norm. So all results are from images :)

Himanshu Kumar profil fotoğrafı
Himanshu Kumar10 ay önce

Abhishek, you're spot on! It's a waste to discard all that extra data; maybe we're missing out on some real gems, no?

Robert Scoble profil fotoğrafı
Robert Scoble10 ay önce

Imitation learning is a dead end. World models are AGI. Learn everything you can about them

Kyowoon Lee profil fotoğrafı
Kyowoon Lee10 ay önce

🤩

Jakie PLA profil fotoğrafı
Jakie PLA10 ay önce

RISE's approach to leveraging non-expert data is brilliant! Solving sparse coverage while maintaining task robustness - this could transform robotics pipelines 🛠️

Min Chon Chi profil fotoğrafı
Min Chon Chi10 ay önce

Interesting; RISE is a good mnemonic for imitation learning work.

Benzer Videolar

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 görüntüleme • 1 yıl önce

Major program launch: Data Analytics Professional Certificate! This large, five-course sequence takes you all the way to being job-ready as a data analyst, and shows how to use Generative AI as a thought partner to enhance your work in this role. Offered by on Coursera, this is taught by Sean Barnes, Ph.D., a Data Science & Engineering Leader at Netflix. Analyzing data remains one of the most important skills in where the world is going with AI. This comprehensive certificate takes you all the way to being job-ready. Each course comes with practical projects demonstrated in real-world contexts, such as analyzing sales data for a Korean bakery, video game sales trends across different regions, or identifying factors impacting customer retention for a communications company. You'll also work on estimating fire distribution for forest fire prevention, analyzing how a diamond's properties affect its market value, and developing predictive models for retail sales analysis, carbon emissions, and coral reef conservation. Here's some of what you'll learn: - How to define data and categorize it into its many types such as discrete & continuous numerical, structured & unstructured, time series, categorical, and know what insights can be derived from the different types of data categories. - How to differentiate between data-related job roles and their responsibilities, and how data flows through an organization from the moment of capture to decision-making. - How to perform data processing functions and apply conditional formatting in spreadsheets to extract business value from your data using statistical calculations and best practices for visualizing and interpreting data. - How to use LLMs for stakeholder analysis, data exploration, and data visualization. - Best practices for using LLMs for as a thought partner to data analysis work By the end of this professional certificate program, you will have learned core statistical concepts, analysis techniques, and visualization methodologies that will serve as the foundation for working as a data analyst. The world needs more data analysts, especially ones who know how to use modern generative AI. With data science roles projected to grow 36% by 2033, the skills taught in this program create new professional opportunities in data. Sign up here!

Andrew Ng

85,107 görüntüleme • 1 yıl önce