Загрузка видео...

Не удалось загрузить видео

На главную

Having humans annotate data to pre-train robots is expensive and time-consuming! Introducing SPRINT: A pre-training approach using LLMs and offline RL to equip robots w/ many language-annotated skills while minimizing human annotation effort! URL: 🧵👇

24,665 просмотров • 3 лет назад •via X (Twitter)

Комментарии: 7

Фото профиля Jesse Zhang
Jesse Zhang3 лет назад

Labeling demonstrations with natural language instructions in hindsight is standard, but it is tedious and expensive to scale. We propose automatically (1) **relabeling** language instructions and (2) **chaining** trajectories together to generate more training data.

Фото профиля Jesse Zhang
Jesse Zhang3 лет назад

(1) Relabeling: If we have two skills, "Put mug in coffee machine" and "Press brew button," we could call this "Make Coffee." In SPRINT, we do this relabeling **automatically** by prompting an LLM to summarize nearby instructions. This gives us 2-2.5X more pre-training data!

Фото профиля Jesse Zhang
Jesse Zhang3 лет назад

(2) Chaining: Offline RL can "stitch" trajectories to learn new behaviors. We carefully relabel rewards with offline RL and modified language instructions to allow stitching even with language conditioning!

Фото профиля Jesse Zhang
Jesse Zhang3 лет назад

Results Overall, this allows us to achieve 2-8X better zero-shot long horizon task execution in ALFRED, a realistic household simulator, and on a real robot setup! SPRINT agents also fine-tune more efficiently to new tasks in unseen environments! ALFRED results:

Фото профиля Jesse Zhang
Jesse Zhang3 лет назад

Real Robot Results With offline fine-tuning, SPRINT achieves superior performance on new, long-horizon manipulation tasks in previously unseen environments!

Фото профиля Jesse Zhang
Jesse Zhang3 лет назад

For more details about SPRINT and experiment results, please see our paper or website. Paper: Website: Work done in collaboration with @KarlPertsch, @JiahuiZhang_32, @JosephLim_AI. @JiahuiZhang_32 is applying for PhD this year!

Фото профиля OliviaLi
OliviaLi3 лет назад

It sounds great, so that humans don't have to complete such a large amount of work every day, just let the robot do it

Похожие видео

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 просмотров • 1 год назад