正在加载视频...

视频加载失败

With RL, the robot can learn very precise tasks, like fastening a zip tie, and can actually do it more consistently and more quickly than even human teleoperation.

19,075 次观看 • 6 个月前 •via X (Twitter)

6 条评论

Physical Intelligence 的头像
Physical Intelligence6 个月前

While the whole model takes a long time to train, with RLT we can adapt individual precise stages with as little as 15 minutes of robot data.

Physical Intelligence 的头像
Physical Intelligence6 个月前

We use RLT to fine-tune the most precise and critical stage of delicate tasks, such as using a screwdriver to attach a cover to one of our robot arms.

Physical Intelligence 的头像
Physical Intelligence6 个月前

The key idea with RL tokens (RLT) is to compress our model’s (e.g., π-0.6) internal representations into a concise feature vector, which can be used by a very small actor and critic network that trains in real time even as the robot is practicing the task.

Physical Intelligence 的头像
Physical Intelligence6 个月前

We developed an RL method for fine-tuning our models for precise tasks in just a few hours or even minutes. Instead of training the whole model, we add an “RL token” output to π-0.6, our latest model, which is used by a tiny actor and critic to learn quickly with RL.

Physical Intelligence 的头像
Physical Intelligence6 个月前

To learn more about RLT, check out our blog post:

TimelessTLD 的头像
TimelessTLD5 个月前

up for grabs 🤖

相关视频