正在加载视频...
视频加载失败
We developed an RL method for fine-tuning our models for precise tasks in just a few hours or even minutes. Instead of training the whole model, we add an “RL token” output to π-0.6, our latest model, which is used by a tiny actor and critic to learn quickly... show more
0 条评论
暂无评论
原始帖子的评论将显示在这里
