正在加载视频...
视频加载失败
New research from Databricks: LLMs Can Learn to Reason via Off-Policy RL Optimal Advantage-based Policy Optimization with Lagged Inference policy (OAPL) shows you don’t need strict on-policy training to improve reasoning. It matches or beats Group Relative Policy Optimization (GRPO), stays stable with large policy lag, and uses ~3×... show more
12,678 次观看 • 5 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
