正在加载视频...
视频加载失败
This is the Claude Code moment for Reinforcement Learning. Every frontier lab knows the secret: post-training is where the magic happens. If you can eval a task, you can use RL to benchmax your model. But building and scaling the RL loop has remained a dark art reserved for... show more
13,309 次观看 • 15 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
