正在加载视频...
视频加载失败
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
2,572,973 次观看 • 8 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里


