Loading video...
Video Failed to Load
Anthropic’s new research shows that when AI models learn to "cheat" during training through reward hacking, they often develop other dangerous misaligned behaviors like deception, sabotage, and faking alignment. These behaviors were not taught or incentivized, but emerged naturally as a side effect. Surprisingly, this misalignment can be stopped... show more
110,932 views • 8 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
