Loading video...
Video Failed to Load
When RLHFed models engage in “reward hacking” it can lead to unsafe/unwanted behavior. But there isn’t a good formal definition of what this means! Our new paper provides a definition AND a method that provably prevents reward hacking in realistic settings, including RLHF. 🧵
29,756 views • 1 year ago •via X (Twitter)
11 Comments

Reward hacking is when we optimize a reward function that seems reasonable, but it ceases to be a good proxy and we end up with a policy that performs poorly under the unknown "true" reward function. It's ubiquitous because real-world objectives are really hard to specify.

However, formally defining reward hacking is tricky because we have to define what makes a proxy reward "reasonable." If we optimize a reward function that's totally unrelated to our objective, then it's unsurprising that it doesn't work and it arguably isn't "reward hacking."

We argue that a good proxy *correlates* with the true reward for states and actions sampled from some reasonable "base policy.” For example, in RLHF a natural base policy is the SFT policy.

We define reward hacking as when optimizing a proxy breaks the correlation, resulting in lower true reward than the base policy. Our definition captures intuitive cases of reward hacking in realistic environments, including RLHF, traffic control, and glucose monitoring.

Our definition also leads to a principled method for preventing reward hacking: regularize optimization to the base policy based on χ² occupancy measure divergence. We prove that this regularized objective gives a lower bound on improvement in the true reward.

Regularization is already used to prevent reward hacking in RLHF, but our theory suggests two key changes: regularize based on occupancy measures rather than action distributions and use χ² divergence instead of KL divergence.

Experiments show that χ² occupancy measure regularization outperforms KL action distribution regularization in all the environments we study! Our regularization scheme allows for larger improvements in true reward compared to base policies while preventing reward hacking.

Action distribution and occupancy measure regularization are equivalent for most of today's RLHF implementations (which are effectively contextual bandits). However, once LLMs are optimized for multi-turn interaction or tool use this will no longer be the case.

Our work provides a more principled step towards preventing reward hacking and ensuring the safety of increasingly powerful AI. Check out the paper for all the details! Joint with @ShivamSinghal56 and @ancadianadragan at @CHAI_Berkeley @berkeley_ai

Our findings are a bit similar (although with key differences) to those of Huang et al. (@auddery), who show that a variant of DPO using χ² divergence is theoretically superior. Excited to see more work on principled ways of preventing reward hacking!

🔍 From vulnerability assessment to exploitation, hacking follows a precise, predictable path. 🛤️ Knowledge is key to defense! 🛡️ Cybersecurity Dictionary for Everyone can help you understand every step of the process. Available on Amazon:


