
Cassidy Laidlaw
@cassidy_laidlaw • 1,209 subscribers
PhD student at UC Berkeley studying RL and AI safety. Also at https://t.co/OrEPAiR8b0
Videos

We built an AI assistant that plays Minecraft with you. Start building a house—it figures out what you’re doing and jumps in to help. This assistant *wasn't* trained with RLHF. Instead, it's powered by *assistance games*, a better path forward for building AI assistants. 🧵
Cassidy Laidlaw490,307 views • 1 year ago

When RLHFed models engage in “reward hacking” it can lead to unsafe/unwanted behavior. But there isn’t a good formal definition of what this means! Our new paper provides a definition AND a method that provably prevents reward hacking in realistic settings, including RLHF. 🧵
Cassidy Laidlaw29,756 views • 1 year ago
No more content to load