Video yükleniyor...
Video Yüklenemedi
New (shorter) lecture! Over-optimization, foundations of reward hacking, sycophancy, verbosity, etc. In recording this, I realized that rubrics are going to be prone to overopt in a way like reward models, where RLVR is its own thing. This is mostly fundamentals, history, and reflections! 00:00 Intro & Why We... show more
33,139 görüntüleme • 15 gün önce •via X (Twitter)
0 Yorum
Yorum bulunmuyor
Orijinal gönderinin yorumları burada görünecek

