Loading video...
Video Failed to Load
Lecture 10 of my course! Nominally on regularization in RL, so I discuss the evolving role of the KL penalty in RL, but also a set of nice RL papers that explain what RL helps models generalize better than SFT -- with theory supporting it. When going through these,... show more
89,468 views • 1 month ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

