Loading video...
Video Failed to Load
It's time to rethink RL. Translating real world use into model improvements requires redesigning post-training algorithms for non-verifiable, per token rewards. At AI Engineer 's World Fair, we share our insights into scaling algorithms like SDPO for continual learning.
65,758 views • 13 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
