Loading video...
Video Failed to Load
1/ Au revoir, RLVR. New work: EBFT (Energy-Based Fine-Tuning), a post-training method that directly optimizes the long-horizon behavior of model generations, addressing SFT’s deployment-time error amplification without relying on sparse, task-specific rewards.
267,394 views • 5 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
