Loading video...
Video Failed to Load
1/ Au revoir, RLVR. New work: EBFT (Energy-Based Fine-Tuning), a post-training method that directly optimizes the long-horizon behavior of model generations, addressing SFT’s deployment-time error amplification without relying on sparse, task-specific rewards.
267,650 views • 6 months ago •via X (Twitter)
12 Comments

2/ The problem: SFT trains on true prefixes, but the deployed model conditions on its own outputs. Minor mistakes can snowball. We measure this gap with a sequence level, feature-matching loss: how far apart in embedding space are model generations and ground-truth completions?

3/ SFT (red) barely improves over the base model (grey), and RLVR (orange) actually makes it worse. EBFT (blue) directly optimizes this loss and achieves the lowest feature-matching loss at every completion length despite training with rollouts of only 8 tokens.

4/ How EBFT works: match features of generations to the ground-truth completion. Sample short rollouts, embed them alongside ground-truth with a frozen feature network, then reward each sample based on how well it matches ground-truth features. No reward model, no verifier.

5/ EBFT delivers strong OOD generalization. When trained on Python, SFT degrades transfer to other languages, while EBFT improves it. And on noisier translation domains, EBFT outperforms both SFT and RLVR– demonstrating its promise beyond clean, verifiable settings.

6/ Work done with amazing collaborators: Samy Jelassi, @MujinKwun, @rosieyzh, Yuanzhi Li, @nfusi , @du_yilun, @cdomingoenrich

7/ Tagging folks who might be interested: @ahatamiz1, @YejinChoinka, @natolambert, @_lewtun, @aviral_kumar2, @tengyuma, @Zanette_ai, @johnschulman2, @agarwl_, @ericzelikman

8/ links: Blog post: Paper: Code: Website:

Can you talk about the relationship to inverse RL? Eg, we had done exactly this in continuous control ( and makes me wonder how feature matching compares with IQ-learn for LLMs as seen here:

Error amplification in SFT is painfully real in production. Small distribution gaps snowball fast on long generations. The reward-free angle here is compelling, but how does the energy function handle tasks with multiple valid output paths? That's usually where these approaches get hairy.

sft drift is real. optimizing long-horizon without relying on clunky reward models is the truth. rip rlvr, you won't be missed.

really like this direction. reward design for long-horizon tasks is genuinely painful and most approaches just try to make the reward better. sidstepping it is a different bet entirely.

Energy-based for long-horizon. Clever fix for SFT's error snowball. How's the training stability compared to RLVR? Graphs look solid but 1.5B's pretty small.
