Загрузка видео...
Не удалось загрузить видео
🧵 More stable PPO for LLM post-training. Introducing EasyPPO: just fix the critic. ❌ No new actor loss ❌ No new policy algorithm ✅ Actor update unchanged EasyPPO delivers better results with zero training collapse across all our experiments.
39,068 просмотров • 2 дней назад •via X (Twitter)
Комментарии: 19

All experiments use Qwen3.5-9B with rollout batch sizes from 512 to 1024, across three very different LLM RL settings: FrontierCS (continuous reward), AIME24 (binary reward), and Search-R1 (agentic). EasyPPO consistently achieves better performance than several PPO baselines.

So where does PPO instability actually come from? We found several surprisingly simple issues on the critic side. The first one: overlong filtering for gradients. If you come from critic-free methods such as GRPO, it is natural to drop truncated rollouts everywhere. But PPO has a critic — and dropping those samples from critic training changes what the value model learns. The critic then only sees the non-truncated subset, even as the policy starts producing more truncated rollouts. Our fix is simple: filter overlong samples for the actor, but keep them for the critic.

The second issue: some prompts produce much more variable returns than others. This variance increases the critic’s gradient magnitude. High-variance prompts can contribute disproportionately large gradients, letting a small number of prompts dominate the critic update.

The fix is simple: normalize each prompt’s critic loss by its return standard deviation. Now, the return-noise term contributes the same constant amount to the critic’s gradient scale across prompts.

And this is exactly what we see in practice. Without noise normalization, critic gradient norms grow strongly with the prompt’s return standard deviation. After normalization, this dependence is largely removed across training steps. High-variance prompts no longer dominate the critic update.

That’s EasyPPO. We also study how critic minibatch size affects PPO stability, along with more analysis and ablations. More details here: 📄 Paper: 💻 Code: 🌐 Website:

Huge thanks to Xuanyi Zhou for leading this project (he will be applying to PhD programs for Fall 2027), and to everyone on the team, @huanzhimao, @DachengLi177, @wenhaocha1, @YichuanM, @karthik_r_n, @alvinkcheung, and @profjoeyg, who made EasyPPO possible.

This is super cool. Congrats ! How about getting rid of the critic altogether ?

Thanks! I think getting rid of the critic makes a lot of sense for single-turn or cheap-to-evaluate tasks—FlashREINFORCE is a great example: But GRPO or other critic free method is hard to scale when rollouts (evaluation) are expensive. In math/coding, another rollout is a few more checks. In RSI, it can be a training run costing thousands. In biology, you may not even have multiple identical systems to roll out in parallel.

Yes, to be honest, GRPO is a terrible way of doing RL at scale. And I like your work as a step up from the PPO itself. But I still think you may be able to get rid of critic one way or another, for agentic long-horizon tasks. We just need more research in this direction to figure it out.

Yes totally agree, PPO definitely has more challenges in a complex agentic setting than critic-free method, e.g., subagents can break your MDP assumption. We're still working on more agentic settings in both ways (critic-free and new critic architecture)

Great work @MangQiuyang and team!

Thanks, Tarun!

nice work

Thanks!

Hit critic divergence in a small RLHF run last month. Reward climbed then crashed around step 800. Never isolated if it was value loss scale or bootstrap quality. Did your collapses tend to start early or drift in mid-run?

"Just fix the critic" can mean three different things though. Is it the loss, the target, or how often it updates? Curious which one actually killed the collapse.

I think instabilities come from the different components of critic. The noise from the return and the normalized loss is the most interesting part of this work to me

So more the loss shaping than the update frequency? If normalizing the loss alone kills most of the noise, that's a smaller fix than I'd have guessed.
