正在加载视频...

视频加载失败

🧵 More stable PPO for LLM post-training. Introducing EasyPPO: just fix the critic. ❌ No new actor loss ❌ No new policy algorithm ✅ Actor update unchanged EasyPPO delivers better results with zero training collapse across all our experiments.

39,141 次观看 • 2 天前 •via X (Twitter)

19 条评论

Qiuyang Mang 的头像
Qiuyang Mang2 天前

All experiments use Qwen3.5-9B with rollout batch sizes from 512 to 1024, across three very different LLM RL settings: FrontierCS (continuous reward), AIME24 (binary reward), and Search-R1 (agentic). EasyPPO consistently achieves better performance than several PPO baselines.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

So where does PPO instability actually come from? We found several surprisingly simple issues on the critic side. The first one: overlong filtering for gradients. If you come from critic-free methods such as GRPO, it is natural to drop truncated rollouts everywhere. But PPO has a critic — and dropping those samples from critic training changes what the value model learns. The critic then only sees the non-truncated subset, even as the policy starts producing more truncated rollouts. Our fix is simple: filter overlong samples for the actor, but keep them for the critic.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

The second issue: some prompts produce much more variable returns than others. This variance increases the critic’s gradient magnitude. High-variance prompts can contribute disproportionately large gradients, letting a small number of prompts dominate the critic update.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

The fix is simple: normalize each prompt’s critic loss by its return standard deviation. Now, the return-noise term contributes the same constant amount to the critic’s gradient scale across prompts.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

And this is exactly what we see in practice. Without noise normalization, critic gradient norms grow strongly with the prompt’s return standard deviation. After normalization, this dependence is largely removed across training steps. High-variance prompts no longer dominate the critic update.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

That’s EasyPPO. We also study how critic minibatch size affects PPO stability, along with more analysis and ablations. More details here: 📄 Paper: 💻 Code: 🌐 Website:

Qiuyang Mang 的头像
Qiuyang Mang2 天前

Huge thanks to Xuanyi Zhou for leading this project (he will be applying to PhD programs for Fall 2027), and to everyone on the team, @huanzhimao, @DachengLi177, @wenhaocha1, @YichuanM, @karthik_r_n, @alvinkcheung, and @profjoeyg, who made EasyPPO possible.

Ali Hatamizadeh 的头像
Ali Hatamizadeh2 天前

This is super cool. Congrats ! How about getting rid of the critic altogether ?

Qiuyang Mang 的头像
Qiuyang Mang2 天前

Thanks! I think getting rid of the critic makes a lot of sense for single-turn or cheap-to-evaluate tasks—FlashREINFORCE is a great example: But GRPO or other critic free method is hard to scale when rollouts (evaluation) are expensive. In math/coding, another rollout is a few more checks. In RSI, it can be a training run costing thousands. In biology, you may not even have multiple identical systems to roll out in parallel.

Ali Hatamizadeh 的头像
Ali Hatamizadeh2 天前

Yes, to be honest, GRPO is a terrible way of doing RL at scale. And I like your work as a step up from the PPO itself. But I still think you may be able to get rid of critic one way or another, for agentic long-horizon tasks. We just need more research in this direction to figure it out.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

Yes totally agree, PPO definitely has more challenges in a complex agentic setting than critic-free method, e.g., subagents can break your MDP assumption. We're still working on more agentic settings in both ways (critic-free and new critic architecture)

Tarun Suresh @ COLM 2026 的头像
Tarun Suresh @ COLM 20262 天前

Great work @MangQiuyang and team!

Qiuyang Mang 的头像
Qiuyang Mang2 天前

Thanks, Tarun!

Xiangxin Zhou 的头像
Xiangxin Zhou2 天前

nice work

Qiuyang Mang 的头像
Qiuyang Mang2 天前

Thanks!

Gregor 的头像
Gregor2 天前

Hit critic divergence in a small RLHF run last month. Reward climbed then crashed around step 800. Never isolated if it was value loss scale or bootstrap quality. Did your collapses tend to start early or drift in mid-run?

AI Quanting 的头像
AI Quanting2 天前

"Just fix the critic" can mean three different things though. Is it the loss, the target, or how often it updates? Curious which one actually killed the collapse.

Qiuyang Mang 的头像
Qiuyang Mang2 天前

I think instabilities come from the different components of critic. The noise from the return and the normalized loss is the most interesting part of this work to me

AI Quanting 的头像
AI Quanting2 天前

So more the loss shaping than the update frequency? If normalizing the loss alone kills most of the noise, that's a smaller fix than I'd have guessed.

相关视频

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 次观看 • 1 年前