Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🧵 More stable PPO for LLM post-training. Introducing EasyPPO: just fix the critic. ❌ No new actor loss ❌ No new policy algorithm ✅ Actor update unchanged EasyPPO delivers better results with zero training collapse across all our experiments.

39,068 Aufrufe • vor 2 Tagen •via X (Twitter)

19 Kommentare

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

All experiments use Qwen3.5-9B with rollout batch sizes from 512 to 1024, across three very different LLM RL settings: FrontierCS (continuous reward), AIME24 (binary reward), and Search-R1 (agentic). EasyPPO consistently achieves better performance than several PPO baselines.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

So where does PPO instability actually come from? We found several surprisingly simple issues on the critic side. The first one: overlong filtering for gradients. If you come from critic-free methods such as GRPO, it is natural to drop truncated rollouts everywhere. But PPO has a critic — and dropping those samples from critic training changes what the value model learns. The critic then only sees the non-truncated subset, even as the policy starts producing more truncated rollouts. Our fix is simple: filter overlong samples for the actor, but keep them for the critic.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

The second issue: some prompts produce much more variable returns than others. This variance increases the critic’s gradient magnitude. High-variance prompts can contribute disproportionately large gradients, letting a small number of prompts dominate the critic update.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

The fix is simple: normalize each prompt’s critic loss by its return standard deviation. Now, the return-noise term contributes the same constant amount to the critic’s gradient scale across prompts.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

And this is exactly what we see in practice. Without noise normalization, critic gradient norms grow strongly with the prompt’s return standard deviation. After normalization, this dependence is largely removed across training steps. High-variance prompts no longer dominate the critic update.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

That’s EasyPPO. We also study how critic minibatch size affects PPO stability, along with more analysis and ablations. More details here: 📄 Paper: 💻 Code: 🌐 Website:

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

Huge thanks to Xuanyi Zhou for leading this project (he will be applying to PhD programs for Fall 2027), and to everyone on the team, @huanzhimao, @DachengLi177, @wenhaocha1, @YichuanM, @karthik_r_n, @alvinkcheung, and @profjoeyg, who made EasyPPO possible.

Profilbild von Ali Hatamizadeh
Ali Hatamizadehvor 2 Tagen

This is super cool. Congrats ! How about getting rid of the critic altogether ?

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

Thanks! I think getting rid of the critic makes a lot of sense for single-turn or cheap-to-evaluate tasks—FlashREINFORCE is a great example: But GRPO or other critic free method is hard to scale when rollouts (evaluation) are expensive. In math/coding, another rollout is a few more checks. In RSI, it can be a training run costing thousands. In biology, you may not even have multiple identical systems to roll out in parallel.

Profilbild von Ali Hatamizadeh
Ali Hatamizadehvor 2 Tagen

Yes, to be honest, GRPO is a terrible way of doing RL at scale. And I like your work as a step up from the PPO itself. But I still think you may be able to get rid of critic one way or another, for agentic long-horizon tasks. We just need more research in this direction to figure it out.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

Yes totally agree, PPO definitely has more challenges in a complex agentic setting than critic-free method, e.g., subagents can break your MDP assumption. We're still working on more agentic settings in both ways (critic-free and new critic architecture)

Profilbild von Tarun Suresh @ COLM 2026
Tarun Suresh @ COLM 2026vor 2 Tagen

Great work @MangQiuyang and team!

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

Thanks, Tarun!

Profilbild von Xiangxin Zhou
Xiangxin Zhouvor 2 Tagen

nice work

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

Thanks!

Profilbild von Gregor
Gregorvor 1 Tag

Hit critic divergence in a small RLHF run last month. Reward climbed then crashed around step 800. Never isolated if it was value loss scale or bootstrap quality. Did your collapses tend to start early or drift in mid-run?

Profilbild von AI Quanting
AI Quantingvor 2 Tagen

"Just fix the critic" can mean three different things though. Is it the loss, the target, or how often it updates? Curious which one actually killed the collapse.

Profilbild von Qiuyang Mang
Qiuyang Mangvor 2 Tagen

I think instabilities come from the different components of critic. The noise from the return and the normalized loss is the most interesting part of this work to me

Profilbild von AI Quanting
AI Quantingvor 2 Tagen

So more the loss shaping than the update frequency? If normalizing the loss alone kills most of the noise, that's a smaller fix than I'd have guessed.

Ähnliche Videos

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 Aufrufe • vor 1 Jahr