
Arip
@machinestein • 1,373 subscribers
Shorts
Videos

Why does critic-free RL work for LLMs? This question has been in the air for a long time. To answer it, we propose the “Value-Gradient Hypothesis.” GRPO and PPO-style methods often improve reasoning without relying on a learned value model. Our paper’s central claim is that critic-free does not mean value-free: the actor’s backward pass can still carry a useful credit-assignment signal.
Arip42,486 görüntüleme • 2 ay önce
Daha fazla içerik yok.