Loading video...
Video Failed to Load
We found a new way to get language models to reason. 🤯 No RL, no training, no verifiers, no prompting. ❌ With better sampling, base models can achieve single-shot reasoning on par with (or better than!) GRPO while avoiding its characteristic loss in generation diversity.
278,167 views • 10 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

