If you're using GRPO in TRL, you should really switch to the new async trainer. In our benchmarks, it's ~2-4x faster 🔥
29,624 görüntüleme
The ML Intern can now ping you on Slack when it's finished training models, generating datasets, or compiling an analysis of the 1000 ablations you launched 💸