Загрузка видео...
Не удалось загрузить видео
Okay okay, spent my weekend gooning around learning GRPO math. Here's some takes. Essentially, this is me yapping through a recap of smaller details on how GRPO is implemented, what Dr. GRPO changes, why, DAPO, connections to PPO, aggregating batches... Reading list below.
123,113 просмотров • 1 год назад •via X (Twitter)
Комментарии: 10

More coherent version coming to @interconnectsai this week. I know this format won't be for everyone, but I hope some of you love it! RLHF Book: DeepSeekMath paper: Where does ratio come from in PPO? DAPO: DAPO announcement: My DAPO recap: Dr. GRPO: Dr. GRPO announcement: TRL GRPO implementation: Unbiased GRPO implementation: Thread on GRPO implementation on x:

Watch on YouTube:

Thanks to many authors and folks for discussing / proposing questions, @QGallouedec , @ethayarajh , @zzlccc, @hamishivi @vwxyzjn @danielhanchen -- have distilled a lot from y'all in the last 72hours. No, this content isn't really meant for you, you already know this shit :)

I know this is B tier production quality but A tier nerding out.

What's your crypto exit strategy? I track 30+ indicators that have successfully marked the top of a bull run in prior cycles. My logic is simple. ✅ Hold blue-chip assets during the bull run ✅ Avoid a -70% drawdown by selling near the top. Check it out 👇

I actually thought Dr GRPO's 1/max_length for eg 1/4096 was to counteract gradient accumulation causing imbalanced losses (like when doing CE masked mean) It makes sense to remove the std() since 1/large(std) reduces impacts of super hard problems - 1/100 (hard) vs 1/1 (easy)

is phrasing still a thing?

you're not the first to goon over RL

Awesome video - not 2nd tier quality, but 1st tier :)

That's the gooning I like, that all smart people must do
