Загрузка видео...

Не удалось загрузить видео

На главную

Okay okay, spent my weekend gooning around learning GRPO math. Here's some takes. Essentially, this is me yapping through a recap of smaller details on how GRPO is implemented, what Dr. GRPO changes, why, DAPO, connections to PPO, aggregating batches... Reading list below.

123,113 просмотров • 1 год назад •via X (Twitter)

Комментарии: 10

Фото профиля Nathan Lambert
Nathan Lambert1 год назад

More coherent version coming to @interconnectsai this week. I know this format won't be for everyone, but I hope some of you love it! RLHF Book: DeepSeekMath paper: Where does ratio come from in PPO? DAPO: DAPO announcement: My DAPO recap: Dr. GRPO: Dr. GRPO announcement: TRL GRPO implementation: Unbiased GRPO implementation: Thread on GRPO implementation on x:

Фото профиля Nathan Lambert
Nathan Lambert1 год назад

Watch on YouTube:

Фото профиля Nathan Lambert
Nathan Lambert1 год назад

Thanks to many authors and folks for discussing / proposing questions, @QGallouedec , @ethayarajh , @zzlccc, @hamishivi @vwxyzjn @danielhanchen -- have distilled a lot from y'all in the last 72hours. No, this content isn't really meant for you, you already know this shit :)

Фото профиля Nathan Lambert
Nathan Lambert1 год назад

I know this is B tier production quality but A tier nerding out.

Фото профиля Filip Brnadic
Filip Brnadic1 год назад

What's your crypto exit strategy? I track 30+ indicators that have successfully marked the top of a bull run in prior cycles. My logic is simple. ✅ Hold blue-chip assets during the bull run ✅ Avoid a -70% drawdown by selling near the top. Check it out 👇

Фото профиля Daniel Han
Daniel Han1 год назад

I actually thought Dr GRPO's 1/max_length for eg 1/4096 was to counteract gradient accumulation causing imbalanced losses (like when doing CE masked mean) It makes sense to remove the std() since 1/large(std) reduces impacts of super hard problems - 1/100 (hard) vs 1/1 (easy)

Фото профиля xlr8harder
xlr8harder1 год назад

is phrasing still a thing?

Фото профиля tom
tom1 год назад

you're not the first to goon over RL

Фото профиля Daniel Han
Daniel Han1 год назад

Awesome video - not 2nd tier quality, but 1st tier :)

Фото профиля Itsram10
Itsram101 год назад

That's the gooning I like, that all smart people must do

Похожие видео

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a higher-dimensional feedback channel. Every major AI lab trains models with RL today (OpenAI, Anthropic, DeepSeek). And their key bottleneck has always been the reward functions. GRPO by DeepSeek worked well for math and code because the environment gave a binary signal. But for real agent tasks, someone still has to hand-code the scoring function. That takes days and breaks every time the pipeline changes. RULER (implemented in OpenPipe ART, 10k stars) addresses the exact problem Karpathy identified. The reward criteria are defined in plain English, and an LLM evaluates each trajectory against that description to provide feedback for training. I trained a Qwen3 1.4B agent that plays 2048 using GRPO with this exact workflow. In this case, the agent saw the board, picked a direction, and RULER evaluated the outcome, all from this natural language definition. You can see the full implementation on GitHub and try it yourself. Here's the ART Repo: (don't forget to star it ⭐ ) Just like RLHF replaced manual rankings and GRPO replaced the critic model, natural language rewards are replacing hand-coded scoring functions. RL reward engineering is now prompt engineering. I wrote a full walkthrough covering RL for LLM agents, from RLHF to GRPO to RULER, in the article below.

Avi Chawla

350,213 просмотров • 2 месяцев назад

New Course: Reinforcement Fine-Tuning LLMs with GRPO! Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with Predibase by Rubrik, and taught by Travis Addair, its Co-Founder and CTO, and Arnav Garg, its Senior Engineer and Machine Learning Lead. Reasoning models have been one of the most important developments in LLMs. Reinforcement Fine-Tuning (RFT) uses rewards to encourage LLMs to find solutions to multi-step reasoning tasks such as solving math problems and debugging code - without needing pre-existing training examples like in traditional supervised fine-tuning. Group Relative Policy Optimization (GRPO) is a reinforcement fine-tuning algorithm gaining rapid adoption. Developed by the DeepSeek team and used to train the R1 reasoning model, GRPO uses reward functions that you can write in Python to assign rewards to model responses. It’s beneficial for tasks with verifiable outcomes and can work well even with fewer than 100 training examples. It can also significantly improve the reasoning ability of smaller LLMs, making applications faster and more cost effective. In this course, you’ll take a technical deep dive into RFT with GRPO. You’ll learn to build reward functions that you can use in the GRPO training process to guide an LLM toward better performance on multi-step reasoning tasks. In detail, you’ll: - Learn when reinforcement fine-tuning is a better fit than supervised fine-tuning, especially for tasks involving multi-step reasoning or limited labeled data. - Understand how GRPO uses programmable reward functions as a more scalable alternative to the human feedback required for other reinforcement learning algorithms, such as RLHF and DPO. - Frame the Wordle game as a reinforcement fine-tuning problem and see how an LLM can learn to plan, analyze feedback, and improve its strategy over time. - Design reward functions that power the reinforcement fine-tuning process. - Learn techniques for evaluating more subjective tasks, such as rating the quality of a text summary, using an LLM as a judge. - Understand why reward hacking happens and how to avoid it by adding penalty functions to discourage undesirable behaviors. - Learn the four key components of the loss calculation in the GRPO algorithm: token probability distribution ratios, advantages, clipping, and KL-divergence. - Launch reinforcement fine-tuning jobs using Predibase’s hosted training services. By the end of this course, you’ll be able to build and fine-tune LLMs using reinforcement learning to improve reasoning without relying on large labeled datasets or subjective human feedback. Please sign up here:

Andrew Ng

86,457 просмотров • 1 год назад

i swear why would you do this to him.. on HIS BIRTHDAY?! 😢 🐈‍⬛ some of you guys are just like, you actually don't care about what the situation is going on, right? like, you guys keep like, saying like, 6, 7, 6 or 7, 6 or 7, enhypen is what, enhypen is what. like, i'm okay. i actually don't care, so i'm okay. enhyoen is enhypen. and like i actually don't care so i'm okay, i always try my best to ignore that so i'm okay but don't do that to other members especially at their birthday like i'm okay but please don't do that to other others. i really don't want them to feel that in their birthday okay just like um like thinking is totally like free but if you're saying that in someone's birthday like i'm okay, i'm saying this again. i'm okay but i feel good, i feel good with all my fans. today's was such a happy birthday for me, but i'm really like nervous that what if other members.. what would feel like depressed in this situation? like i'm really scared for that so hopefully i don't want that happen in the in the future in other members birthday like in this kind of things like so the only thing i want to say is like let's see the situation like um let's look around okay, let's just look around like what is going on, i hopefully like yeah.. that's it like i'm fine i'm okay. i'm okay. i don't actually care. like, i'm fine. but i just want other members to be happy with no risks, okay? that's the reason.

ɪᴄɪᴇʟ. 🐚

79,503 просмотров • 3 месяцев назад

My Girlfriend Gave Me An HIV; Not Sure Where To Go From Here ( Should We Stay Together?) Caller (Tim):so my question is, my girlfriend gave me a... a quite serious STD, and my parents found out about this. They are now against our relationship saying that it's a way that God is showing that this relationship was not meant to be. Um, I've been fighting this like against my parents. And my question is, has... is there too much that has happened for this relationship to work? Uh, or am I trying too hard to make it work? Dr. John Delony: Bro, there's a lot going on here, man. Whole lot. Um, before I dig in, let's... let me just... let's just have a conversation so I can get to know you for a second. Is that cool? Caller: Okay, yeah. Dr. John Delony: Um, you sound a little bit nervous. Are you nervous, or is this just weird? Or, I realize you're calling a stranger being like, "Yeah, I got an STI from this woman that I love." This whole thing's kind of weird. Caller: Well, no. This is basically the first person outside of this situation that I've talked to, so... Dr. John Delony: Gotcha. Well, dude, I'm... I'm grateful for the trust. It sounds like there's a lot, lot going on here. Um, tell me about you. How old are you? What do you do for a living? All that kind of stuff. Caller: Uh, 27. Um, uh, going to... uh, studying nursing. Dr. John Delony: Okay. Caller: And, uh, just about to graduate here. Dr. John Delony: So, congratulations on that. Um, do you live at home? Caller: No, I don't. Dr. John Delony: Okay. So, what kind of STI did you get? Caller: Uh, HIV. Dr. John Delony: Okay. So, you've got HIV. Okay, so I was wondering if this was... The... my first question was why in the world do your parents know? Um, if I'm 27, the last people on planet Earth I would call and tell I have an STI is my parents. But, um, HIV is super serious. What is your, um, health prognosis? Caller: it's... so, I guess the situation gets a little bit more difficult because, my dad does computer diagnostics, and he was able to accidentally kind of find out I have this, and so I am... Dr. John Delony: Accidentally kind of? What does that mean? Caller: Meaning he wasn't expecting it. He wasn't looking for it. It was just kind of came up when, uh, he was doing the... I regularly do checkups. He's a... a doctor. Dr. John Delony: Oh, so he's a doctor at the place where you went to get checked out, and he discovered through the computer system that his son has HIV? Caller: Yes. it's a... alternative medicine that he does, and so it's... it's a very different kind from modern medicine, so it's just the way that he found out. Um, I just go to my dad for all his, checkups. Dr. John Delony: Yeah. What a mess. Okay, so tell me about this person that you're dating. Caller: Um, so... um... I guess what... uh, what would you like to know about her? Dr. John Delony: Well, just tell me about her. Caller: Well, uh, she's a... she's a nurse. Um, and, uh, Illinois, and, um... This is the second, uh, serious relationship she's had. Dr. John Delony: Okay. Caller: Uh, for me, it's... um, maybe the first serious relationship for me. Dr. John Delony: Okay. Caller: Um, and so, she's... I mean, she grew up in a Christian family. She's... um... Dr. John Delony: Did she knowingly... knowingly give you HIV? Caller: No, she didn't know about it. Dr. John Delony: Okay. Is she the first person you've ever slept with? Caller: Yes. Dr. John Delony: Okay. How long have y'all been together?

Hecto Crypto | NetLink ⛓

14,886 просмотров • 17 дней назад