
Nathan Lambert
@natolambert • 100,011 subscribers
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote https://t.co/alRXKINTwE
Shorts
Videos

Lecture 10 of my course! Nominally on regularization in RL, so I discuss the evolving role of the KL penalty in RL, but also a set of nice RL papers that explain what RL helps models generalize better than SFT -- with theory supporting it. When going through these, it's so interesting how seasonal problems in ML are. Lots of problems from controlling reward models overopt will rhyme as we try to control rubrics for agents. 00:00 Intro & the role of regularization 02:50 The KL penalty in RL 10:53 RL as a reverse KL loss 21:09 Why RL generalizes better than SFT 25:15 Other regularization tools Just a few videos left as I get to the end of the course. Thanks all, and keep sending questions. Spread the word if you have a second.
Nathan Lambert89,468 görüntüleme • 1 ay önce

The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it: * Has potential for high real world impact * Clearly used extensively at frontier labs * Almost no empirical literature exists * More accessible on academic compute This lecture covers what character training is, reviews model specs, constitutions, the differences, the motivations in real world events, some example research papers I like, and open questions in how it relates to post-training/model use generally. Hopefully this brings more people into the field (and reach out if you have questions). It is one of the more research-y chapters in my book, but one that I felt needed the reference. There is still so little, educational content on the topic online. 0:00 Intro 6:22 Part 1: Fundamentals — character, constitutions, and model specs 19:21 Part 2: Character training in practice 23:23 Part 3: Character elicitation without gradient steps 28:03 Part 4: Open questions (and the end of the course) 32:27 The course, complete Thanks for watching. No need to like and subscribe now that the course is done, you definitely wouldn't! h/t to Sharan for leading the technical work I got to do in the space, and Zafir Stojanovski for investing a lot of attention at this book chapter.
Nathan Lambert57,041 görüntüleme • 1 ay önce

New lecture for the book! Nominally about synthetic data, but mostly is a walk through of the distillation literature from the Hinton 2015 paper to multi-teach on-policy distillation of today! At 7.4 hours of video in my post-training brain dump and counting :) It was fun to stare at the math long enough and talk through the 3-4 core changes that needed to be made to the original formulation to have on-policy distillation be ready for the mainstream like it is today (and in RL frameworks). Otherwise, I include a bit of a history lesson for how synthetic data generally slowly took over all post-training data research (it wasn't always the case)! Then I do some 101 review on constitutional AI, rubrics, and other popular methods. 00:00 The emergence of synthetic data 10:50 Background on teacher-student knowledge-distillation 24:47: On-policy distillation (OPD, MOPD, and OPSD) 37:11 Constitutional AI & AI Feedback 45:50 Rubrics as rewards & conclusions Ofc, watch on YouTube etc.
Nathan Lambert93,028 görüntüleme • 2 ay önce

My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on Florian Brand's insights). This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for. 00:00 Intro: frontier evaluation is harder than ever 03:39 Part 1: The eras of post-training evaluation 17:36 Part 2: An intro to agentic evals 21:19 Part 3: Can you trust the number? 30:41 Takeaways & conclusion Thanks for watching! Just one more lecture after this :)
Nathan Lambert39,123 görüntüleme • 1 ay önce

Lecture 11 is a tool-use/function calling/agentic 101! I almost skipped this one, as this chapter started as the only skill-specific topic in the book, but since writing it tool-use has only become more foundational to modern models. This is a tour from the basics -- why LLMs need tools -- to some of the cutting edge challenges scaling agentic RL. Of course, there's plenty of history and fundamentals along the way. Personally, I'm excited to get much more deeply into this area as a researcher again soon. Just a few more lectures left, but I'll likely keep making a few more videos now that I have pipelines that are lightweight and fun. 00:00 Introduction & Motivations 07:29 Part 1: Why Language Models Need Tools (+ related work) 15:30 Part 2: Infra - How Tool Calls Actually Work 24:17 Part 3: Training for Tool Use & OpenThoughts-Agent 36:38 Takeaways Thanks for watching & sharing questions.
Nathan Lambert43,949 görüntüleme • 1 ay önce

Another quick lecture -- I've been asked many times for prereq's to my book and what you should know, so built a little lecture (with GLM 5.2) to cover some more basics. Topics include: 00:00 Introduction & Course Prerequisites 01:37 Language Models Overview 02:47 The LM Head 04:29 Softmax & Log-Probabilities 06:13 Anatomy of an LM Training Example 06:37 Computing LLM Probabilities (+Phoebe the Dog) 09:52 Three Common Masks in Post-Training 11:03 A Small Decoding Review 12:14 Training an LM: Cross-Entropy 13:23 Optimization & Fine-Tuning 13:55 Pretraining to Midtraining to SFT Pipeline 15:25 Probability Essentials: KL Divergence & Entropy 19:36 Sigmoid & Pairwise Likelihood 20:29 Reinforcement Learning Framing (MDP) 22:28 Transitioning Tools into Post-Training 23:12 Recommended Resources & Wrap-Up Happy learning and I'm still taking questions from during the course for Q&A videos.
Nathan Lambert52,406 görüntüleme • 2 ay önce

New (shorter) lecture! Over-optimization, foundations of reward hacking, sycophancy, verbosity, etc. In recording this, I realized that rubrics are going to be prone to overopt in a way like reward models, where RLVR is its own thing. This is mostly fundamentals, history, and reflections! 00:00 Intro & Why We Care About Over-optimization 04:20 Part 1: Over-optimization & Goodhart's Law 09:57 Signatures of Over-Optimization & Misalignment 14:47 Part 2: Beyond "Just Style" 20:13 Llama 4 & Gaming the Leaderboards 23:04 Course Recap & Conclusion Primarily on Chapter 14.
Nathan Lambert33,228 görüntüleme • 1 ay önce

New lecture! This one is a recap of a bunch of history of preferences, the nature of rewards, how RLHF is formulated, which were once seen as central problems in the field. How much as changed. Still... super interesting to understand our optimization tools today. Books coming soon :D 00:00 Intro & context 07:34 A short history of preferences (from Aristotle to the VNM Utility Theorem) 20:17 A brief overview of preference data (from the last two years of my practice) 31:11 Open questions in RLHF data Lecture 8, covering Chapters 10 & 11 of my book.
Nathan Lambert28,657 görüntüleme • 1 ay önce

New podcast with finbarr! We survey the latest post-training recipes, from GLM 5.1, Kimi K2.6, DeepSeek V4, Xiaomi MiMo V2.5, Nemotron Ultra, etc. and discuss: - Why the industry slowly shifted to multi-teacher on-policy distillation (MOPD). - What an Olmo-style recipe would need improvements in - How post-training works / suits larger organizational efforts - Career advice in the foothills of the singularity - and other topics I heard y'all wanted me to start doing this, so making some time when I'm in funemployment! Chapters: 00:00 Introduction & Olmo reflections 06:28 Post-train recipes review (history) 23:00 2026’s model recipes (MiMo Flash, DeepSeek V4, GLM 5, Kimi K2.6, etc.) 39:05 Open-ended post-training discussions 48:22 Career advice in the LLM race Links below, please follow Interconnects AI and like and subscribe and buy my book?
Nathan Lambert43,127 görüntüleme • 2 ay önce

Okay okay, spent my weekend gooning around learning GRPO math. Here's some takes. Essentially, this is me yapping through a recap of smaller details on how GRPO is implemented, what Dr. GRPO changes, why, DAPO, connections to PPO, aggregating batches... Reading list below.
Nathan Lambert123,160 görüntüleme • 1 yıl önce

Here's a recent talk I gave recapping the last 6-12 months of AI progress, why getting perfect models is hard, how labs are likely approaching the next phase of training (for agents), and other interesting tidbits across the reasoning landscape. Topics: 00:00 Introduction & the state of reasoning 05:50 Hillclimbing imperfect evals 09:18 Technical bottlenecks 13:02 Sycophancy 18:08 The Goldilocks Zone 19:28 What comes next? (hint, planning) 26:40 Q&A YouTube etc in replies. Thanks Kyle Corbitt and OpenPipe for hosting me.
Nathan Lambert89,542 görüntüleme • 1 yıl önce

DPO Debate: Is RL needed for RLHF? All things as we cannot settle if DPO or RL is better. At least it is a good exercise. 1. Derivations in the DPO paper. Hint, the authors are good at math 2. cDPO, IPO, and related equations 3. Speculation on potential oddities of DPO vs RL 4. Reminders on the state of open RLHF tldr: we have more limitations with data and tooling and evaluation than optimizer choice Slides: Recent blog post of mine on DPO (more next Wed.): DPO Paper: On youtube:
Nathan Lambert100,027 görüntüleme • 2 yıl önce

Here's my conversation with Lucas Atkins and the team at Arcee.ai on their path to training and releasing Trinity Large today. From going all in on open models built end to end in the US 6 months ago to having the model in hand is no easy feet. I loved this conversation on how to design a startup around open models and take a bold step to scale it up. I'm openly an Arcee fan, watching them take risk and pull it off. We discuss: - The state (and future) of open vs. closed models, - The business of selling open models for on-prem deployments, - The story of Arcee AI & going “all-in” on this training run, - The ATOM project, - Building frontier model training teams in 6 months, - and other great topics. I really loved this one, and think you well too. Chapters: 00:00:00 Intro: Arcee AI, Trinity Models & Trinity Large 00:08:26 Transitioning a Company to Pre-training 00:13:00 Technical Decisions: Muon and MoE 00:18:41 Scaling and MoE Training Pain 00:23:14 Post-training and RL Strategies 00:28:09 Team Structure and Data Scaling 00:31:31 The Trinity Manifesto: US Open Weights 00:42:31 Specialized Models and Distillation 00:47:12 Infrastructure and Hosting 400B 00:50:53 Open Source as a Business Moat 00:56:31 Predictions: Best Model in 2026 01:02:29 Lightning Round & Conclusions More great open model builder podcasts coming soon!
Nathan Lambert26,872 görüntüleme • 7 ay önce

I re-recorded the post-training part of our NeurIPS tutorial on language models, added some more slides, and wrote up a mini state of the union on Interconnects. Enjoy! Links in QT. 00:00 Introduction 10:00 Prompts & Skill Selection 14:19 Instruction Finetuning 21:45 Preference Finetuning 36:17 Reinforcement Finetuning 45:28 Open Questions 52:02 Wrap Up
Nathan Lambert49,310 görüntüleme • 1 yıl önce

Another friday afternoon talk in attempt of open science. This is a re-record of a talk I gave at two Chantham House Rules workshops this fall: The History and Risks of Reinforcement Learning and Human Feedback (yes, the and is clever and intentional). This tries to broaden the discussion around RLHF to include more stakeholders and diverse ideas about how to integrate or measure values (can you at all?!?). I'm trying to re-record every talk that I give that isn't to a public audience 🫡😅, good practice at least. Slides: Youtube: Paper:
Nathan Lambert54,036 görüntüleme • 2 yıl önce
Daha fazla içerik yok.