Loading video...

Video Failed to Load

Go Home

Another quick lecture -- I've been asked many times for prereq's to my book and what you should know, so built a little lecture (with GLM 5.2) to cover some more basics. Topics include: 00:00 Introduction & Course Prerequisites 01:37 Language Models Overview 02:47 The LM Head 04:29 Softmax...

52,668 views • 3 months ago •via X (Twitter)

15 Comments

Nathan Lambert's profile picture
Nathan Lambert3 months ago

Lecture 0:

Nathan Lambert's profile picture
Nathan Lambert3 months ago

(video is ft. phoebe, second half with her on my lap)

Zach Mueller's profile picture
Zach Mueller3 months ago

I’ve really gotta sit down and spend some time with these. Btw is there a space you’ve made yet for students taking the videos to yap? (Discord etc)

Nathan Lambert's profile picture
Nathan Lambert3 months ago

Yeah there’s a book discord, link on site / description (don’t past discord links on x or get got)

Zach Mueller's profile picture
Zach Mueller3 months ago

Bootiful

Juergentron9000's profile picture
Juergentron90003 months ago

5.2 has been blowing my socks off. I've never considered selling a kidney for a Mac studio to run it locally. Until now.

kim's profile picture
kim3 months ago

thank you so much. learning A LOT from everything you’re sharing: from to rlhf book.

Dipanshu Kushwaha's profile picture
Dipanshu Kushwaha3 months ago

That sounds super helpful! It’s great to see you breaking down the essentials for everyone. Can’t wait to dive into it!

Vishal's profile picture
Vishal3 months ago

This is amazing , thanks 🤩

veloX's profile picture
veloX3 months ago

KL divergence ve entropy başlığında bir detay: ikisi de bilgi teorisinin temeli ama çoğu kişi entropy'yi "belirsizlik" olarak, KL divergence'ı ise "mesafe" olarak düşünüyor. Aslında ikisi de olasılık dağılımlarının "ekonomisi" hakkında konuşuyor — entropy kaynağın ne kadar "pahalı" olduğunu, KL divergence ise yanlış modelle kodlama yaparken ne kadar "harcadığını" ölçüyor. Post-training'de bu fark kritik: cross-entropy sadece loss, KL divergence ise modelin orijinal bilgisinden ne kadar sapmaya izin verdiğini kontrol ediyor.

Verma's profile picture
Verma3 months ago

Great Thanks for your effort and posting this prerequisite, Nathan!

Eren | AI x Markets's profile picture
Eren | AI x Markets3 months ago

If it does one annoying thing for me every day, I am listening.

Dr. Xi Zeng's profile picture
Dr. Xi Zeng3 months ago

The timestamped prereq lecture is a good filter. If someone cannot point to the LM head screen in the video, I trust their agent product claim a little less.

aitization 𝕏 's profile picture
aitization 𝕏 3 months ago

Interesting

Akshat B's profile picture
Akshat B3 months ago

Is the LM head really the biggest bottleneck for training efficiency now? I am reading that the softmax bottleneck in that last layer can suppress 95% of the gradient norm during pretraining. It is good you covered this because it makes some patterns almost impossible to learn.

Related Videos

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 views • 1 year ago

#명재현 댄스 챌린지 모음 zip 00:00 Candy 00:26 특 00:51 # menow 01:20 최애의 아이 01:36 NIGHT DANCER 01:57 Seven 02:21 Cream Soda 02:47 RINDAMAN 03:23 In Bloom 03:41 ALOHA 04:12 해금 04:29 Killin' Me Good 04:56 Bubble 05:19 블루체크 05:36 Ready or Not 05:49 Perfume 06:04 겐토 06:29 ビートDEトーヒ 07:00 3D 07:41 Worth It 07:52 B.O.M.B 08:28 첫 눈 08:43 War Cry 08:58 Perfect Night 09:37 Like I Do 09:48 Honey 10:05 DASH 10:27 Standing Next To You 10:46 Hype Boy 11:13 첫 만남은 계획대로 되지 않아 11:41 사랑스러워 12:07 때깔 12:22 Legit 12:51 Water 13:04 Two of Hearts 13:12 EASY 13:32 Smoke Remix 14:09 TAP 14:46 Martini Blue 15:03 BODY 15:26 PICK IT UP 15:41 LIGHTHOUSE 16:00 EENIE MEENIE 16:22 Nectar 16:45 C'est La Vie 17:07 쩔어 17:26 꿈빛 파티시엘 17:56 Impossible 18:14 포켓몬 18:29 SWEAT 18:51 빠나나날라 19:41 SPOT! 20:05 포철고 20:28 빛 20:56 MAESTRO 21:21 Girls Never Die 21:45 KICK IT 22:12 Armageddon 22:39 Taxi Blurr 23:00 Nothin' on You 23:20 Feel the POP 23:39 Shooting Star 24:00 진격의 방탄 24:16 ZOMBIE 24:30 LIKE THAT 24:45 ABCD 25:01 Boom Boom Bass 25:16 Yo Bunny 25:31 BADVILLAIN 25:51 올드스쿨 뉴스쿨 26:08 Supernatural 26:31 How Sweet 27:08 내일에서 기다릴게 27:31 Peaches 27:46 GOOD SO BAD 28:13 별별별 28:41 CRAZY 29:03 슈퍼슈퍼 29:32 껌 29:52 I Need A Girl 30:14 SAD SONG 30:32 네모네모 30:51 옴브리뉴 31:12 반가워, 나의 첫사랑 31:33 혀끝 32:00 Hit the Floor 32:16 기억사탕 32:34 Cherish

먕

91,257 views • 1 year ago

Today, we're joined by Aakanksha Chowdhery, member of technical staff at Reflection, to explore the fundamental shifts required to build true agentic AI. While the industry has largely focused on post-training techniques to improve reasoning, Aakanksha draws on her experience leading pre-training efforts for Google’s PaLM and early Gemini models to argue that pre-training itself must be rethought to move beyond static benchmarks. We explore the limitations of next-token prediction for multi-step workflows and examine how attention mechanisms, loss objectives, and training data must evolve to support long-form reasoning and planning. Aakanksha shares insights on the difference between context retrieval and actual reasoning, the importance of "trajectory" training data, and why scaling remains essential for discovering emergent agentic capabilities like error recovery and dynamic tool learning. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 02:26 - Reflection 04:54 - Limitations of post-training for building agents 07:31 - Rethinking pre-training in agents 10:51 - Scaling 11:27 - Evolving attention mechanisms for agentic capabilities 12:39 - Memory as a tool 14:13 - Loss objectives and training data 15:50 - Fine-tuning loss in agent performance 19:37 - Training data 21:29 - Augmenting dominant training data source 24:11 - Overcoming challenges in training on synthetic data 25:47 - Benchmarks 30:44 - Scaling laws in large models versus small models 33:20 - Long-form versus short-form reasoning 37:57 - Agent’s ability to recover from failure 40:15 - Hallucinations and failure recovery 43:53 - Tool use in agents 46:38 - Coding agents 48:37 - How researchers can contribute to agentic AI

The TWIML AI Podcast

45,470 views • 9 months ago

📁 2024 #이정현 챌린지 연말결산 00:00 Like I Do (1) 00:19 3D 00:36 흠뻑 00:58 첫 만남은 계획대로 되지 않아 01:22 Grillz 01:41 으르렁 01:55 JACKPOT 02:10 뉴스쿨 올드스쿨 02:37 사랑 없는 노래 02:59 Two of Hearts 03:10 때깔 03:27 전방향 미소년 03:37 전야 03:55 To. X 04:08 my alter ego 04:19 Everytime We Touch 04:34 Donkey Donkey Donk 04:49 What You Won’t Do for Love 04:59 TAP 05:29 Mongkol Challenge 05:51 Boo Thang 06:07 WISH 06:19 올 듯 말 듯 06:45 Smoothie 07:13 Boyfriend 07:29 Swag 07:50 꽁꽁 얼어붙은 한강 위로 08:04 SWEAT 08:25 Earth, Wind & Fire 08:40 포켓몬 댄스 08:53 아낀다 09:06 MILLIONS 09:21 ZOMBIE 09:33 WORK 09:51 Wet & Wild (WATERBOMB) 10:14 LIKE THAT 10:30 내가 S면 넌 나의 N이 되어줘 10:51 Small girl 11:07 STUPIDZ 11:27 Shooting Stars 11:47 Stay With Me 12:06 Treasure 12:20 Open Always Wins 12:45 Underground 13:00 Feels 13:19 Deja vu 13:34 삐그덕 13:48 태양의 후예 14:03 Chk Chk Boom 14:29 다라리 14:35 VOICEMAIL 14:53 Jus Know 15:11 Supernatural 15:36 MCNasty 15:53 Lay It Down 16:09 Loyal 16:33 MILLION DOLLAR BABY 16:49 ZOO 17:17 doodle 17:34 시카노코 (1) 17:45 Spin 17:59 시카노코 (2) 18:10 So Hot 18:30 Big Dawgs 18:44 샹하이 로맨스 19:03 Young, Wild & Free 19:31 Tell Me 19:43 꿈빛 파티시엘 19:59 APT. 20:05 Boyfriend 20:20 Hit the Floor 20:34 혀끝 21:00 LOVE, MONEY, FAME 21:28 I Don’t Care 21:47 TRIGGER 22:04 No Doubt 22:24 머리어깨무릎발 22:34 igloo 22:53 너와의 모든 지금 23:17 Maps 23:34 Knock Knock 23:42 Like I Do (2) 24:04 Santa Tell Me 24:16 Must Have Love #EVNNE #LEEJEONGHYEON #이정현

이브..

25,809 views • 1 year ago

내가 보려고 모은 #조이 역대 브릿지 파트 모음.zip (Ice Cream Cake~Sweet Dreams) 💚 * 레드벨벳 앨범만 포함했습니다 00:00 Somethin Kinda Crazy 00:05 Take It Slow 00:13 사탕 (Candy) 00:20 Dumb Dumb 00:22 Campfire 00:31 Lady’s Room 00:33 Day 1 00:43 Coold World 00:51 My Dear 00:55 Bad Dracula 00:58 Happily Ever After 01:03 빨간 맛 (Red Flavor) 01:09 여름빛 (Mojito) 01:13 I Just 01:17 Kingdom Come 01:23 두 번째 데이트 (My Second Date) 01:27 Attaboy 01:36 Perfect 10 01:42 달빛소리 01:53 Bad Boy 01:59 Time To Love 02:09 #Cookie Jar 02:13 Aitai-tai 02:20 Power Up 02:23 한 여름의 크리스마스 (With You) 02:30 Mr.E 02:35 Mosquito 02:41 RBB (Really Bad Boy) 02:47 So Good 02:51 멋있게 (Sassy Me) 02:54 Swimming Pool 02:57 Sayonara 03:04 짐살라빔 (Zimzalabim) 03:12 Sunny Side Up! 03:16 Milkshake 03:21 친구가 아냐 (Bing Bing) 03:29 안녕, 여름 (Parade) 03:32 LP 03:41 음파음파 (Umpah Umpah) 03:48 카풀 (Carpool) 03:57 Love Is The Way 04:05 Jumpin’ 04:11 눈 맞추고, 손 맞대고 (Eyes Locked, Hand Locked) 04:17 Psycho 04:24 In & Out 04:34 Remember Forever 04:44 Milky Way 04:52 Queendom 05:00 Pose 05:07 Knock On Wood 05:15 Better Be 05:25 Feel My Rhythm 05:29 Rainbow Halo 05:36 Beg For Me 05:43 In My Dreams 05:48 Marionette 05:50 WILDSIDE 05:54 Jackpot 06:00 Snap Snap 06:08 ‘Cause it’s you 06:14 Color of Love 06:16 Birthday 06:19 BYE BYE 06:22 롤러코스터 (On A Ride) 06:29 Zoom 06:40 Celebrate 06:48 Chill Kill 06:56 Knock Knock (Who’s There?) 06:59 Nightmare 07:03 One Kiss 07:10 Wings 07:21 풍경화 (Scenery) 07:32 Sunflower 07:38 Love Arcade 07:44 Bubble 07:51 Night Drive 07:57 Sweet Dreams

맑음쨩 🍏

20,678 views • 1 year ago

We don't know what most microbial genes do. Can genomic language models help? there's only one way to find out! this is a 1 hour and 42 minute interview with an MIT professor (the famous Yunha Hwang) chatting about these questions, her work in solving them at Tatta Bio, and more. zoomer captions are back too Links in reply! Timestamps: 00:00:00 - Clips + sponsor roll from the wonderful LatchBio 00:02:07 – Introduction 00:02:23 – Why do microbial genomes matter 00:04:07 – Deep learning acceptance in metagenomics 00:05:25 – The case for genomic “context” over sequence matching 00:06:43 – OMG: the only ML-ready metagenomic dataset 00:09:27 – gLM2: A multimodal genomic language model 00:11:06 – What do you do with the output of genomic language models? 00:17:41 – How will OMG evolve? 00:20:26 – Why train on only microbial genomes, as opposed to all genomes? 00:22:58 – Do we need more sequences or more annotations? 00:23:54 – Is there a conserved microbial genome ‘language’? 00:28:11 – What non-obvious things can this genomic language model tell you? 00:33:08 – Semantic deduplication and evaluation 00:37:33 – How does benchmarking work for these types of models? 00:41:31 – Gaia: A genomic search engine 00:44:18 – Even ‘well-studied’ genomes are mostly unannotated 00:50:51 – Using agents on Gaia 00:54:53 – Will genomic language models reshape the tree of life? 00:59:18 – Current limitations of genomic language models 01:08:54 – Directed evolution as training data 01:12:35 – What is Tatta Bio? 01:19:02 – Building Google for genomic sequences (SeqHub) 01:25:46 – How to create communities around scientific OSS 01:29:06 – What’s the purpose in the centralization of the software? 01:35:37 – How will the way science is done change in 10 years?

owl

44,279 views • 9 months ago

Gemini 3, scaling laws and the 'finite data' era: my conversation with Sebastian Borgeaud, research engineer at Google DeepMind and a pre-training lead for Gemini 3 00:00 – Cold intro: “We’re ahead of schedule” + AI is now a system 00:58 – Oriol Vinyals's “secret recipe”: better pre- + post-training 02:09 – Why AI progress still isn’t slowing down 03:04 – Are models actually getting smarter? 04:36 – Two–three years out: what changes first? 06:34 – AI doing AI research: faster, not automated 07:45 – Frontier labs: same playbook or different bets? 10:19 – Post-transformers: will a disruption happen? 10:51 – DeepMind’s advantage: research × engineering × infra 12:26 – What a Gemini 3 pre-training lead actually does 13:59 – From Europe to Cambridge to DeepMind 18:06 – Why he left RL for real-world data 20:05 – From Gopher to Chinchilla to RETRO (and why it matters) 20:28 – “Research taste”: integrate or slow everyone down 23:00 – Fixes vs moonshots: how they balance the pipeline 24:37 – Research vs product pressure (and org structure) 26:24 – Gemini 3 under the hood: MoE in plain English 28:30 – Native multimodality: the hidden costs 30:03 – Scaling laws aren’t dead (but scale isn’t everything) 33:07 – Synthetic data: powerful, dangerous? 35:00 – Reasoning traces: what he can’t say (and why) 37:18 – Long context + attention: what’s next 38:40 – Retrieval vs RAG vs long context 41:49 – The real boss fight: evals (and contamination) 42:28 – Alignment: pre-training vs post-training 43:32 – Deep Think + agents + “vibe coding” 46:34 – Continual learning: updating models over time 49:35 – Advice for researchers + founders 53:35 – “No end in sight” for progress + closing

Matt Turck

51,484 views • 9 months ago

2025 #지훈 챌린지 & 릴스 모음💝 00:00 sticky 00:16 샹하이 로맨스 00:40 Boom Boom Base (remix ver.) 01:01 오늘만 I LOVE YOU 01:30 💥💥💥 01:44 Make A Wish 02:20 뱅(Bang) ! 02:39 Kiss Kiss Shy Shy 02:52 wall🪼❤️‍🔥 03:07 🍉🍓🍊👌 03:24 Sweet Dreams 03:52 Lucifer 04:21 WE LIKE 2 PARTY 04:51 かわいいだけじゃだめですか? 05:08 Ayo 05:36 BEBE 05:56 NIN 06:10 Come on!⚡️ 06:24 GO 지훈🎉 GO 지훈🥳 06:40 🪼💸 07:07 BANDIT 07:32 KNOW ABOUT ME 07:53 뭘봐 08:15 Cool & hot 08:38 SOURPATCH 08:55 Chains 09:24 이뻐이뻐 09:55 행복 10:22 move like that🎵 10:40 Shake It To The Max✌️ 10:52 boy for the 42🩵 11:08 JIHOON Calling📞 11:13 I Feel Good 11:34 HANDS UP 12:12 ROCKSTAR 12:36 Thunder 12:59 Bad Desire 13:17 시끄러!! 13:26 Touch It🪼 13:43 good vibes☀️ 13:52 Killin' It Girl 14:16 ECHO! 14:41 🚘🕶️ 14:56 Buzzin💭 15:07 FAMOUS 15:18 빌려온 고양이 15:35 Domino 15:51 The Blue Sun 16:07 POP 16:32 シュガ リード ライブ 16:52 훈제란즈✌️✌️ 17:02 되고파 너의 17:17 🪼vs🪼 17:34 before sleep 🛌 17:44 Walk Like👟 17:54 🌼 18:02 Barbie Girl 18:11 HOWL 18:23 Tempo 18:43 🪼🔄 18:52 Dougie🤙 19:10 What You Want 19:37 😎😎 19:51 Bang🪼 19:58 ✌️✌️ 20:10 EZ 20:29 M.O. 20:55 I‘m right here 💐 21:12 산책 21:24 body 21:42 SPAGHETTI 22:15 Blue Valentine 22:50 Hollywood Action 23:17 BURNING UP 23:40 하얀 그리움 24:20 Break your legs👟 24:29 Let‘s 🪩 24:47 Talk to You 25:16 My Body 25:37 MISMATCH 26:02 Do you wanna dance?🎵 26:12 🦾 26:41 All I Want for Christmas Is You 26:58 Hmm🕶️ 27:08 ˏ₍⸜̠̇⸝̠̇₎ˎ ̀⁽⸌̠̇⸍̠̇⁾ ́ˏ₍⸜̠̇⸝̠̇₎ˎ ̀⁽⸌̠̇⸍̠̇⁾ ́ 27:22 You like it? 27:28 락 (樂) 27:48 Love Shot

🏝️

106,001 views • 9 months ago