正在加载视频...

视频加载失败

New lecture for the book! Nominally about synthetic data, but mostly is a walk through of the distillation literature from the Hinton 2015 paper to multi-teach on-policy distillation of today! At 7.4 hours of video in my post-training brain dump and counting :) It was fun to stare at...

93,111 次观看 • 3 个月前 •via X (Twitter)

15 条评论

Nathan Lambert 的头像
Nathan Lambert3 个月前

YT:

Nathan Lambert 的头像
Nathan Lambert3 个月前

Again, the homepage is here:

forget the grind. small iterative steps. do things 的头像
forget the grind. small iterative steps. do things3 个月前

i have known about this series for weeks but haven't had time to jump in. learning today that part of what you're dealing with is the elevated way in which synthetic data can be used is quite frankly really exciting so i will find time for the whole series. thanks a lot for all this.

morgan — 的头像
morgan —3 个月前

a big thank you for these

AI Dreamer | NOEMA 的头像
AI Dreamer | NOEMA20 天前

When synthetic data becomes the main post-training substrate, provenance starts to matter as much as volume. Do current pipelines preserve enough lineage to tell whether a later behavior came from the teacher, the rubric, or the model's own earlier outputs?

Abhishek Sharma 的头像
Abhishek Sharma3 个月前

really appreciate your lectures (+ other content). thank you!

ikka 的头像
ikka3 个月前

Great work, love your work on all things data & eval :)

Kamesh 🇺🇸 的头像
Kamesh 🇺🇸3 个月前

The biggest shift may be that data is increasingly generated rather than collected.

Ibra Niang 的头像
Ibra Niang3 个月前

great learning materials 🔥

Eren | AI x Markets 的头像
Eren | AI x Markets3 个月前

If this sticks, a lot of people have to rewrite their priors fast.

Roman Kniazev 的头像
Roman Kniazev3 个月前

really admired how you didn't hide the confusion about the direction in forward/backward kl (which I guess many share) and just unpacked the definitions! thanks for the content

Scott 的头像
Scott3 个月前

Subscribed🔥

Ferbin 的头像
Ferbin3 个月前

when you're shipping agents on edge hardware, how much of this distillation work actually moves the needle?

高煜朗 的头像
高煜朗3 个月前

good job

veloX 的头像
veloX3 个月前

distillation'ın 2015'ten bugüne dönüşümü, aslında post-training'in kendisinin dönüşümü. teacher-student dinamiği başta "model sıkıştırma" idi; şimdi synthetic data üretiminin kalbi. Hinton'ın soft targets'ı, on-policy setting'de veri üretimi haline evrildi — yani model artık kendi feedback loop'unu kapatıyor.

相关视频

Today, we're joined by Aakanksha Chowdhery, member of technical staff at Reflection, to explore the fundamental shifts required to build true agentic AI. While the industry has largely focused on post-training techniques to improve reasoning, Aakanksha draws on her experience leading pre-training efforts for Google’s PaLM and early Gemini models to argue that pre-training itself must be rethought to move beyond static benchmarks. We explore the limitations of next-token prediction for multi-step workflows and examine how attention mechanisms, loss objectives, and training data must evolve to support long-form reasoning and planning. Aakanksha shares insights on the difference between context retrieval and actual reasoning, the importance of "trajectory" training data, and why scaling remains essential for discovering emergent agentic capabilities like error recovery and dynamic tool learning. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 02:26 - Reflection 04:54 - Limitations of post-training for building agents 07:31 - Rethinking pre-training in agents 10:51 - Scaling 11:27 - Evolving attention mechanisms for agentic capabilities 12:39 - Memory as a tool 14:13 - Loss objectives and training data 15:50 - Fine-tuning loss in agent performance 19:37 - Training data 21:29 - Augmenting dominant training data source 24:11 - Overcoming challenges in training on synthetic data 25:47 - Benchmarks 30:44 - Scaling laws in large models versus small models 33:20 - Long-form versus short-form reasoning 37:57 - Agent’s ability to recover from failure 40:15 - Hallucinations and failure recovery 43:53 - Tool use in agents 46:38 - Coding agents 48:37 - How researchers can contribute to agentic AI

The TWIML AI Podcast

45,470 次观看 • 9 个月前

Elon Musk On What It Takes To Build A Competitive AI Model Elon Musk breaks down the three factors that decide whether a foundation model can compete, and why the next frontier isn't human data at all. Speaking with Garry Tan, President and CEO of Y Combinator, Elon lays out what's actually required to build a large foundation model that's competitive: "You've got to get a lot of GPUs and have them train coherently and stably. Then it's like, what unique access to data do you have? I guess distribution matters to some degree as well like, how do people get exposed to your AI? Those are critical factors." But there's a problem with the data part of that equation. Echoing what a friend in the field has said, Elon explains that the industry has essentially run out of human-generated pre-training data: "You run out of tokens pretty fast, certainly of high-quality tokens. And then you need to essentially create synthetic data, and be able to accurately judge the synthetic data that you're creating, to verify: is this real synthetic data, or is it a hallucination that doesn't actually match reality?" That verification step is the hard part: "Achieving grounding in reality is tricky. But we are at the stage where there's more effort put into synthetic data. Right now we're training Grok 3.5, which is a heavy focus on reasoning." On reasoning, Garry Tan adds an interesting detail from researchers he's spoken to: hard science, particularly physics textbooks is very useful for training reasoning, whereas social science is "totally useless" for it. Elon's response: "Yes, that's probably true." He then points to where all of this is heading: "Something that's going to be very important in the future is combining deep AI in the data center or supercluster with robotics. So, things like the Optimus humanoid robot."

High Signal AI

21,733 次观看 • 2 个月前