正在加载视频...
视频加载失败
New lecture for the book! Nominally about synthetic data, but mostly is a walk through of the distillation literature from the Hinton 2015 paper to multi-teach on-policy distillation of today! At 7.4 hours of video in my post-training brain dump and counting :) It was fun to stare at... show more
15 条评论

YT:

Again, the homepage is here:

i have known about this series for weeks but haven't had time to jump in. learning today that part of what you're dealing with is the elevated way in which synthetic data can be used is quite frankly really exciting so i will find time for the whole series. thanks a lot for all this.

a big thank you for these

When synthetic data becomes the main post-training substrate, provenance starts to matter as much as volume. Do current pipelines preserve enough lineage to tell whether a later behavior came from the teacher, the rubric, or the model's own earlier outputs?

really appreciate your lectures (+ other content). thank you!

Great work, love your work on all things data & eval :)

The biggest shift may be that data is increasingly generated rather than collected.

great learning materials 🔥

If this sticks, a lot of people have to rewrite their priors fast.

really admired how you didn't hide the confusion about the direction in forward/backward kl (which I guess many share) and just unpacked the definitions! thanks for the content

Subscribed🔥

when you're shipping agents on edge hardware, how much of this distillation work actually moves the needle?

good job

distillation'ın 2015'ten bugüne dönüşümü, aslında post-training'in kendisinin dönüşümü. teacher-student dinamiği başta "model sıkıştırma" idi; şimdi synthetic data üretiminin kalbi. Hinton'ın soft targets'ı, on-policy setting'de veri üretimi haline evrildi — yani model artık kendi feedback loop'unu kapatıyor.

