Загрузка видео...
Не удалось загрузить видео
Recently met Sasha Rush and he started giving me an impromptu lecture on how targeted on-policy self-distillation works. I asked him if I could record it on my iPhone. The basic idea is this: if the model made a mistake at some point in the rollout (for example, calling... show more
431,015 просмотров • 4 месяцев назад •via X (Twitter)
Комментарии: 35

@srush_nlp lfg, we need a full blackboard episode with sasha 🙏

@srush_nlp You should invite @willcb next !

@srush_nlp It always amazes me how quickly you can get an idea when someone explains it to you live vs trying to read about it in papers

@jxmnop @srush_nlp Always in awe of blackboard explanations! Somehow, concepts explained via explains on blackboards/whiteboards get deeply ingrained in mind + are crystal clear

@srush_nlp does he want a wife

@srush_nlp !!!

Here is how I would explain the confusion at 13min. Imagine you are doing a sequence of moves playing tennis. Student produces moves (or tokens) 1,2,3, only one rollout. OPD works as follows: At every given move (e.g. current token 1) *Nadal gets inside your brain* and produces a distribution over next move p(2/1). (logProbs for token2 conditioned on token1). Then, you update your neurons so that your distribution for next move looks more like Nadal's. Then, position 2 (token 2) is still your next bad move (NOT what Nadal would do). Nadal takes over your brain again, and computes P(3|2,1). Again you want to update your brain so that your distribution for token 3 (given your bad moves 1,2) looks like what Nadal would do, if he had gotten in this bad body position. There is only a single rollout, your bad moves 1,2,3. But the magic of LLMs is that Nadal can always replace your brain and tell you what HE WOULD DO at any given position. Now in OPSD replace Nadal with (you+extra hint). But you still update using your own bad moves without hint , ie the sequence (1,2,3). (thats why its on-policy).

I could watch a whole series of lectures like this. I find this far more engaging than watching a slide-deck. The imperfect nature, and the natural pauses allows me to focus more on the material and absorb it rather than passively watching a polished presentation with picture-perfect slides.

@srush_nlp the textual feedback bit just reminded me of GEPA. That’s literally the value add of GEPA that you can optimize on not just scalar rewards but detailed textual feedback. What if I just GEPA optimize the student based on textual feedback generated by another student? @srush_nlp?

@srush_nlp wait @srush_nlp so u guys localize the loss to the tokens in the student rollout that had a problem u are trying to fix? ie u mask out the rest of the student sequences?

@srush_nlp After having initially added the Self-distillation trainers into TRL, we have been working on refactoring them to be able to scale them as well as making it less cumbersome to add new variants: see and

@srush_nlp Egocentric data, but for expert lectures

@srush_nlp we need a full episode on distillation and RL by @srush_nlp :)

📍 Agent error repair turns AI capex into operating leverage Agent error repair flips AI capex from brute-force scaling into a measurable cost curve. If a bad tool call can be isolated, scored, and patched at the decision point, the model stops paying the same inference tax across every future rollout. That turns agent failure into amortizable technical debt, not permanent waste. The market consequence is sharper than “better agents.” Logs become the asset. Companies with dense workflow traces can lower cost per completed task while slower rivals keep buying GPUs to mask brittle behavior. The check is simple: tool-call error rate, recovery rate after correction, and completed task value per inference dollar. If those move together, the productivity case gets durable. If not, capex keeps eating the margin.

@srush_nlp So, path planning

@srush_nlp @CatAstro_Piyush

@srush_nlp Is kv caching of the prefix before the hint injection part of the reason this is able to be faster? Prefill is still required for any parts after the injection but no decode is that right @srush_nlp ?

@srush_nlp i wish x had auto-captions like youtube, it would help tremendously

@srush_nlp Exactly how you would teach someone. Nothing new.

@jonas_nelle @srush_nlp I personally love this format thank you for sharing.

@jxmnop @srush_nlp Cool

@srush_nlp hey dwarkesh, loved this. sent u a dm :)

@srush_nlp road markers , detours , road closed

@srush_nlp The model that creates those text amendments sounds a lot like a teacher... 🤔

@srush_nlp You can also do a version of this at test-time!

@srush_nlp You know you’ve made it when you have a blackboard at home

@jxmnop @srush_nlp 4:35 are we comparing to what teacher would have done or rather how teacher would score student rollout?

@srush_nlp Blackboards have so much aura

@mohammad2012191 @srush_nlp I remember him supporting Israel and its war crimes. I don't know if he still holds onto that because that made him absolutely vile.

@srush_nlp wonder if readers can differentiate between forward pass and a rollout. easily confusing.

@srush_nlp What happens if the second model has bias, is not a SOTA as first one, hallucinated & gave incorrect recommendation . I understand trajectory → reward = 0. But your bet is The bet is: “A slightly wrong local signal is better than an extremely sparse global signal.”

@srush_nlp sounds like a deep dive. self-distillation’s tricky.

@srush_nlp Now every student sitting in the front row of the lecture records video lecture as part of egocentric data collection 😆

@srush_nlp Maybe the actual text feedback comes from angry users :) ... (more likely than maybe even)

@srush_nlp Nice
