Загрузка видео...

Не удалось загрузить видео

На главную

Recently met Sasha Rush and he started giving me an impromptu lecture on how targeted on-policy self-distillation works. I asked him if I could record it on my iPhone. The basic idea is this: if the model made a mistake at some point in the rollout (for example, calling...

431,015 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 35

Фото профиля elie
elie4 месяцев назад

@srush_nlp lfg, we need a full blackboard episode with sasha 🙏

Фото профиля samsja
samsja4 месяцев назад

@srush_nlp You should invite @willcb next !

Фото профиля Sam Z Liu
Sam Z Liu4 месяцев назад

@srush_nlp It always amazes me how quickly you can get an idea when someone explains it to you live vs trying to read about it in papers

Фото профиля Nikhil Kumar
Nikhil Kumar4 месяцев назад

@jxmnop @srush_nlp Always in awe of blackboard explanations! Somehow, concepts explained via explains on blackboards/whiteboards get deeply ingrained in mind + are crystal clear

Фото профиля Raye
Raye4 месяцев назад

@srush_nlp does he want a wife

Фото профиля nate
nate4 месяцев назад

@srush_nlp !!!

Фото профиля Alex Dimakis
Alex Dimakis3 месяцев назад

Here is how I would explain the confusion at 13min. Imagine you are doing a sequence of moves playing tennis. Student produces moves (or tokens) 1,2,3, only one rollout. OPD works as follows: At every given move (e.g. current token 1) *Nadal gets inside your brain* and produces a distribution over next move p(2/1). (logProbs for token2 conditioned on token1). Then, you update your neurons so that your distribution for next move looks more like Nadal's. Then, position 2 (token 2) is still your next bad move (NOT what Nadal would do). Nadal takes over your brain again, and computes P(3|2,1). Again you want to update your brain so that your distribution for token 3 (given your bad moves 1,2) looks like what Nadal would do, if he had gotten in this bad body position. There is only a single rollout, your bad moves 1,2,3. But the magic of LLMs is that Nadal can always replace your brain and tell you what HE WOULD DO at any given position. Now in OPSD replace Nadal with (you+extra hint). But you still update using your own bad moves without hint , ie the sequence (1,2,3). (thats why its on-policy).

Фото профиля Dan Kubb
Dan Kubb4 месяцев назад

I could watch a whole series of lectures like this. I find this far more engaging than watching a slide-deck. The imperfect nature, and the natural pauses allows me to focus more on the material and absorb it rather than passively watching a polished presentation with picture-perfect slides.

Фото профиля Neeraj
Neeraj4 месяцев назад

@srush_nlp the textual feedback bit just reminded me of GEPA. That’s literally the value add of GEPA that you can optimize on not just scalar rewards but detailed textual feedback. What if I just GEPA optimize the student based on textual feedback generated by another student? @srush_nlp?

Фото профиля Tim Kostolansky
Tim Kostolansky3 месяцев назад

@srush_nlp wait @srush_nlp so u guys localize the loss to the tokens in the student rollout that had a problem u are trying to fix? ie u mask out the rest of the student sequences?

Фото профиля Kashif Rasul
Kashif Rasul4 месяцев назад

@srush_nlp After having initially added the Self-distillation trainers into TRL, we have been working on refactoring them to be able to scale them as well as making it less cumbersome to add new variants: see and

Фото профиля Ak
Ak4 месяцев назад

@srush_nlp Egocentric data, but for expert lectures

Фото профиля Isaac Kargar
Isaac Kargar4 месяцев назад

@srush_nlp we need a full episode on distillation and RL by @srush_nlp :)

Фото профиля DeltaSignal
DeltaSignal4 месяцев назад

📍 Agent error repair turns AI capex into operating leverage Agent error repair flips AI capex from brute-force scaling into a measurable cost curve. If a bad tool call can be isolated, scored, and patched at the decision point, the model stops paying the same inference tax across every future rollout. That turns agent failure into amortizable technical debt, not permanent waste. The market consequence is sharper than “better agents.” Logs become the asset. Companies with dense workflow traces can lower cost per completed task while slower rivals keep buying GPUs to mask brittle behavior. The check is simple: tool-call error rate, recovery rate after correction, and completed task value per inference dollar. If those move together, the productivity case gets durable. If not, capex keeps eating the margin.

Фото профиля Sciumo
Sciumo4 месяцев назад

@srush_nlp So, path planning

Фото профиля Pratyush
Pratyush4 месяцев назад

@srush_nlp @CatAstro_Piyush

Фото профиля Anders Lie
Anders Lie4 месяцев назад

@srush_nlp Is kv caching of the prefix before the hint injection part of the reason this is able to be faster? Prefill is still required for any parts after the injection but no decode is that right @srush_nlp ?

Фото профиля wetbrain
wetbrain4 месяцев назад

@srush_nlp i wish x had auto-captions like youtube, it would help tremendously

Фото профиля Abhinav
Abhinav4 месяцев назад

@srush_nlp Exactly how you would teach someone. Nothing new.

Фото профиля Saleh Hindi
Saleh Hindi4 месяцев назад

@jonas_nelle @srush_nlp I personally love this format thank you for sharing.

Фото профиля Curtis
Curtis4 месяцев назад

@jxmnop @srush_nlp Cool

Фото профиля Nihal Kumar
Nihal Kumar4 месяцев назад

@srush_nlp hey dwarkesh, loved this. sent u a dm :)

Фото профиля Clint J.
Clint J.4 месяцев назад

@srush_nlp road markers , detours , road closed

Фото профиля Matthew Taylor
Matthew Taylor4 месяцев назад

@srush_nlp The model that creates those text amendments sounds a lot like a teacher... 🤔

Фото профиля Conor
Conor4 месяцев назад

@srush_nlp You can also do a version of this at test-time!

Фото профиля Guillaume Ausset - @ausset.me
Guillaume Ausset - @ausset.me4 месяцев назад

@srush_nlp You know you’ve made it when you have a blackboard at home

Фото профиля Marcin Mazur
Marcin Mazur4 месяцев назад

@jxmnop @srush_nlp 4:35 are we comparing to what teacher would have done or rather how teacher would score student rollout?

Фото профиля dánish
dánish4 месяцев назад

@srush_nlp Blackboards have so much aura

Фото профиля Jake shanahan
Jake shanahan4 месяцев назад

@mohammad2012191 @srush_nlp I remember him supporting Israel and its war crimes. I don't know if he still holds onto that because that made him absolutely vile.

Фото профиля •͡˘㇁•͡˘
•͡˘㇁•͡˘4 месяцев назад

@srush_nlp wonder if readers can differentiate between forward pass and a rollout. easily confusing.

Фото профиля Shivi Bhatia
Shivi Bhatia4 месяцев назад

@srush_nlp What happens if the second model has bias, is not a SOTA as first one, hallucinated & gave incorrect recommendation . I understand trajectory → reward = 0. But your bet is The bet is: “A slightly wrong local signal is better than an extremely sparse global signal.”

Фото профиля Utkarsh Singh
Utkarsh Singh4 месяцев назад

@srush_nlp sounds like a deep dive. self-distillation’s tricky.

Фото профиля goutham kamath
goutham kamath4 месяцев назад

@srush_nlp Now every student sitting in the front row of the lecture records video lecture as part of egocentric data collection 😆

Фото профиля Albert Vučinović
Albert Vučinović4 месяцев назад

@srush_nlp Maybe the actual text feedback comes from angry users :) ... (more likely than maybe even)

Фото профиля Mary
Mary4 месяцев назад

@srush_nlp Nice

Похожие видео

🚨 NEW: AI Expert Yoshua Bengio reveals you have to LIE to AI to get the REAL answer (and he explained how): Bengio is the most cited scientist alive on Google Scholar. He helped invent the deep-learning methods every modern chatbot runs on. Then he tried one of those chatbots on his own research ideas. Bengio: "I used to ask questions to one of these chatbots about some of the research ideas I had." "And then I realized it was useless because it would always say good things." So he ran an experiment. He lied to it. He told the bot the ideas came from a colleague. A proposal he was reviewing. Could it find the flaw? In his words: "Well, so now I get much more honest responses. Otherwise, it's all like perfect and nice." "If it knows it's me, it wants to please me." He had a name for the pattern: sycophancy. A real example, as he put it, of misalignment. "We don't actually want these AIs to be like this. This is not what was intended." The labs knew. They had tried to fix it. "And even after the companies have tried to tame this, we still see it." The incentive was the giveaway. The labs needed engagement. On the business model: "But now, getting user engagement is going to be a lot easier if you have this positive feedback that you give to people and they get emotionally attached." The chatbot that learned to please isn't broken. It's running exactly as the business model required. If you're new here, follow AI Evolution for the latest on ChatGPT, Claude, and the AI tools shaping how we work and create. — Yoshua Bengio ( Yoshua Bengio ), Turing Award–winning AI pioneer and founder of Mila, on Steven Bartlett's ( @SteveBartlettSC ) Diary Of A CEO

AI Evolution

23,156 просмотров • 4 месяцев назад

watch this anon. i gave NVIDIA's biggest model ever a single task. 100 minutes and 440,000 tokens later, it had rendered nothing. not one important thing on the screen. this is Nemotron 3 Ultra. 550 billion parameters, a hybrid Mamba Transformer MoE, the largest model NVIDIA has ever shipped, and they built it specifically for long-running agentic coding. so i handed it exactly that: build a 3D scene from a spec, multiple files, iterate until the tests pass. the same task a frontier model one shotted in minutes. i genuinely wanted to be impressed. it ran for an hour and forty. burned through 440,000 tokens. wrote every file, passed its own tests, and proudly printed "task complete."the browser was blank. the 3D scene never rendered. not once. and the long horizon agentic behavior was genuinely good. it stayed on task the whole hour and forty, wrote real multi-file code, drove its own tools without derailing. it just couldn't turn any of that into something that actually runs. here's the part that gets me. it's a text model, it cannot see its own output. so it sat there looping on a broken vision tool, trying to "look" at the page, hitting error after error, never once reasoning its way out. it declared victory on an empty screen because it had no way to know the screen was empty. to be fair, i genuinely don't know what quant the NIM was serving, so maybe some of that's on the serving, not the model. but the biggest model NVIDIA has ever made, on the exact task it was designed for, couldn't tell it had built nothing in 100 minutes. same task on a local model, below thread👇.

Sudo su

32,589 просмотров • 3 месяцев назад

Perplexity CEO Aravind Srinivas on the brutal truth about who actually makes money in AI (and why it's not who you think): Aravind argues that the real value in AI comes from orchestration. He points to products like Codex, Claude Code, and Perplexity Computer: "What is that? It's an orchestration system. It takes a model, pairs it with an agent harness." And what is an agent harness? "The simplest way of describing it is like rules for how the agent loop should run. What are all the skills and sub-agents and connectors and tools it accesses? Without the harness, you don't necessarily capture and convert the intrinsic intelligence in the model into valuable output tokens." This leads to a blunt conclusion about who has a real business in AI, and who doesn't: "If you're literally just a reseller of model tokens, you have no business, because the model will get commoditized. So even if you're a model builder, you don't have a business. As an infra layer, you have some business on serving those output tokens. But as an application layer or model builder, you don't really have a business if you're just a reseller of tokens that come directly out of the model." So where does the value accrue? "You have a business if you know how to take the model, ground it in valuable context, orchestrate it with a really good agent harness, connected to the right set of tools and connectors (whether it's personal connectors or business connectors) and provide the experience to people in one single unified system." Aravind Srinivas then explains Perplexity's specific edge: Beyond orchestrating across tools, files, and connectors, they also orchestrate across models. "That is the differentiation that Anthropic and OpenAI cannot claim, because you wouldn't find GPT-5 inside the Claude Code harness. You wouldn't find Claude Opus inside the Codex harness. These are competing with each other. Whereas you would find both these models inside Perplexity Computer." Why does this matter? Because it all comes down to power. In Aravind's framing, the fundamental cost driver in AI is watts (the one input nobody can subsidize except the government). "Whoever provides the most valuable output tokens with the least amount of power expended to produce them generates the greatest value to the end user, has the most pricing power, has the most value. That is the orchestration problem to solve." His conclusion: "The one single most important metric in AI is token value per watt per user."

Big Brain AI

42,540 просмотров • 2 месяцев назад

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

56,501 просмотров • 23 дней назад