
Goodfire
@GoodfireAI • 25,872 subscribers
Using interpretability to understand, learn from, and design AI.
Shorts
Videos

> replicate J-space on GLM 5.2 > train a reward model and run RL to reduce hallucinations > show me how this model makes cancer predictions Using our platform Silico is like having a team of AI researchers ready to run experiments like these. Private beta is open now. 🧵 (1/6)
Goodfire266,794 次观看 • 17 天前

Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9)
Goodfire183,872 次观看 • 1 个月前

Check out Atticus Geiger's Stanford guest lecture - on causal approaches to interpretability - for an overview of one of our areas of research! 01:51 - Activation steering (e.g. Golden Gate Claude) 10:23 - Causal mediation analysis (understanding the contribution of an intermediate component) 21:42 - Causal abstraction methods (explaining a complex causal system with a simple one) 54:54 - Lookback mechanisms: a case study in designing counterfactuals This is the first of three guest lectures we'll be posting from Surya Ganguli's course.
Goodfire36,522 次观看 • 8 个月前

Our last Stanford guest lecture - Ekdeep Singh Lubana on what counts as an explanation & a neuro-inspired "model systems approach" to interp Plus, how in-context learning and many-shot jailbreaking are explained by LLM representations changing in-context (as a case study for that approach) 00:33 - What counts as an explanation? 04:47 - Levels of analysis & standard interpretability approaches 18:19 - The "model systems" approach to interp [Case study on in-context learning] 23:36 - How LLM representations change in-context 44:10 - Modeling ICL with rational analysis 1:10:54 - Conclusion & questions Thanks again to Surya Ganguli for having us in his class!
Goodfire31,410 次观看 • 7 个月前

Today, we’re releasing our research preview ( to let you look inside your AI. We've created a desktop interface that helps you understand and control Llama 3's behavior. You can 1) see Llama 3's internal features (the internal building blocks of its responses) and 2) precisely adjust these features to create new Llama variants. Try it out and share your findings with #GoodfireAI
Goodfire67,161 次观看 • 1 年前

Another Stanford interpretability guest lecture: Jack Merullo on "computational motifs" - the algorithmic primitives of transformers that show up again and again across circuits/tasks/models e.g. induction heads, binding vectors, helical representation comparisons, copy suppresion heads, etc. 00:53 - Intro: defining "computational motifs" 05:48 - Induction heads (a classic motif) 08:31 - Motifs in the Indirect Object Identification circuit 44:33 - More examples 51:15 - Challenges and open problems 1:03:12 - Conclusion & questions
Goodfire16,634 次观看 • 7 个月前
没有更多内容可加载