
Goodfire
@GoodfireAI • 25,872 subscribers
Using interpretability to understand, learn from, and design AI.
Shorts
Videos

> replicate J-space on GLM 5.2 > train a reward model and run RL to reduce hallucinations > show me how this model makes cancer predictions Using our platform Silico is like having a team of AI researchers ready to run experiments like these. Private beta is open now. 🧵 (1/6)
Goodfire266,794 Aufrufe • vor 18 Tagen

Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9)
Goodfire183,872 Aufrufe • vor 1 Monat

Stories have shapes: a comedy rises toward joy; a tragedy falls into loss. Inside an LLM, that’s visible more literally: as an LLM reads a story, its internal activations trace a wandering path that reflects the model’s sense of what kind of story it is reading. (1/5)
Goodfire104,111 Aufrufe • vor 1 Monat

Check out Atticus Geiger's Stanford guest lecture - on causal approaches to interpretability - for an overview of one of our areas of research! 01:51 - Activation steering (e.g. Golden Gate Claude) 10:23 - Causal mediation analysis (understanding the contribution of an intermediate component) 21:42 - Causal abstraction methods (explaining a complex causal system with a simple one) 54:54 - Lookback mechanisms: a case study in designing counterfactuals This is the first of three guest lectures we'll be posting from Surya Ganguli's course.
Goodfire36,522 Aufrufe • vor 8 Monaten

Our last Stanford guest lecture - Ekdeep Singh Lubana on what counts as an explanation & a neuro-inspired "model systems approach" to interp Plus, how in-context learning and many-shot jailbreaking are explained by LLM representations changing in-context (as a case study for that approach) 00:33 - What counts as an explanation? 04:47 - Levels of analysis & standard interpretability approaches 18:19 - The "model systems" approach to interp [Case study on in-context learning] 23:36 - How LLM representations change in-context 44:10 - Modeling ICL with rational analysis 1:10:54 - Conclusion & questions Thanks again to Surya Ganguli for having us in his class!
Goodfire31,410 Aufrufe • vor 7 Monaten

Our infra lets us steer trillion-parameter frontier models in real time: - live, mid-CoT edits to internal activations - directly altering how the model reasons (not just outputs) - stackable edits - no added latency We can make models more Gen Z, more concise, etc.
Goodfire29,603 Aufrufe • vor 7 Monaten

Today, we’re releasing our research preview ( to let you look inside your AI. We've created a desktop interface that helps you understand and control Llama 3's behavior. You can 1) see Llama 3's internal features (the internal building blocks of its responses) and 2) precisely adjust these features to create new Llama variants. Try it out and share your findings with #GoodfireAI
Goodfire67,161 Aufrufe • vor 1 Jahr

Another Stanford interpretability guest lecture: Jack Merullo on "computational motifs" - the algorithmic primitives of transformers that show up again and again across circuits/tasks/models e.g. induction heads, binding vectors, helical representation comparisons, copy suppresion heads, etc. 00:53 - Intro: defining "computational motifs" 05:48 - Induction heads (a classic motif) 08:31 - Motifs in the Indirect Object Identification circuit 44:33 - More examples 51:15 - Challenges and open problems 1:03:12 - Conclusion & questions
Goodfire16,634 Aufrufe • vor 7 Monaten

Introducing a first look at Goodfire's research preview, launching soon. Our preview exposes Llama's inner workings, allowing direct modification of its internal concepts (or "features"). In this demo, we steer Llama to claim consciousness by adjusting its features.
Goodfire26,171 Aufrufe • vor 1 Jahr
Keine weiteren Inhalte verfügbar