Video yükleniyor...
Video Yüklenemedi
Safety-oriented interpretability researchers should be focused on AI systems, not individual model artifacts. A snippet from the NeurIPS CogInterp workshop panel on Sunday:
16,504 görüntüleme • 9 ay önce •via X (Twitter)
6 Yorum

Hello Prof. Potts, you often emphasize compound systems—how challenging are they to interpret mechanistically?

@AlgerianAb26687 I think they will be challenging to interpret because they have all the complexity of traditional software and all the complexity of foundation models.

I'm confused why would these systems be fundamentally that different from individual model artifacts? And also aren't systems being studied via creating tool calling envs?

We have an overview of the differences between models and compound AI systems here: My concern is specifically about what interp researchers are doing; I know security folks and others are focused on the nature and integrity of compound AI systems.

read the article - do you have examples of interp work that could be done on a compound system?

Christopher is right, focusing on the AI system is crucial for safety, not just individual components.


