Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Safety-oriented interpretability researchers should be focused on AI systems, not individual model artifacts. A snippet from the NeurIPS CogInterp workshop panel on Sunday:

16,504 görüntüleme • 9 ay önce •via X (Twitter)

6 Yorum

‏عبد الهادي‏ profil fotoğrafı
‏عبد الهادي‏9 ay önce

Hello Prof. Potts, you often emphasize compound systems—how challenging are they to interpret mechanistically?

Christopher Potts profil fotoğrafı
Christopher Potts9 ay önce

@AlgerianAb26687 I think they will be challenging to interpret because they have all the complexity of traditional software and all the complexity of foundation models.

arya profil fotoğrafı
arya9 ay önce

I'm confused why would these systems be fundamentally that different from individual model artifacts? And also aren't systems being studied via creating tool calling envs?

Christopher Potts profil fotoğrafı
Christopher Potts9 ay önce

We have an overview of the differences between models and compound AI systems here: My concern is specifically about what interp researchers are doing; I know security folks and others are focused on the nature and integrity of compound AI systems.

arya profil fotoğrafı
arya9 ay önce

read the article - do you have examples of interp work that could be done on a compound system?

Himanshu Kumar profil fotoğrafı
Himanshu Kumar9 ay önce

Christopher is right, focusing on the AI system is crucial for safety, not just individual components.

Benzer Videolar