Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Safety-oriented interpretability researchers should be focused on AI systems, not individual model artifacts. A snippet from the NeurIPS CogInterp workshop panel on Sunday:

16,504 Aufrufe • vor 9 Monaten •via X (Twitter)

6 Kommentare

Profilbild von ‏عبد الهادي‏
‏عبد الهادي‏vor 9 Monaten

Hello Prof. Potts, you often emphasize compound systems—how challenging are they to interpret mechanistically?

Profilbild von Christopher Potts
Christopher Pottsvor 9 Monaten

@AlgerianAb26687 I think they will be challenging to interpret because they have all the complexity of traditional software and all the complexity of foundation models.

Profilbild von arya
aryavor 9 Monaten

I'm confused why would these systems be fundamentally that different from individual model artifacts? And also aren't systems being studied via creating tool calling envs?

Profilbild von Christopher Potts
Christopher Pottsvor 9 Monaten

We have an overview of the differences between models and compound AI systems here: My concern is specifically about what interp researchers are doing; I know security folks and others are focused on the nature and integrity of compound AI systems.

Profilbild von arya
aryavor 9 Monaten

read the article - do you have examples of interp work that could be done on a compound system?

Profilbild von Himanshu Kumar
Himanshu Kumarvor 9 Monaten

Christopher is right, focusing on the AI system is crucial for safety, not just individual components.

Ähnliche Videos