Загрузка видео...

Не удалось загрузить видео

На главную

Safety-oriented interpretability researchers should be focused on AI systems, not individual model artifacts. A snippet from the NeurIPS CogInterp workshop panel on Sunday:

16,504 просмотров • 9 месяцев назад •via X (Twitter)

Комментарии: 6

Фото профиля ‏عبد الهادي‏
‏عبد الهادي‏9 месяцев назад

Hello Prof. Potts, you often emphasize compound systems—how challenging are they to interpret mechanistically?

Фото профиля Christopher Potts
Christopher Potts9 месяцев назад

@AlgerianAb26687 I think they will be challenging to interpret because they have all the complexity of traditional software and all the complexity of foundation models.

Фото профиля arya
arya9 месяцев назад

I'm confused why would these systems be fundamentally that different from individual model artifacts? And also aren't systems being studied via creating tool calling envs?

Фото профиля Christopher Potts
Christopher Potts9 месяцев назад

We have an overview of the differences between models and compound AI systems here: My concern is specifically about what interp researchers are doing; I know security folks and others are focused on the nature and integrity of compound AI systems.

Фото профиля arya
arya9 месяцев назад

read the article - do you have examples of interp work that could be done on a compound system?

Фото профиля Himanshu Kumar
Himanshu Kumar9 месяцев назад

Christopher is right, focusing on the AI system is crucial for safety, not just individual components.

Похожие видео