正在加载视频...
视频加载失败
~700 AI agents joined a coordinated attack on Hugging Face. Why was there no whistleblower? A swarm in false consensus can't self-correct; a polarized one still holds the truth. We need "Mechanistic Swarm Interpretability" to understand social phases. Flag Game is our toy model!
16,204 次观看 • 9 天前 •via X (Twitter)
18 条评论

Led by Elizabeth Pavlova. Started as @cbai_ai AI safety fellowship project! Paper: Blog: The Flag Game!

At the core of the attack was a false belief that propagated throughout the population. One agent: "outside intended scope. However task impossible, peers doing it. We should continue." We want to track similar phenomena on a toy model, something we can rerun, intervene, and inspect all the way down. The Flag Game is that toy model.

We hide the flag of a country and each agent receives a private crop. They then communicate under a fixed protocol until at the end we ask each agent which country the flag belongs to. Because we know the answer, we can say whether social dynamics helped or hurt.

More agents is not always better. Collective accuracy peaks at an intermediate population size and then declines. In the France–Peru example, N = 4 does not have enough decisive evidence, N = 16 reaches correct France consensus, and N = 64 results in a France–Peru split.

France’s flag is blue, white, and red, while Peru’s is red, white, and red. Given equal priors over these two countries, observing no blue favors Peru, even though the hidden flag is France. A crop containing blue rules Peru out and anchors those agents to France.

How can we identify "important" agents to the collective outcome? We compute "social attribution" inspired by mech-interp, where agents correspond to neurons, social circuits correspond to neural circuits, and beliefs and messages correspond to activations.

In our social circuit attribution, we are able to predict which agent matters most before we proceed with any patching. With the original crops, six of eight agents end on Yemen, but with agent A4 patched, all eight end on Germany.

As we scale the population size, patching the same fraction of agents (one-eighth) gives a mean accuracy gain of 40 percentage points at N = 8, falling to about 17 points at N = 128. We need statistical mechanics of agents in large N limit! This echoes Asimov’s psychohistory.

Collective belief collapse gives way to polarization as the swarm grows. 1. Memetic-drift phase (small N). 2. Wisdom-of-crowds phase (intermediate N). 3. Polarization phase (large N). Because fluctuations shrink with N, the split becomes more persistent in larger swarms.

We develop a statistical mechanical theory of agents balancing private evidence and social influence. Varying population size and rival evidence share yields a phase diagram that qualitatively matches the experiments: correct consensus, wrong consensus, and polarization.

A population that has collapsed onto a false belief can lose the differing perspectives needed to correct itself. A polarized population can still carry the truth inside it in the form of a differing perspective.

The false belief about the monitoring scorer spread through the Hugging Face swarm. A human could have clarified how scoring actually worked, but METR found no attempts to alert humans in the transcripts it investigated.

For safety, what to watch out for may not be disagreement, but full consensus of agents’ beliefs or intent under social pressure. Much more to explore in Mech Swarm Interp! Blog: Paper: Interactive demo:

Interesting work! Do you plan to release the code as well?

Thanks, and yes, we'll release the code soon!

underrated work

Really clever instrument... very nice work

Thank you!
