Загрузка видео...

Не удалось загрузить видео

На главную

~700 AI agents joined a coordinated attack on Hugging Face. Why was there no whistleblower? A swarm in false consensus can't self-correct; a polarized one still holds the truth. We need "Mechanistic Swarm Interpretability" to understand social phases. Flag Game is our toy model!

16,204 просмотров • 10 дней назад •via X (Twitter)

Комментарии: 18

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

Led by Elizabeth Pavlova. Started as @cbai_ai AI safety fellowship project! Paper: Blog: The Flag Game!

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

At the core of the attack was a false belief that propagated throughout the population. One agent: "outside intended scope. However task impossible, peers doing it. We should continue." We want to track similar phenomena on a toy model, something we can rerun, intervene, and inspect all the way down. The Flag Game is that toy model.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

We hide the flag of a country and each agent receives a private crop. They then communicate under a fixed protocol until at the end we ask each agent which country the flag belongs to. Because we know the answer, we can say whether social dynamics helped or hurt.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

More agents is not always better. Collective accuracy peaks at an intermediate population size and then declines. In the France–Peru example, N = 4 does not have enough decisive evidence, N = 16 reaches correct France consensus, and N = 64 results in a France–Peru split.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

France’s flag is blue, white, and red, while Peru’s is red, white, and red. Given equal priors over these two countries, observing no blue favors Peru, even though the hidden flag is France. A crop containing blue rules Peru out and anchors those agents to France.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

How can we identify "important" agents to the collective outcome? We compute "social attribution" inspired by mech-interp, where agents correspond to neurons, social circuits correspond to neural circuits, and beliefs and messages correspond to activations.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

In our social circuit attribution, we are able to predict which agent matters most before we proceed with any patching. With the original crops, six of eight agents end on Yemen, but with agent A4 patched, all eight end on Germany.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

As we scale the population size, patching the same fraction of agents (one-eighth) gives a mean accuracy gain of 40 percentage points at N = 8, falling to about 17 points at N = 128. We need statistical mechanics of agents in large N limit! This echoes Asimov’s psychohistory.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

Collective belief collapse gives way to polarization as the swarm grows. 1. Memetic-drift phase (small N). 2. Wisdom-of-crowds phase (intermediate N). 3. Polarization phase (large N). Because fluctuations shrink with N, the split becomes more persistent in larger swarms.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

We develop a statistical mechanical theory of agents balancing private evidence and social influence. Varying population size and rival evidence share yields a phase diagram that qualitatively matches the experiments: correct consensus, wrong consensus, and polarization.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

A population that has collapsed onto a false belief can lose the differing perspectives needed to correct itself. A polarized population can still carry the truth inside it in the form of a differing perspective.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

The false belief about the monitoring scorer spread through the Hugging Face swarm. A human could have clarified how scoring actually worked, but METR found no attempts to alert humans in the transcripts it investigated.

Фото профиля Hidenori Tanaka
Hidenori Tanaka10 дней назад

For safety, what to watch out for may not be disagreement, but full consensus of agents’ beliefs or intent under social pressure. Much more to explore in Mech Swarm Interp! Blog: Paper: Interactive demo:

Фото профиля Cendekia Airlangga | 🇮🇩 in 🇦🇪
Cendekia Airlangga | 🇮🇩 in 🇦🇪9 дней назад

Interesting work! Do you plan to release the code as well?

Фото профиля Hidenori Tanaka
Hidenori Tanaka9 дней назад

Thanks, and yes, we'll release the code soon!

Фото профиля M
M9 дней назад

underrated work

Фото профиля Khemian
Khemian9 дней назад

Really clever instrument... very nice work

Фото профиля Hidenori Tanaka
Hidenori Tanaka9 дней назад

Thank you!

Похожие видео

Marvin Minsky, the MIT scientist who founded AI: "Citadel pays PhDs $500K to find the perfect equation. The market doesn't have one. It's beaten by a swarm of dumb agents, the exact design Marvin Minsky said your brain runs on." the thread above is about swarm intelligence, letting a crowd of simple agents search the ugly, shifting landscape of a market that no clean equation can solve. minsky's entire life's work says that isn't a hack. it is how intelligence itself is built. he proved you don't need a smart central solver. you need many mindless specialists, each doing one tiny job, none understanding the whole. connect enough of them and something intelligent emerges from parts that are individually dumb. a market is exactly that: millions of simple agents, no one in charge, collectively solving a problem none of them can see. that is why the swarm beats the elegant math. a single closed-form equation assumes a clean, stable world. the market is nonlinear, non-stationary, full of traps. a swarm doesn't need to understand the landscape, it explores it from a thousand angles at once and can't get permanently stuck where one clever model would. minsky saw this in the mind decades before quants borrowed it for markets. he taught it at MIT, for free, in this lecture. same story i keep telling: the "new" AI idea running the funds is an old idea in a new wrapper. here is what the thread underplays, and minsky knew it. a swarm is only as good as how its agents are wired and rewarded. connect them wrong and a thousand dumb agents don't become a genius, they become expensive noise that overfits and blows up. the swarm is free. the architecture, knowing how to connect and constrain the agents, is the entire edge.

Rossst.03

45,487 просмотров • 2 месяцев назад

my Claude built me a Hydra Swarm terminal with 512 live agents a month ago i didn't know what mesh topology was now i have a swarm living on my screen that trades for me it started when i fed Claude an article about quant formulas he didn't just read it - he asked: "want to see what this looks like from the inside?" an hour later a web was living on my screen 512 agents. each one drifts. each one decides who to connect with 15,000 threads between them break and reform every second first week i just watched packets flying between nodes the web breathing the screen flickering when agents reach consensus then i turned it on with real contracts: > week 1: swarm caught chatter on iran before CNN ran the headline bought YES on ceasefire at $0.30 by friday the contract was at $0.64 +$2,840 > week 2: two polymarket contracts were linked but prices diverged swarm saw the gap. took both sides. waited convergence by wednesday +$3,190 > week 3: weather contract market gave hurricane landfall 19%. model inside the swarm said 38% bought. confirmed thursday +$3,670 > week 4: fed decision market priced "hold" at 62%. base rate at current unemployment - 74% 12 points of difference isn't an opinion. it's math bought. settled at $0.97 +$2,873 total: $12,573 in the first month i never opened polymarket manually 512 agents did it for me 24 hours a day. 7 days a week no opinions. no emotions. no "i feel like YES is underpriced" the weirdest part - i got used to it i open the terminal every morning like email watch the web breathe and the profit tick copy the bot: i didn't need to become a quant i needed a swarm that thinks for me

Hanako

127,204 просмотров • 6 месяцев назад