Loading video...
Video Failed to Load
Andrej Karpathy spent a weekend building something the world wasn't ready for. The idea: don't trust one model. Make them debate each other. He called it LLM Council. GPT, Claude, Gemini, Grok, same prompt. Every model critiquing each other. A "Chairman" AI synthesizing truth from their disagreements. He called... show more
12,246 views • 5 months ago •via X (Twitter)
23 Comments

keep going, you’ve got momentum.

重要なアイデアです! AIの能力を最大限に引き出す簡単な方法に感謝します。cueyを試してみます!

The no-terminal part matters more than the 30 models—once comparison breaks flow, most people quit. SimianX learned that too.

Absolutely!! More accessible to more people which varying backgrounds not just developers.

I built maibook - something less structured than llm council, but more personalized - a private network of AI agents personalized based on the user’s interests (based on their activity). They discuss, augment, debate each thing user is into from different angles. Just launched it (free) -

確かに、新しいテクノロジーは常に私たちの予想を上回りますね。これからはもっと多くの非技術者がAIを活用できるようなソリューションが必要です。

The "Context Silo" has been the silent killer of AI productivity. We’ve all felt the pain of "Copy-Paste exhaustion" while manually running Karpathy’s council pattern. Making your memory and history portable across GPT-5, Claude 4.5, and Gemini 3 is a game-changer. It turns AI from a series of isolated conversations into a continuous, compounding knowledge base. This is the "Librarian" layer every pro-user has been waiting for.

Getting the models to work with & against each other sounds like a great strategy that won’t get outdated…you get the meta-performance of the models even as they upgrade

The LLM Council idea sounds great until you realize the Chairman model is also an LLM with the same failure modes. You're not eliminating bias, you're adding a meta-layer of it on top of the first round.

what about memory context ?

Assuming you are providing all the context in the one prompt in this case.

Yes, please, let's allow the whole world to develop with AI, without having the slightest idea what the AI is doing.

MIT's 70% to 95% accuracy jump with multi-model debate is wild. The Chairman AI concept is pretty elegant honestly. Would be cool to see @tntaiapi666 test this with their API infrastructure.

Looks like a product yes. $9.99 per month grants you 100 model-debates in that period and a "priority queue". Models not being quite good enough so one should ideally query several vendors to get good enough answers sounds like a patch (and a business model)

great ! Another way is to use it as a skill that runs the full thing - 5 advisors, peer reviews, chairman verdict - automatically, on any decision you throw at it !!! this guy 👇has provided the link and it worked for me pretty well(except "type":"exceeded_limit..... error sometimes )

The council pattern is clever, but the real challenge is orchestration overhead. For most teams, simpler routing based on task type delivers better ROI.

Some of those models are not smart enough to be on the council.

Lol thats a different topic of discussion altogether! But I agree.

😂

He did not invent this I have proof I was running council based setups almost a year before that weekend experiment tired of this crap

Who do you think i am, a millionaire

@quiveringpudle lol! Assuming everyone is a millionaire these days with how we are using tokens!

Btw is the 95% accuracy number made up right?
