Loading video...
Video Failed to Load
🚨1 year ago today Sam Altman returned to OpenAI after a boardroom power struggle. Like many, we were blindsided: WHAT HAPPENED behind closed doors to the most successful entrepreneur of his generation? We built SIM-1 to find out (w help from GPT4-o1): what happened at OpenAI?🚨
31,294 views • 1 year ago •via X (Twitter)
18 Comments

Collaborating with our genius friends @TREE_Industries we built SIM-1 framework inspired by @polynoamial 'Human-level play in Diplomacy...' (image below) @demishassabis 2003 'Republic: The Revolution' simulation wargame @joon_s_pk seminal Generative Agents paper @DrJimFan work on 3D simulations @GuangyuRobert work on AI behavior in Simulations They are true seminal academic researchers, whereas we're just a small startup, but we look up to them!

We ran 20 AI simulations of what exactly happened 1 year ago today at OpenAI. Each agent, @satyanadella @DarioAmodei @pmarca @elonmusk @ilyasut @adamdangelo @miramurati etc - had 200+ turns across 15+ rounds and 5 days to: 1. Control OpenAI as new CEO or 2. Hire its staff or 3. Start a new, well-funded company in the chaos In this Simulation #11 video, GPT4o1 and SIM-1 chose @elonmusk as the new CEO of OpenAI in an ironic full circle moment.

The brilliant @CharlieFink @Forbes covered it today

As did the redoubtable @deantak at @VentureBeat and @GamesBeat

In each round, all agents take a turn and are able to use exhortation, inspiration and even deception to achieve their goals. Agents can call each other, release misinformation, try to manage public relations etc and an adjudicator agent analyzes the winner each round. In the final round, at the end of the 200+ turns, the adjudicator agent also selects the overall winner - which ranged widely.

Why build an AI framework that allows agents to deceive, be aggressive & ruthless? Our long term quest is to create a realistic AI person. Today, AI Agents trying to imitate humans are 'nice, but dull' - unrealistic to the messiness of human psychology. Agents today lack what Jung called 'The Shadow' - the darker side of the psyche. Many AI agents of course only fulfill tasks - they don't need to be realistic, but for our dream of an AI that can fully imitate a human, research should begin to approach the darker side of humanity - if it wants to be realistic. The light and dark in complementary relation are what make us human. We built SIM-1 to have a framework for more realistic AI people.

In trying to design a wargame that captured best current practices, we spoke to ex-officials at the Department of Defense who have participated in Pentagon Wargames. This gave us the adjudicator agent - an agent that decides round winners & keeps the wargame on track to conclusion, & even is able to introduce unexpected complexity to unsettle agents.

Our research was also inspired by @demishassabis Simulation videogame 'Republic', built shortly before Hassabis founded @GoogleDeepMind, an under-explored part of his story. 'Republic' the simulation was ahead of its time in 2003, but it showed the possibilities of a complex, aggressive simulated political wargame. Players took control of a labor organizer/journalist/officer and were able to use different game theory approaches to influence opinion and take over the fictional republic of Novistrana. It was limited by the technology of 2003 but the underlying principles were ones we were able to build on for this wargame.

The below diagram gives a sense of how the SIM-1 framework dynamically updates context each turn so that agents can react to events & each other's previous turns. SIM-1 understands the difference between private phonecalls (information only participants know) vs wider press statements and how rumors can spread. Each run of 20 rounds lasts for 4 hours - which can be a lot to consume, so every round SIM-1 also passes highlights to SHOW-1, which then generates a scene.

We used Showrunner and SHOW-1 to generate scenes from the simulation since ... watching 4 hours can get tedious! Here are highlights (this only shows a few turns out of the total 200) from one of the Simulations where Sam wins, passed to SHOW-1 to make a mini episode of the AITV show Exit Valley, set in Sim Francisco.

In our research, we also found some weaknesses in using LLMs for wargames. For wargames, personality is important, but a history of past decisions is even more important to predicting future behavior. A model trained exclusively on all the decisions an individual has made (rather than what they say about themselves) might lead to more accurate wargames in future and be a fruitful area for research.

SIM-1 also hallucinated multiple times - occasionally giving choices of unrealistic alliances, contradicting itself, or hallucinating that an agent currently worked at OpenAI when they were considering whether to join. For future simulations we are working to control these.

SIM-1 research has implications both for true wargames & for entertainment. In the area of geopolitical wargames, we are applying this research to a simulation of a wargame that will be on all our minds in the coming years : a potential Chinese invasion of Taiwan and the reaction of policymakers. Simulating key American players like a future President Trump, Secretary of State Marco Rubio set in Sim Washington DC SIM-1 would simulate Cai Qi (first ranked Secretary to the secretariat of the Communist Party), Xi Jinping, Dong Jun, Chinese Defense Secretary, Li Qiang Prime Minister Lai Ching-Te, leader of Taiwan, Japan’s leader Ishiba, PM Keir Starmer, French President Macron and President Putin and North Jorean leader Kim Jong Un and Elon Musk.

We also think there are implications in SIM-1 for entertainment. Last year, @gdb spoke about the possibilities of AI generating a 'do-over' of the final Season of Game of Thrones. The promise of SIM-1 would be that instead of humans manually generating a new season scene by scene with AI - AI would coherently simulate realistic decision-making by the characters in the war for the throne, foreshadowing decisions over a longer range than scene by scene and staying true to realistic character strategies.

The future of AI & media isn't text to video - but 'SIM-TO-VIDEO'. In other words - if you want AI to 'fix' the ending of Game of Thrones you need to build Westeros and populate with all the characters competing for the throne with the ability to exhort, manipulate, deceive, inspire and surprise. More 'Realistic characters' doesn't mean photoreal deepfake faces - rather, it means realistic in the way actors, novelists, playwrights build complex characters, able to execute long-range plans and react in realtime to others' deception in a fog of war.

We'll be releasing the paper itself to this thread and we would love ideas on where to go with it - we're not a team of academics but an applied research startup which has worked since our time working at Oculus on the problem of realistic AI characters. We built Showrunner to visualize what is happening in Simulations, and to allow users to direct agents. It's early days for that - but if you're excited for the journey sign up at Today is about Sim Francisco but for those who are curious to make scenes (rather than to build simulations) and who have signed up in the past check your email later today we've invited some of the earliest cohorts who signed up. Please remember this is still a CONFIDENTIAL Alpha.

For those who say "why focus on comedy - be serious!" we think the admonition in the great film 'The Manchurian Candidate' is instructive How to make realistic AI people?: "With humor, my dear Zilkov, always with a little humor" In the race to create realistic AI people we're betting that agents with humor, irony, and a hint of Jung's shadow will be more realistic - but let's see!

We are certain though that creating a realistic AI person will ultimately be a WORK OF ART as much as or more than a FEAT OF ENGINEERING. The artists, comedians, storytellers and those with a true sense of humor might just get to realistic 'AI people' sooner than the more dry approaches: Because sometimes "the funniest outcome is the most likely" as @elonmusk always reminds us!
