Loading video...

Video Failed to Load

Go Home

Gave Prime Intellect's new self-improving agent an E2B sandbox and told it to play Factorio. Factorio is a long-horizon test of world models: the agent has to learn a new system, predict the effects of its actions, and adapt when its assumptions are wrong. That requires a persistent sandbox...

23,085 views • 2 months ago •via X (Twitter)

19 Comments

Teknium 🪽's profile picture
Teknium 🪽2 months ago

@factoriogame Their harness self improved into cheating in the items to get the win condition fyi…

Vasek Mlejnsky's profile picture
Vasek Mlejnsky2 months ago

Docs:

Oscar Moxon's profile picture
Oscar Moxon2 months ago

@PrimeIntellect @factoriogame 🔥🔥

michael froehlich's profile picture
michael froehlich2 months ago

@PrimeIntellect @factoriogame factorio is the final eval

Sawyer Hood's profile picture
Sawyer Hood2 months ago

@PrimeIntellect @factoriogame @maxbittker factorio bench when?

Mish Ushakov's profile picture
Mish Ushakov2 months ago

@PrimeIntellect @factoriogame okay this is next level

Diego Alejandro's profile picture
Diego Alejandro2 months ago

@PrimeIntellect @factoriogame Did it finish already? Curious for more games and even live streaming

Filip Sanda's profile picture
Filip Sanda2 months ago

@PrimeIntellect @factoriogame now, this is a great case for a demo, lovely!

Federico Ulfo's profile picture
Federico Ulfo2 months ago

@PrimeIntellect @factoriogame how does it compare with running GPT 5.6 Sol or Fable 5 + goals?

homanp's profile picture
homanp2 months ago

@PrimeIntellect @factoriogame Very cool

Revofusion's profile picture
Revofusion2 months ago

@PrimeIntellect @factoriogame using e2b a lot lately, did it do better on a 2nd run with prime agent?

Mikayel's profile picture
Mikayel2 months ago

@PrimeIntellect @factoriogame Ok but can this play AoE for me

Roko ʕ •ᴥ•ʔっ🪄✨🐍's profile picture
Roko ʕ •ᴥ•ʔっ🪄✨🐍2 months ago

@PrimeIntellect @factoriogame hot hot hot Lets try underclass next @ClevrPwn

llevy's profile picture
llevy2 months ago

@PrimeIntellect @factoriogame

AI Alpha By Mario's profile picture
AI Alpha By Mario2 months ago

@PrimeIntellect @factoriogame Factorio is such a good stress test—small planning mistakes compound fast. Curious how often it had to revise its strategy

Dylan Fetch's profile picture
Dylan Fetch2 months ago

@PrimeIntellect @factoriogame Very interesting. I'm in the early stages of a similar project with Open Transporation Tycoon Deluxe.

Sentient Yoghurt's profile picture
Sentient Yoghurt2 months ago

@PrimeIntellect @factoriogame Would love to see the factory layout. Bet it's either elegant or complete spaghetti.

AI Mastery Guide's profile picture
AI Mastery Guide2 months ago

@PrimeIntellect @factoriogame Adapting when its own assumptions break is the real test here.

Anders's profile picture
Anders2 months ago

@PrimeIntellect @factoriogame How did you find using it? I read it is using Pi at the heart so not too sure about stability. I value the stability of Hermes Agent.

Related Videos

Introducing Headlong, an open source microharness for persistent agents: self-guided agents that think continuously. Most agent harnesses are reactive: you send a task, the agent completes it, and then it sits frozen until the next request. Cron jobs and heartbeats wake it up to run a checklist and put it back to sleep. A Headlong agent is never asleep. It keeps generating thoughts about whatever it decides is interesting, in a self-guided loop inspired by human inner monologue. Your message doesn't start a session. It's one more observation that lands in the agent's thought stream, and the agent decides if and when to reply. Headlong is built on the idea of persistent agency: continuous inner thought generation between external interactions. The agent sets its own interests and priorities, comes up with its own projects, and sometimes pings you unprompted with progress. To keep our prototype as simple and small as possible, we implemented Headlong as a microharness: a complete agent harness in under 10K lines of Bash, organized as a handful of small executables. It includes a loop that generates the next thought, shellm (a recursive language model written in Bash), a trajectory stored as a DAG of jsonl files, and context as a projection of that trajectory. We've been running one Headlong agent internally at Laude for several weeks. The whole team talks to it over Slack and Telegram, and every conversation lands in its single stream of thought. It works in its own fork of Headlong and we've pulled over 50 of its commits into main. One night, with nobody talking to it, it went back to check whether a recall process it had built was actually wired into its mind, found that it wasn't, diagnosed and fixed the bug, and verified the fix end to end. 48 minutes, no human asked for the fix or was in the loop at any point. Every step is a timestamped line in its log. Things broke too, and we wrote those up. Background thinking costs us $1 to $2 an hour, our agent stopped its own service three times by accident, and self-delegation died on day one. Details in the post. One line installs everything and starts an agent. Use a dedicated sandbox and spend-capped API key; it runs real shell commands and thinks around the clock. Headlong is research software, be careful! curl -fsSL | bash Launch post: Repo: Headlong is a Laude Institute / MIT collaboration.

Andy Konwinski

361,750 views • 1 month ago

New short course: Long-Term Agentic Memory with LangGraph. Learn to build an agent with long-term memory in this course developed in collaboration with taught by its Co-Founder and CEO, Harrison Chase! Personal assistance and productivity tasks have become important use cases for agents. An important feature of an AI assistant, such as a coding or calendar assistant, is its ability to keep improving over time from its experience. Agent memory is the key capability that enables this. To add memory to an agent, you must first figure out what to store and what to retrieve when it is time to use the information. Additionally, you’ll have to decide when to update the stored information. For example, you might update in each iteration loop of the agent or perform updates in the background, with a helper agent. In this course, you will learn a mental framework to build agents with long-term memory. You'll create a useful email assistant that can respond, ignore, and notify using writing, scheduling, and memory-management tools. You’ll develop your agent's memory by adding facts to its memory store, provide examples to learn the user's preferences, and optimize system prompts to evolve instructions based on previous responses. In detail, you’ll: - Learn how the three types of memory--semantic, episodic, and procedural–and the two update mechanisms–via hot path and in the background–apply to your agents. - Build an email agent with writing, scheduling, and availability tools, along with a router that triages incoming email and handles it accordingly by ignoring, responding, or notifying the user. - Add tools to your email agent that allow it to operate on semantic memory by learning facts about the user, storing them in a long-term memory store, and searching over them in future interactions. - Incorporate episodic memory, in the form of few-shot examples, in the triage step of your agents to help them learn and update user preferences. - Add procedural memory as system prompts, optimized with feedback to improve the instructions the agent follows. Learn how to approach memory in agents, and start building agents with long-term memory with LangGraph! Please sign up here:

Andrew Ng

132,058 views • 1 year ago