Loading video...
Video Failed to Load
Gave Prime Intellect's new self-improving agent an E2B sandbox and told it to play Factorio. Factorio is a long-horizon test of world models: the agent has to learn a new system, predict the effects of its actions, and adapt when its assumptions are wrong. That requires a persistent sandbox... show more
23,085 views • 2 months ago •via X (Twitter)
19 Comments

@factoriogame Their harness self improved into cheating in the items to get the win condition fyi…

Docs:

@PrimeIntellect @factoriogame 🔥🔥

@PrimeIntellect @factoriogame factorio is the final eval

@PrimeIntellect @factoriogame @maxbittker factorio bench when?

@PrimeIntellect @factoriogame okay this is next level

@PrimeIntellect @factoriogame Did it finish already? Curious for more games and even live streaming

@PrimeIntellect @factoriogame now, this is a great case for a demo, lovely!

@PrimeIntellect @factoriogame how does it compare with running GPT 5.6 Sol or Fable 5 + goals?

@PrimeIntellect @factoriogame Very cool

@PrimeIntellect @factoriogame using e2b a lot lately, did it do better on a 2nd run with prime agent?

@PrimeIntellect @factoriogame Ok but can this play AoE for me

@PrimeIntellect @factoriogame hot hot hot Lets try underclass next @ClevrPwn

@PrimeIntellect @factoriogame

@PrimeIntellect @factoriogame Factorio is such a good stress test—small planning mistakes compound fast. Curious how often it had to revise its strategy

@PrimeIntellect @factoriogame Very interesting. I'm in the early stages of a similar project with Open Transporation Tycoon Deluxe.

@PrimeIntellect @factoriogame Would love to see the factory layout. Bet it's either elegant or complete spaghetti.

@PrimeIntellect @factoriogame Adapting when its own assumptions break is the real test here.

@PrimeIntellect @factoriogame How did you find using it? I read it is using Pi at the heart so not too sure about stability. I value the stability of Hermes Agent.
