Loading video...
Video Failed to Load
Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games. We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon. On simpler games, agents get close to the human world... show more
17,716 views • 17 days ago •via X (Twitter)
14 Comments

We set up a demo at where you can pick any model and any game and play them in the browser.

Most game benchmarks measure “can the agent complete the game”? That metric basically saturates the moment frontier models complete the game. Human speedrunning communities have spent decades pushing the same games well past completion by discovering new routes and techniques, which makes frame count to the milestone of a game a natural objective.

We find that agents often inherit a slower route than the best human players, and route freely when not. Additionally, giving a model a seed makes the task tractable for weaker agents, since several models failed to reach the goal from scratch at all and many models cannot reach the goal on more complicated games like Pokémon. More at:

Website: Paper: Blog: Benchmark:

If problems like this are interesting to you, we're hiring:

Great benchmark! Thanks for mentioning pokeagent in the paper.

We love Pokeagent!

I care more about the clock after the first mistake. A clean run measures planning; recovery speed tells me whether the agent can operate when the plan breaks.

Speed instead of completion is the right axis — completion saturates the day it's cleared. One asymmetry in the 4x: a world record is a maximum over decades of attempts by thousands of runners. Is the agent's time a best-of-N or a mean, and what is N?

Speedrunning is a fascinating benchmark for agent intelligence.

This is so sick!

SpeedrunBench across 10 titles, including Mario Kart and Pokemon, is a useful test. Measuring agent speed, not completion, shows careful benchmark design. Follow back if open; X needs it before my DM with a handful-only offer, examples, and quote for patronus . ai

This is awesome.

speedrun routes are public and they have dates on them, so a chunk of this tracks training cutoff rather than planning. a model trained a year later inherits a better route for nothing. games whose record route landed after every cutoff would separate the two

