Loading video...

Video Failed to Load

Go Home

Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games. We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon. On simpler games, agents get close to the human world...

17,716 views • 17 days ago •via X (Twitter)

14 Comments

PatronusAI's profile picture
PatronusAI17 days ago

We set up a demo at where you can pick any model and any game and play them in the browser.

PatronusAI's profile picture
PatronusAI17 days ago

Most game benchmarks measure “can the agent complete the game”? That metric basically saturates the moment frontier models complete the game. Human speedrunning communities have spent decades pushing the same games well past completion by discovering new routes and techniques, which makes frame count to the milestone of a game a natural objective.

PatronusAI's profile picture
PatronusAI17 days ago

We find that agents often inherit a slower route than the best human players, and route freely when not. Additionally, giving a model a seed makes the task tractable for weaker agents, since several models failed to reach the goal from scratch at all and many models cannot reach the goal on more complicated games like Pokémon. More at:

PatronusAI's profile picture
PatronusAI17 days ago

Website: Paper: Blog: Benchmark:

PatronusAI's profile picture
PatronusAI17 days ago

If problems like this are interesting to you, we're hiring:

Seth Karten's profile picture
Seth Karten17 days ago

Great benchmark! Thanks for mentioning pokeagent in the paper.

PatronusAI's profile picture
PatronusAI17 days ago

We love Pokeagent!

Alex's profile picture
Alex17 days ago

I care more about the clock after the first mistake. A clean run measures planning; recovery speed tells me whether the agent can operate when the plan breaks.

IpezyGJ's profile picture
IpezyGJ12 days ago

Speed instead of completion is the right axis — completion saturates the day it's cleared. One asymmetry in the 4x: a world record is a maximum over decades of attempts by thousands of runners. Is the agent's time a best-of-N or a mean, and what is N?

RAZA | AI EXPLORER's profile picture
RAZA | AI EXPLORER17 days ago

Speedrunning is a fascinating benchmark for agent intelligence.

baba yaga's profile picture
baba yaga16 days ago

This is so sick!

Madison Guyen's profile picture
Madison Guyen15 days ago

SpeedrunBench across 10 titles, including Mario Kart and Pokemon, is a useful test. Measuring agent speed, not completion, shows careful benchmark design. Follow back if open; X needs it before my DM with a handful-only offer, examples, and quote for patronus . ai

newsmartttttttt's profile picture
newsmartttttttt16 days ago

This is awesome.

AI Quanting's profile picture
AI Quanting17 days ago

speedrun routes are public and they have dates on them, so a chunk of this tracks training cutoff rather than planning. a model trained a year later inherits a better route for nothing. games whose record route landed after every cutoff would separate the two

Related Videos

Today, we’re excited to announce our $50M Series B, led by Greenfield Partners, with participation from Lightspeed and Notable Capital. 🚀 At Patronus AI, we develop simulations and evals to train and improve AI. The first phase of AI was built on static benchmarks, but that era is over. As agents are used to solve longer and longer tasks, they need to practice in dynamic, living worlds to get better. Simulations are the critical infrastructure powering this next phase. As a company, we’re behind the most influential research and products in AI evaluation, like FinanceBench, Lynx, and Percival. And things have moved at the speed of light since.⚡ We partner with the world's leading frontier AI labs and enterprises, and our revenue has grown more than 15x over the past year. Additionally, today, we’re introducing a preview of the first Digital World Model for AI agent training and simulation: Patronus-DWM. Digital World Models are language diffusion world models that predict realistic environment behaviors and steer agent actions across digital workflows. Just as physical world models predict how objects move through space, we’re developing the equivalent for the digital world: predicting how agents act in digital workflows, then using that to scale the creation of high-quality training data for LLMs. Digital World Models help us push the frontier of ultra long horizon workflows, and unlock a new class of self-improving RL environments. This is our scalable approach to simulating all of the world’s intelligence. The round was also joined by Datadog, Inc., Samsung Ventures, Gokul Rajaram, Factorial Capital, and a large cohort of amazing AI leaders across Anthropic, OpenAI, Google DeepMind, NVIDIA, Recursive, and more.✨ It has been the ride of a lifetime. But we’re just getting started. The best is yet to come. "Do not go gentle into that good night, Rage, rage against the dying of the light" - Dylan Thomas (1954)

PatronusAI

95,619 views • 2 months ago