Loading video...

Video Failed to Load

Go Home

We’d like to know how far current frontier AIs generalize “out-of-distribution” (that is, how able they are to solve problems they haven’t seen before), to know how fast things will move. So we watched some AIs play Pokemon.

11,012 views • 15 days ago •via X (Twitter)

18 Comments

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

(Almost all real work in the following is by @nc_znc)

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Humans often take 20 hours to complete Pokemon, so it’s a nice long-context task, & early last year it was out of distribution. The task got famous last February. LLMs weren’t very good, needed huge scaffolding, & still took 100s of hours to get anywhere

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

We test frontier models using one standardized scaffold, i.e. without the heavy signposting and borderline cheating of past harnesses.

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Results: obvs newer models do far better. Fable, Opus 5, Sol and Astra all solve the tricky Rock Tunnel puzzle quickly, every time, where GPT-5.5 and Opus 4.8 had usually failed. Astra is pushing up against optimal (maybe 3hrs human-equivalent play, vs a 1.75h world record)

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Clearly a big problem for naive evals is the vast amount of unbelievably granular (frame-by-frame) Pokemon training data on the public internet, and the high probability that Pokemon-specific RL environments are being used now.

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Newer models also show increased recall of important milestones in the game.

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Large spikes in performance could indicate the onset of Pokemon-specific post-training, and Astra is such an example.

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

We thus try to push the AIs out of distribution. For the first time, we test models on random layout variants (i.e. with the map scrambled) to test their generalization. Sol and Opus 5 collapse in performance on Rock Tunnel variants. Astra does not fall at all.

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Finally, we test the models on an obscure Pokemon-like game, the fan-made “Pokemon Brown”. Astra successfully completed the game in 10 thousand steps (very roughly 5 hours of human play), without displaying notable memorization of this new game.

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Overall we see evidence of memorization and shallow generalization in many models, but also reliability and tolerance to task variants (particularly Astra). We thus see evidence both for increasing “general intelligence” and for “whack-a-mole” task-specific post-training

gavin leech (Non-Reasoning)'s profile picture
gavin leech (Non-Reasoning)15 days ago

Lots more colour and fun experiments (newer Claudes dutifully fight all the battles, where all OpenAI models skip as many of them as possible) at the link

jsd's profile picture
jsd15 days ago

This is great!

dylan's profile picture
dylan15 days ago

Cool asf

☃️Darth thromBOOzyt📯's profile picture
☃️Darth thromBOOzyt📯15 days ago

We really should try Challenge Runs or ROM Hacks that require deep planning. There's basically no training data for this (unless they went very heavy on RL which seems ridiculous) - No Damage - No Heals - Pokemon Level down instead up - ... (Speedrunners offer many ideas here)

j d's profile picture
j d15 days ago

there are many fan-made pokemon games and challenge runs, might have to start benchmarking nuzlocke stuff

zbrek's profile picture
zbrek14 days ago

Make them play a modern roguelike like Caves of Qud or Cogmind. If they don't die within 5 minutes I'll concede AGI. If they beat a necro tower in Dwarf Fortress Adventure Mode we have ASI.

Michal Langmajer's profile picture
Michal Langmajer15 days ago

if they already tested it on older versions of those models, doesn't it mean that it gonna be included in the training data for those new versions they now do much better? (= it's not anymore an unknown problem)

TheDigitalRetailHivemindNetwork💹🧲's profile picture
TheDigitalRetailHivemindNetwork💹🧲15 days ago

@scaling01 Like no one ever was, tum tum tudum

Related Videos