Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We’d like to know how far current frontier AIs generalize “out-of-distribution” (that is, how able they are to solve problems they haven’t seen before), to know how fast things will move. So we watched some AIs play Pokemon.

11,012 görüntüleme • 15 gün önce •via X (Twitter)

18 Yorum

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

(Almost all real work in the following is by @nc_znc)

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Humans often take 20 hours to complete Pokemon, so it’s a nice long-context task, & early last year it was out of distribution. The task got famous last February. LLMs weren’t very good, needed huge scaffolding, & still took 100s of hours to get anywhere

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

We test frontier models using one standardized scaffold, i.e. without the heavy signposting and borderline cheating of past harnesses.

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Results: obvs newer models do far better. Fable, Opus 5, Sol and Astra all solve the tricky Rock Tunnel puzzle quickly, every time, where GPT-5.5 and Opus 4.8 had usually failed. Astra is pushing up against optimal (maybe 3hrs human-equivalent play, vs a 1.75h world record)

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Clearly a big problem for naive evals is the vast amount of unbelievably granular (frame-by-frame) Pokemon training data on the public internet, and the high probability that Pokemon-specific RL environments are being used now.

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Newer models also show increased recall of important milestones in the game.

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Large spikes in performance could indicate the onset of Pokemon-specific post-training, and Astra is such an example.

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

We thus try to push the AIs out of distribution. For the first time, we test models on random layout variants (i.e. with the map scrambled) to test their generalization. Sol and Opus 5 collapse in performance on Rock Tunnel variants. Astra does not fall at all.

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Finally, we test the models on an obscure Pokemon-like game, the fan-made “Pokemon Brown”. Astra successfully completed the game in 10 thousand steps (very roughly 5 hours of human play), without displaying notable memorization of this new game.

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Overall we see evidence of memorization and shallow generalization in many models, but also reliability and tolerance to task variants (particularly Astra). We thus see evidence both for increasing “general intelligence” and for “whack-a-mole” task-specific post-training

gavin leech (Non-Reasoning) profil fotoğrafı
gavin leech (Non-Reasoning)15 gün önce

Lots more colour and fun experiments (newer Claudes dutifully fight all the battles, where all OpenAI models skip as many of them as possible) at the link

jsd profil fotoğrafı
jsd15 gün önce

This is great!

dylan profil fotoğrafı
dylan15 gün önce

Cool asf

☃️Darth thromBOOzyt📯 profil fotoğrafı
☃️Darth thromBOOzyt📯15 gün önce

We really should try Challenge Runs or ROM Hacks that require deep planning. There's basically no training data for this (unless they went very heavy on RL which seems ridiculous) - No Damage - No Heals - Pokemon Level down instead up - ... (Speedrunners offer many ideas here)

j d profil fotoğrafı
j d15 gün önce

there are many fan-made pokemon games and challenge runs, might have to start benchmarking nuzlocke stuff

zbrek profil fotoğrafı
zbrek14 gün önce

Make them play a modern roguelike like Caves of Qud or Cogmind. If they don't die within 5 minutes I'll concede AGI. If they beat a necro tower in Dwarf Fortress Adventure Mode we have ASI.

Michal Langmajer profil fotoğrafı
Michal Langmajer15 gün önce

if they already tested it on older versions of those models, doesn't it mean that it gonna be included in the training data for those new versions they now do much better? (= it's not anymore an unknown problem)

TheDigitalRetailHivemindNetwork💹🧲 profil fotoğrafı
TheDigitalRetailHivemindNetwork💹🧲15 gün önce

@scaling01 Like no one ever was, tum tum tudum

Benzer Videolar