正在加载视频...
视频加载失败
We’d like to know how far current frontier AIs generalize “out-of-distribution” (that is, how able they are to solve problems they haven’t seen before), to know how fast things will move. So we watched some AIs play Pokemon.
11,012 次观看 • 15 天前 •via X (Twitter)
18 条评论

(Almost all real work in the following is by @nc_znc)

Humans often take 20 hours to complete Pokemon, so it’s a nice long-context task, & early last year it was out of distribution. The task got famous last February. LLMs weren’t very good, needed huge scaffolding, & still took 100s of hours to get anywhere

We test frontier models using one standardized scaffold, i.e. without the heavy signposting and borderline cheating of past harnesses.

Results: obvs newer models do far better. Fable, Opus 5, Sol and Astra all solve the tricky Rock Tunnel puzzle quickly, every time, where GPT-5.5 and Opus 4.8 had usually failed. Astra is pushing up against optimal (maybe 3hrs human-equivalent play, vs a 1.75h world record)

Clearly a big problem for naive evals is the vast amount of unbelievably granular (frame-by-frame) Pokemon training data on the public internet, and the high probability that Pokemon-specific RL environments are being used now.

Newer models also show increased recall of important milestones in the game.

Large spikes in performance could indicate the onset of Pokemon-specific post-training, and Astra is such an example.

We thus try to push the AIs out of distribution. For the first time, we test models on random layout variants (i.e. with the map scrambled) to test their generalization. Sol and Opus 5 collapse in performance on Rock Tunnel variants. Astra does not fall at all.

Finally, we test the models on an obscure Pokemon-like game, the fan-made “Pokemon Brown”. Astra successfully completed the game in 10 thousand steps (very roughly 5 hours of human play), without displaying notable memorization of this new game.

Overall we see evidence of memorization and shallow generalization in many models, but also reliability and tolerance to task variants (particularly Astra). We thus see evidence both for increasing “general intelligence” and for “whack-a-mole” task-specific post-training

Lots more colour and fun experiments (newer Claudes dutifully fight all the battles, where all OpenAI models skip as many of them as possible) at the link

This is great!

Cool asf

We really should try Challenge Runs or ROM Hacks that require deep planning. There's basically no training data for this (unless they went very heavy on RL which seems ridiculous) - No Damage - No Heals - Pokemon Level down instead up - ... (Speedrunners offer many ideas here)

there are many fan-made pokemon games and challenge runs, might have to start benchmarking nuzlocke stuff

Make them play a modern roguelike like Caves of Qud or Cogmind. If they don't die within 5 minutes I'll concede AGI. If they beat a necro tower in Dwarf Fortress Adventure Mode we have ASI.

if they already tested it on older versions of those models, doesn't it mean that it gonna be included in the training data for those new versions they now do much better? (= it's not anymore an unknown problem)

@scaling01 Like no one ever was, tum tum tudum
相关视频
We know how they are conditioning people to believe as they do.
MAGA Cult Slayer🦅🇺🇸
22,741 次观看 • 5 个月前
