Video wird geladen...
Video konnte nicht geladen werden
Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by Zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
500,588 Aufrufe • vor 3 Monaten •via X (Twitter)
64 Kommentare

To succeed at this game, agents must reason about shape rotation, concurrency, and optimizing against competing tradeoffs. To match the human world record on all puzzles would be an insane feat. Agents played the game entirely through a python REPL with hex coords, no vision.

Opus Magnum is an awesome test bed for model capabilities because each puzzle has an infinite solution space, with some solutions scoring better than others. We can judge quality by comparing against human scores. 1 is Human WR. Here's Fable 5 optimizing a puzzle over its run:

On puzzles that all 3 solved, Fable 5 (high) beat the next best model GPT-5.5 (xhigh) by reaching better solutions in fewer turns, and similarly outpaces Fable 5 (low). This chart shows that Fable 5 (high) on average reached 80% of the human best score on puzzles it solved.

No language model solved all 36 puzzles. Fable 5 and GPT-5.5 performed best, with GLM 5.2 as the best open weights model. No model beat a human world record, though a few matched or got close on the easier puzzles.

Full benchmark site here, where you can explore every model's solutions to different puzzles and read more about the methodology, agent harness, and specific model failure modes

Comparing GPT-5.5 and Claude Fable 5:

@zachtronics fable doesn't look too bad compared to humans in this clip

@zachtronics Fable got super close to the human world record on this puzzle and a number of other puzzles

@zachtronics smart fable

@zachtronics why is gpt-5.5s solution like that? surely that is not economical

@zachtronics it only had a few solutions like that, I only chose the Alchemical Jewel puzzle to show a contrast between what is a good solution and what is a bad solution but I think it would struggle with colliding atoms and make something super spread out to avoid that class of error

@zachtronics interesting, do you know how often super long tracks like that appear in human solutions?

@zachtronics very rarely and it's more like a reddit shitpost when they do it

@zachtronics I'll also say I'd be pretty surprised if opus magnum .solution files are in distribution

@zachtronics this game looks pretty difficult, i might not be agi

@zachtronics It's one of the hardest puzzle games I know of! Only ~10% of steam players complete all of the campaign puzzles, and of those, a much smaller percent approach world record scores.

@zachtronics look @PKUCXK, great environment for visual primitives

@zachtronics oh hell yea i remember when we talked about this SOOOO COOL IM GLAD YOU GOT FABLE BENCHMARKS

@zachtronics I know lol I was about to leave this benchmark thinking the claude models just sucked at it. Really glad I caught Fable before the api went down!

@zachtronics the design it made is so elegant wtf what did opus look like ??

@zachtronics and what happens if we prompt "make this design faster"

@zachtronics I forget which but there was one fable solution where it got either cycle optimal or close to cycle optimal but realized it could trade off 10 cycles for a better gold score

@zachtronics I love this benchmark!!! I’m slightly concerned that solutions will leak though.

@zachtronics We have access to an infinite amount of puzzles in the env, I just picked the campaign ones for relevance to human players and baseline optimal scores. At least currently, most frontier knowledge on the game is in the community discord or binary .solution files.

@a__tomala @zachtronics Their subreddit compiles all top solutions in 4 categories, it should all be in the training data Still interesting to see tho! Curious how they'd do on invented puzzles

solutions are online but whether they are in a format that is understandable to the models is another question. Any trial that found a way to source existing solutions from the internet I reran. on invented puzzles, they do decently, but the gap between them and people is often larger bc they're harder. But it's also harder to objectively evaluate how optimal their solutions are bc they have less human optimization effort behind them.

@zachtronics it's a beautiful and difficult game. i think we need more game benchmarks honestly

@zachtronics Its cool that human record solutions almost all center around using the 6-arm tool as the centerpiece of the solution to do as much as possible. The AIs didn't make that leap for some reason

@zachtronics Yeah some members of the community suggested including more explicit guidance on optimal strategies but I mostly wanted to see what they would reach for on their own and to let them figure out the game independently. There were a few solutions that used the 6-arm but very few

@zachtronics it's not clear to me what makes a solution better. Is it speed? Does using fewer resources matter?

@zachtronics Good question! They are graded on the sum score, which is cycles (speed) + gold (cost of each part used) + area (size). Human: 97 cycles, 160 gold, 41 area, sum 298 Fable: 157 cycles, 110 gold, 43 area, sum 310 So Fable got close, but human traded 50 gold for 60 cycles and won

@zachtronics ok, makes sense, thanks

@zachtronics Oh, I thought it was obvious that Opus 4.8 is a wordcel. He’s like a humanities professor who can also code. But humanities professor first.

@zachtronics Claude 4.8 Opus took a full ten minutes to solve this math question involving rotations. It was really hard for Claude, and the thought process showed Claude used simulations and probabilities. GPT 5.5 answered it very quickly. Yours is an interesting challenge for LLMs.

@zachtronics When do we get numbers for Le Chaton Fat?

@zachtronics finally a decent benchmark for AI models

@zachtronics Ingenius benchmark. Makes me want to go through my Steam library for games that it would be interesting to watch Claude Cowork play. However, I have done this already with Caves of Qud. Very fun to watch. Haiku was the best, because it didn’t wait forever bt moves.

@zachtronics What did the harness look like for Caves of Qud? computer use? Or did you mod it and use an mcp? How did you represent game state to the model? I’ve thought a bit about doing this with Caves of Qud, held off mostly bc navigation is particularly hard for models in other games

@zachtronics It was simply the computer use skill. A different game state representation here would definitely be nice, as a way to reduce input tokens and get faster behavior.

@zachtronics In the past few days, I've cloned and started messing with the Mindcraft library, so I've just learned game state representations alternative to CV exist.

@zachtronics I mention Mindcraft-Mineflayer because with that setup, the agent sees the world-state as an easily parseable JSON. Another useful pattern is A mod for accessing runtime memory of Dwarf Fortress. Idk if memory access tool exists for CoQ.

@zachtronics i'm a longtime fan of spacechem and infinifactory

@zachtronics hahaha what a great benchmark idea! love zachtronics

@zachtronics they automated my youtuber

@repligate @zachtronics I'm super curious about puzzle 14 that fable couldn't solve but many other models got 80+?

@repligate @zachtronics I ran each puzzle 3 trials per model, probably would have gotten it if I ran it more (low effort fable got it) but on some level this is all nondeterministic

@zachtronics i was about to say woa this looks like that alchemy game :D

@zachtronics Are they coming in to each puzzle fresh? What about having them record learnings and building "experience"? I'm guessing the stronger model will learn better techniques more quickly.

That would be an area for future research, and I did have an earlier version of the benchmark where it had 15 generated puzzles (not from the real game) and it did them all sequentially, one hour per puzzle, and wrote down memories that it could use in subsequent puzzles. Ultimately that made it 15 hours per model, whereas independent puzzles in parallel takes 1 hour total wall clock time. It’s a trade off I made for convenience. I could picture a world where it does say all chapter 1 puzzles in parallel, then writes memories/learning that could be useful for future puzzles, then does chapter 2, writes memories, and so on. Would still be 5 hours per model.

@zachtronics GPT 5.5:

@zachtronics gpt 5.5 said lemme just build a whole highway to solve this fable 5 said nah im built diff 💀

@zachtronics Love the game and I love this benchmark.

@zachtronics This actually seems like a good benchmark

@zachtronics Cool benchmark! Fable's lead is quite impressive and a sign of real progress. I wonder if this kind of game could be used to evaluate multi--agent orchestrations.

@zachtronics Love this Rob! Thanks for slipping it in =) I might have to play

@zachtronics Interesting

@zachtronics Yooo Great job man ! I was curious about teaching Hermes play Opus Magnum. 💜

@zachtronics now AI agents doing factorio

@zachtronics Zachtronics games are brilliant. What a good idea for an ai bench.

@zachtronics Opus Magnum isn't really pass/fail though, half the game is optimizing cost vs cycles vs area. did you score solution efficiency or just whether they solved it? that's where I'd expect the models to actually separate

@zachtronics Yup! If you check out the thread / benchmark site you'll see that we graded on the sum score (cycles + gold + area) and normalized it with human optimal scores as 1

@liminal_bardo @zachtronics OH MAN this is such a great benchmark. We should absolutely be using the ZachBench to evaluate reasoning.

@zachtronics Looks interesting,Where can I play the game

@zachtronics Steam!
