Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by Zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.

500,588 görüntüleme • 3 ay önce •via X (Twitter)

64 Yorum

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

To succeed at this game, agents must reason about shape rotation, concurrency, and optimizing against competing tradeoffs. To match the human world record on all puzzles would be an insane feat. Agents played the game entirely through a python REPL with hex coords, no vision.

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

Opus Magnum is an awesome test bed for model capabilities because each puzzle has an infinite solution space, with some solutions scoring better than others. We can judge quality by comparing against human scores. 1 is Human WR. Here's Fable 5 optimizing a puzzle over its run:

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

On puzzles that all 3 solved, Fable 5 (high) beat the next best model GPT-5.5 (xhigh) by reaching better solutions in fewer turns, and similarly outpaces Fable 5 (low). This chart shows that Fable 5 (high) on average reached 80% of the human best score on puzzles it solved.

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

No language model solved all 36 puzzles. Fable 5 and GPT-5.5 performed best, with GLM 5.2 as the best open weights model. No model beat a human world record, though a few matched or got close on the easier puzzles.

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

Full benchmark site here, where you can explore every model's solutions to different puzzles and read more about the methodology, agent harness, and specific model failure modes

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

Comparing GPT-5.5 and Claude Fable 5:

Lisan al Gaib profil fotoğrafı
Lisan al Gaib3 ay önce

@zachtronics fable doesn't look too bad compared to humans in this clip

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics Fable got super close to the human world record on this puzzle and a number of other puzzles

Lisan al Gaib profil fotoğrafı
Lisan al Gaib3 ay önce

@zachtronics smart fable

j⧉nus profil fotoğrafı
j⧉nus3 ay önce

@zachtronics why is gpt-5.5s solution like that? surely that is not economical

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics it only had a few solutions like that, I only chose the Alchemical Jewel puzzle to show a contrast between what is a good solution and what is a bad solution but I think it would struggle with colliding atoms and make something super spread out to avoid that class of error

j⧉nus profil fotoğrafı
j⧉nus3 ay önce

@zachtronics interesting, do you know how often super long tracks like that appear in human solutions?

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics very rarely and it's more like a reddit shitpost when they do it

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics I'll also say I'd be pretty surprised if opus magnum .solution files are in distribution

execute profil fotoğrafı
execute3 ay önce

@zachtronics this game looks pretty difficult, i might not be agi

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics It's one of the hardest puzzle games I know of! Only ~10% of steam players complete all of the campaign puzzles, and of those, a much smaller percent approach world record scores.

Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) profil fotoğrafı
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)3 ay önce

@zachtronics look @PKUCXK, great environment for visual primitives

BOOTOSHI 👑 profil fotoğrafı
BOOTOSHI 👑3 ay önce

@zachtronics oh hell yea i remember when we talked about this SOOOO COOL IM GLAD YOU GOT FABLE BENCHMARKS

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics I know lol I was about to leave this benchmark thinking the claude models just sucked at it. Really glad I caught Fable before the api went down!

BOOTOSHI 👑 profil fotoğrafı
BOOTOSHI 👑3 ay önce

@zachtronics the design it made is so elegant wtf what did opus look like ??

BOOTOSHI 👑 profil fotoğrafı
BOOTOSHI 👑3 ay önce

@zachtronics and what happens if we prompt "make this design faster"

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics I forget which but there was one fable solution where it got either cycle optimal or close to cycle optimal but realized it could trade off 10 cycles for a better gold score

Alex Tomala profil fotoğrafı
Alex Tomala3 ay önce

@zachtronics I love this benchmark!!! I’m slightly concerned that solutions will leak though.

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics We have access to an infinite amount of puzzles in the env, I just picked the campaign ones for relevance to human players and baseline optimal scores. At least currently, most frontier knowledge on the game is in the community discord or binary .solution files.

Nit Ko profil fotoğrafı
Nit Ko3 ay önce

@a__tomala @zachtronics Their subreddit compiles all top solutions in 4 categories, it should all be in the training data Still interesting to see tho! Curious how they'd do on invented puzzles

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

solutions are online but whether they are in a format that is understandable to the models is another question. Any trial that found a way to source existing solutions from the internet I reran. on invented puzzles, they do decently, but the gap between them and people is often larger bc they're harder. But it's also harder to objectively evaluate how optimal their solutions are bc they have less human optimization effort behind them.

kethic profil fotoğrafı
kethic3 ay önce

@zachtronics it's a beautiful and difficult game. i think we need more game benchmarks honestly

champagne papineau profil fotoğrafı
champagne papineau3 ay önce

@zachtronics Its cool that human record solutions almost all center around using the 6-arm tool as the centerpiece of the solution to do as much as possible. The AIs didn't make that leap for some reason

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics Yeah some members of the community suggested including more explicit guidance on optimal strategies but I mostly wanted to see what they would reach for on their own and to let them figure out the game independently. There were a few solutions that used the 6-arm but very few

Breno Brito profil fotoğrafı
Breno Brito3 ay önce

@zachtronics it's not clear to me what makes a solution better. Is it speed? Does using fewer resources matter?

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics Good question! They are graded on the sum score, which is cycles (speed) + gold (cost of each part used) + area (size). Human: 97 cycles, 160 gold, 41 area, sum 298 Fable: 157 cycles, 110 gold, 43 area, sum 310 So Fable got close, but human traded 50 gold for 60 cycles and won

Breno Brito profil fotoğrafı
Breno Brito3 ay önce

@zachtronics ok, makes sense, thanks

Kyle R. McNease, Defective Altruist profil fotoğrafı
Kyle R. McNease, Defective Altruist3 ay önce

@zachtronics Oh, I thought it was obvious that Opus 4.8 is a wordcel. He’s like a humanities professor who can also code. But humanities professor first.

Fedesco profil fotoğrafı
Fedesco3 ay önce

@zachtronics Claude 4.8 Opus took a full ten minutes to solve this math question involving rotations. It was really hard for Claude, and the thought process showed Claude used simulations and probabilities. GPT 5.5 answered it very quickly. Yours is an interesting challenge for LLMs.

Adam Holter profil fotoğrafı
Adam Holter3 ay önce

@zachtronics When do we get numbers for Le Chaton Fat?

Andrew profil fotoğrafı
Andrew3 ay önce

@zachtronics finally a decent benchmark for AI models

Daniel Lisle profil fotoğrafı
Daniel Lisle3 ay önce

@zachtronics Ingenius benchmark. Makes me want to go through my Steam library for games that it would be interesting to watch Claude Cowork play. However, I have done this already with Caves of Qud. Very fun to watch. Haiku was the best, because it didn’t wait forever bt moves.

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics What did the harness look like for Caves of Qud? computer use? Or did you mod it and use an mcp? How did you represent game state to the model? I’ve thought a bit about doing this with Caves of Qud, held off mostly bc navigation is particularly hard for models in other games

Daniel Lisle profil fotoğrafı
Daniel Lisle3 ay önce

@zachtronics It was simply the computer use skill. A different game state representation here would definitely be nice, as a way to reduce input tokens and get faster behavior.

Daniel Lisle profil fotoğrafı
Daniel Lisle3 ay önce

@zachtronics In the past few days, I've cloned and started messing with the Mindcraft library, so I've just learned game state representations alternative to CV exist.

Daniel Lisle profil fotoğrafı
Daniel Lisle3 ay önce

@zachtronics I mention Mindcraft-Mineflayer because with that setup, the agent sees the world-state as an easily parseable JSON. Another useful pattern is A mod for accessing runtime memory of Dwarf Fortress. Idk if memory access tool exists for CoQ.

Byron - profil fotoğrafı
Byron -3 ay önce

@zachtronics i'm a longtime fan of spacechem and infinifactory

Gergő Móricz profil fotoğrafı
Gergő Móricz3 ay önce

@zachtronics hahaha what a great benchmark idea! love zachtronics

adas🧦🌹 profil fotoğrafı
adas🧦🌹3 ay önce

@zachtronics they automated my youtuber

Boyd Kane (quantized) profil fotoğrafı
Boyd Kane (quantized)3 ay önce

@repligate @zachtronics I'm super curious about puzzle 14 that fable couldn't solve but many other models got 80+?

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@repligate @zachtronics I ran each puzzle 3 trials per model, probably would have gotten it if I ran it more (low effort fable got it) but on some level this is all nondeterministic

Solgato profil fotoğrafı
Solgato3 ay önce

@zachtronics i was about to say woa this looks like that alchemy game :D

Young Tai Ahn profil fotoğrafı
Young Tai Ahn3 ay önce

@zachtronics Are they coming in to each puzzle fresh? What about having them record learnings and building "experience"? I'm guessing the stronger model will learn better techniques more quickly.

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

That would be an area for future research, and I did have an earlier version of the benchmark where it had 15 generated puzzles (not from the real game) and it did them all sequentially, one hour per puzzle, and wrote down memories that it could use in subsequent puzzles. Ultimately that made it 15 hours per model, whereas independent puzzles in parallel takes 1 hour total wall clock time. It’s a trade off I made for convenience. I could picture a world where it does say all chapter 1 puzzles in parallel, then writes memories/learning that could be useful for future puzzles, then does chapter 2, writes memories, and so on. Would still be 5 hours per model.

Brother Freeman profil fotoğrafı
Brother Freeman3 ay önce

@zachtronics GPT 5.5:

Brosko profil fotoğrafı
Brosko3 ay önce

@zachtronics gpt 5.5 said lemme just build a whole highway to solve this fable 5 said nah im built diff 💀

beholdling profil fotoğrafı
beholdling3 ay önce

@zachtronics Love the game and I love this benchmark.

SolarFable profil fotoğrafı
SolarFable3 ay önce

@zachtronics This actually seems like a good benchmark

The Guy With A Hat profil fotoğrafı
The Guy With A Hat3 ay önce

@zachtronics Cool benchmark! Fable's lead is quite impressive and a sign of real progress. I wonder if this kind of game could be used to evaluate multi--agent orchestrations.

Jared Smith profil fotoğrafı
Jared Smith3 ay önce

@zachtronics Love this Rob! Thanks for slipping it in =) I might have to play

𝕹յ𝖌𝖍𝖙𝕱𝖆𝖇𝖑𝖊𝖃 profil fotoğrafı
𝕹յ𝖌𝖍𝖙𝕱𝖆𝖇𝖑𝖊𝖃3 ay önce

@zachtronics Interesting

Foola Froos🌞 profil fotoğrafı
Foola Froos🌞3 ay önce

@zachtronics Yooo Great job man ! I was curious about teaching Hermes play Opus Magnum. 💜

farismaulana profil fotoğrafı
farismaulana3 ay önce

@zachtronics now AI agents doing factorio

fuzzyfacts profil fotoğrafı
fuzzyfacts3 ay önce

@zachtronics Zachtronics games are brilliant. What a good idea for an ai bench.

Somi profil fotoğrafı
Somi3 ay önce

@zachtronics Opus Magnum isn't really pass/fail though, half the game is optimizing cost vs cycles vs area. did you score solution efficiency or just whether they solved it? that's where I'd expect the models to actually separate

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics Yup! If you check out the thread / benchmark site you'll see that we graded on the sum score (cycles + gold + area) and normalized it with human optimal scores as 1

AutoGolpe 127b Flash Q2.5 Abliterated profil fotoğrafı
AutoGolpe 127b Flash Q2.5 Abliterated3 ay önce

@liminal_bardo @zachtronics OH MAN this is such a great benchmark. We should absolutely be using the ZachBench to evaluate reasoning.

tower profil fotoğrafı
tower3 ay önce

@zachtronics Looks interesting,Where can I play the game

Rob Haisfield profil fotoğrafı
Rob Haisfield3 ay önce

@zachtronics Steam!

Benzer Videolar