Loading video...
Video Failed to Load
Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by Zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
498,486 views • 2 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
