
Rob Haisfield
@RobertHaisfield • 9,730 subscribers
cofounder of @websim_ai, imagining new internets with our users. GenAI, TfT, BeSci, HCI, UX. Ex-Tana, Edge & Node, Spark Wave
Shorts
Videos

Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by Zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
Rob Haisfield500,588 次观看 • 3 个月前

Fable with low reasoning is roughly comparable to GPT-5.5 medium on the 20 puzzles they both solved. Fable scored higher in 14 of those, and I think Warming Tonic illustrates how different quality alchemical machines look. Solve rate isn't everything, quality is important!
Rob Haisfield41,084 次观看 • 3 个月前
没有更多内容可加载