
Rob Haisfield
@RobertHaisfield • 9,730 subscribers
cofounder of @websim_ai, imagining new internets with our users. GenAI, TfT, BeSci, HCI, UX. Ex-Tana, Edge & Node, Spark Wave
Shorts
Videos

Are AI agents shape rotators? In this new benchmark, we let the models play campaign puzzles in Opus Magnum, a puzzle game by Zachtronics. Ironically, Claude Opus 4.8 performed poorly, being beaten by GPT-5.5, Gemini 3.5 Flash, and GLM 5.2. Claude Fable 5 crushed them all.
Rob Haisfield500,588 görüntüleme • 3 ay önce

I optimized my Rune Mysteries Quest script 75% to 1530 ticks (15 min normal game time, from 58 min). Loop: write a script checkpoint, run it, note learnings and ideas to a shared log file, repeat. Each loop took five min, I let it rip overnight with a claude team.
Rob Haisfield85,854 görüntüleme • 7 ay önce

Fable with low reasoning is roughly comparable to GPT-5.5 medium on the 20 puzzles they both solved. Fable scored higher in 14 of those, and I think Warming Tonic illustrates how different quality alchemical machines look. Solve rate isn't everything, quality is important!
Rob Haisfield41,084 görüntüleme • 3 ay önce
Daha fazla içerik yok.