Загрузка видео...
Не удалось загрузить видео
Game development remains one of the most-requested, and most-challenging, categories on Arena. Wayne Chi, PhD candidate at Carnegie Mellon University and research intern at Arena, just walked us through GameDevBench: a benchmark built from real tutorials that turns game development into verifiable, deterministic tasks. How do top frontier models... show more
17,370 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 10

@iamwaynechi Wow, so frontier Agentss are still bad at going but they can solve millenium maths questions Seems the world ain't ending soon

@iamwaynechi Deterministic tasks a beginner can finish in an hour. That’s a real bench.

@iamwaynechi did the misses look like bad code or “compiled fine, scene still wrong”?

@iamwaynechi nice

@iamwaynechi 游戏开发在 Arena 一直难拿分。用真实教程拆成可验证任务之后,卡点到底是写代码,还是多模态理解,这俩得分开看,别混成一句「模型不行」。

Building it from tutorials is the clever part — a tutorial already encodes the intended end state, so you get ground truth for free on a task that's otherwise pure taste. My bet is the bottleneck isn't code or vision but the loop between them: knowing the sprite is in the wrong place is easy, knowing which of forty lines put it there is the hard one.

@iamwaynechi game dev benchmarks are brutal because "compiles" and "is actually fun" are different bars. we see that gap daily building GameGen — prompt to playable game in minutes.

@iamwaynechi Multimodal context usually breaks way before the code does

@iamwaynechi game dev is one of the hardest for agents cause every engine is a different planet. curious what the answer ends up being

@iamwaynechi GameDevBench from real tutorials — agents that survive PowerPoint demos just met a category that doesn’t. Multimodal coding heat with a curriculum attached.
