Загрузка видео...

Не удалось загрузить видео

На главную

Game development remains one of the most-requested, and most-challenging, categories on Arena. Wayne Chi, PhD candidate at Carnegie Mellon University and research intern at Arena, just walked us through GameDevBench: a benchmark built from real tutorials that turns game development into verifiable, deterministic tasks. How do top frontier models...

17,370 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 10

Фото профиля Drex
Drex1 месяц назад

@iamwaynechi Wow, so frontier Agentss are still bad at going but they can solve millenium maths questions Seems the world ain't ending soon

Фото профиля Ava Nakamura
Ava Nakamura1 месяц назад

@iamwaynechi Deterministic tasks a beginner can finish in an hour. That’s a real bench.

Фото профиля okimraise.eth
okimraise.eth1 месяц назад

@iamwaynechi did the misses look like bad code or “compiled fine, scene still wrong”?

Фото профиля Harkirat Behl
Harkirat Behl1 месяц назад

@iamwaynechi nice

Фото профиля VastPlan
VastPlan1 месяц назад

@iamwaynechi 游戏开发在 Arena 一直难拿分。用真实教程拆成可验证任务之后,卡点到底是写代码,还是多模态理解,这俩得分开看,别混成一句「模型不行」。

Фото профиля Ricci Research
Ricci Research1 месяц назад

Building it from tutorials is the clever part — a tutorial already encodes the intended end state, so you get ground truth for free on a task that's otherwise pure taste. My bet is the bottleneck isn't code or vision but the loop between them: knowing the sprite is in the wrong place is easy, knowing which of forty lines put it there is the hard one.

Фото профиля GameGen
GameGen1 месяц назад

@iamwaynechi game dev benchmarks are brutal because "compiles" and "is actually fun" are different bars. we see that gap daily building GameGen — prompt to playable game in minutes.

Фото профиля MASA
MASA1 месяц назад

@iamwaynechi Multimodal context usually breaks way before the code does

Фото профиля Uno Alpha
Uno Alpha1 месяц назад

@iamwaynechi game dev is one of the hardest for agents cause every engine is a different planet. curious what the answer ends up being

Фото профиля Artiment Index
Artiment Index1 месяц назад

@iamwaynechi GameDevBench from real tutorials — agents that survive PowerPoint demos just met a category that doesn’t. Multimodal coding heat with a curriculum attached.

Похожие видео

Chamath Palihapitiya believes AGI may already exist inside leading AI labs and the bigger story is that advanced intelligence is becoming cheaper and more widely available (Save this). Chamath Palihapitiya argues that the public may be focused too much on benchmark rankings, while frontier labs are already developing models capable of complex reasoning, coding, research, and tool use. The main question is how quickly companies will release these systems and how much access they will provide. AGI has not been officially confirmed and strong benchmark results do not necessarily prove that a model can perform every intellectual task like a human. However, AI capabilities are improving quickly, while the cost of running advanced models continues to fall. That combination is important because cheaper AI can be used by more businesses for customer service, software development, research, marketing, financial analysis, and automation. Competition is also accelerating among OpenAI, Anthropic, Google, xAI, Meta, and open source developers because as more companies release capable models, users gain more choices and prices continue to decline. This creates a powerful cycle in which better models attract more users, more usage generates more revenue and data, and lower prices encourage companies to apply AI to additional tasks. The biggest challenge is moving from impressive demonstrations to measurable business results. Companies still need to redesign workflows, train employees, protect sensitive information, and prove that AI spending is producing a real return on investment. AI agents could create the next major increase in demand because they can plan tasks, use tools, check their work, retry failed actions and operate for long periods without constant human supervision. Even if each AI task becomes cheaper, total usage could grow much faster as businesses use models across more departments and this could increase demand for GPUs, high bandwidth memory, networking equipment, electricity, cooling systems, and data centers.

Milk Road AI

13,501 просмотров • 1 месяц назад

Before software engineers even begin writing code, they have to set the stage of the entire development process. This process requires engineers to make complex tradeoffs between requirements, system design, and implementations details. Current IDEs that rely on AI features, like chat and inline coding, can help engineers get the job done quickly on small development tasks. Still, engineers spend much more time on larger projects—even after the initial code is generated—by conducting rigorous testing and creating documentation. This is where today’s AI IDEs can do more to accelerate the development lifecycle—and this is why we built Kiro. Kiro is an AI IDE that helps you go from prototype to production with spec-driven development and agent hooks. From simple to complex tasks, Kiro works alongside you to turn prompts into detailed specs, then into working code, docs, and test so what you build is exactly what you want and ready to share with your team. After a developer builds the code with Kiro, Kiro’s agent hooks help engineers solve challenging problems and automate tasks like generating documentation and unit tests. Kiro brings structure and mature engineering practices to AI coding, so you can go from concept to application while being in the driver’s seat every step of the way. Kiro is free during preview, and supports Mac, Windows, and Linux, and most popular programming languages. We're excited for you to try it out and let us know what you think ➡️

Swami Sivasubramanian

154,343 просмотров • 1 год назад