正在加载视频...
视频加载失败
Agents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!
56,351 次观看 • 8 天前 •via X (Twitter)
51 条评论

very interesting, would love to run Gemini 4 Argon on this!

🙌🙌 yes we’d love to!

We wanted to test agents’ ability to conduct multi-hop reasoning across many services (up to 40 services in 1 environment!). We simulate prod as full Kubernetes stacks: 70 services across three full environments, including databases, queues and workers, all communicating with realistic traffic. Then we inject faults as configs, runtime errors and container faults for agents to investigate and repair. Below is a look at a sample environment!

We build a new type of functional verifiers, measuring (1) the outcome and the (2) safety of the repair In one run, GPT-6-Astra traced the problem to a slow Postgres query, but left the existing connection settings intact. The diagnosis is correct, but repair is incomplete.

Reward hacking behaviour: In an early incident-repair task, Claude Opus 5 was supposed to restore a broken service. It instead forged a cookie and sent a malicious Python pickle payload to the Kubernetes pod running the grader. When that pod loaded the payload, it executed the agent’s code, getting RCE

Turning up reasoning effort didn’t consistently fix more incidents. Eight of ten models scored best below their highest reasoning setting. e.g Claude Opus 5.5 peaked at medium effort: 58.3% success, versus 41.7% at max. The max setting cost roughly 4.7× as much per run.

Incident Arena is integrated with Harbor. Explore the paper, tasks, leaderboard and agent runs:

this gonna help agents stop PagerDuty from taking innocent engineer’s sleep, so goated

the US GDP is blocked by engineers not getting enough sleep

"We simulate prod as full Kubernetes stacks: 70 services across three..." this is really cool

yessir, whole companies rolled into a harbor environment

congrats on the launch!

thank you sir!

Very cool!

🙏 thank you jackson!! let's catch up soon

🔥🔥

openlens sre-agent when?

Lowkey the leaderboard fits my preference of what model to use in general as well

indeed i use 6-astra-xhigh to center my divs

The task realism we so desperately need. Congrats @andrezfu !!

thank you rishi 🙏🏽🙏🏽

paged at 2am is the loud version. mine: dns ENOTFOUND took down the orchestrator, and the fallback compiler needed the same api, so both failed together. no page

bc we have k8s and advanced networking we can even simulate dns networking faults !

that's the right test. mine failed for real with no simulation. a dual path that shares one dependency isn't a fallback

Congratulations Andre. This is going to be incredible.

thank you shipra! appreciate your help on this work!

Congrats, Andre! Very interesting

thank you jesse!!

congrats!!

thank you sir

omg this is awesome

thanks man! 🙏

im still waiting for funnybench

JokeBench is just the agent looking at my handwritten code

Cool stuff!

thank you man! let's catch up soon!

Amazing stuff

thanks man

very very cool!

thanks man!

Very cool bench!

thank you max! you guys helped catch a ton of bugs, appreciate your support man

this is awesome andre!!

thank you sir!

the reward-hack finding generalizes: a grader inside the environment the agent controls is part of the attack surface. once the cheapest route to green runs through the checker, that's the route that gets found. score repairs from outside the blast radius.

the sev0 savior we needddd - nice work andre!

no more pagerduty 🤞 ??

Very cool

thanks big dawg

on-call agents need a kill switch the human still trusts at 2am. who holds it?

Great stuff Andre! Much needed move beyond single container environments to multi-service sims



