Loading video...

Video Failed to Load

Go Home

Agents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!

56,351 views • 8 days ago •via X (Twitter)

51 Comments

Logan Kilpatrick's profile picture
Logan Kilpatrick8 days ago

very interesting, would love to run Gemini 4 Argon on this!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

🙌🙌 yes we’d love to!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

We wanted to test agents’ ability to conduct multi-hop reasoning across many services (up to 40 services in 1 environment!). We simulate prod as full Kubernetes stacks: 70 services across three full environments, including databases, queues and workers, all communicating with realistic traffic. Then we inject faults as configs, runtime errors and container faults for agents to investigate and repair. Below is a look at a sample environment!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

We build a new type of functional verifiers, measuring (1) the outcome and the (2) safety of the repair In one run, GPT-6-Astra traced the problem to a slow Postgres query, but left the existing connection settings intact. The diagnosis is correct, but repair is incomplete.

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

Reward hacking behaviour: In an early incident-repair task, Claude Opus 5 was supposed to restore a broken service. It instead forged a cookie and sent a malicious Python pickle payload to the Kubernetes pod running the grader. When that pod loaded the payload, it executed the agent’s code, getting RCE

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

Turning up reasoning effort didn’t consistently fix more incidents. Eight of ten models scored best below their highest reasoning setting. e.g Claude Opus 5.5 peaked at medium effort: 58.3% success, versus 41.7% at max. The max setting cost roughly 4.7× as much per run.

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

Incident Arena is integrated with Harbor. Explore the paper, tasks, leaderboard and agent runs:

SpaceSquid's profile picture
SpaceSquid8 days ago

this gonna help agents stop PagerDuty from taking innocent engineer’s sleep, so goated

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

the US GDP is blocked by engineers not getting enough sleep

sensho's profile picture
sensho8 days ago

"We simulate prod as full Kubernetes stacks: 70 services across three..." this is really cool

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

yessir, whole companies rolled into a harbor environment

Daniel Wang's profile picture
Daniel Wang8 days ago

congrats on the launch!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you sir!

Jackson Clark's profile picture
Jackson Clark8 days ago

Very cool!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

🙏 thank you jackson!! let's catch up soon

Cameron Witkowski's profile picture
Cameron Witkowski8 days ago

🔥🔥

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

openlens sre-agent when?

kyle's profile picture
kyle8 days ago

Lowkey the leaderboard fits my preference of what model to use in general as well

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

indeed i use 6-astra-xhigh to center my divs

Rishi Desai's profile picture
Rishi Desai8 days ago

The task realism we so desperately need. Congrats @andrezfu !!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you rishi 🙏🏽🙏🏽

Shivam Patel's profile picture
Shivam Patel8 days ago

paged at 2am is the loud version. mine: dns ENOTFOUND took down the orchestrator, and the fallback compiler needed the same api, so both failed together. no page

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

bc we have k8s and advanced networking we can even simulate dns networking faults !

Shivam Patel's profile picture
Shivam Patel8 days ago

that's the right test. mine failed for real with no simulation. a dual path that shares one dependency isn't a fallback

Shipra Jha's profile picture
Shipra Jha8 days ago

Congratulations Andre. This is going to be incredible.

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you shipra! appreciate your help on this work!

Jesse Shulman's profile picture
Jesse Shulman8 days ago

Congrats, Andre! Very interesting

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you jesse!!

Vishal Jain's profile picture
Vishal Jain8 days ago

congrats!!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you sir

Gergő Móricz's profile picture
Gergő Móricz8 days ago

omg this is awesome

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thanks man! 🙏

Logan's profile picture
Logan8 days ago

im still waiting for funnybench

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

JokeBench is just the agent looking at my handwritten code

Adit Agarwal's profile picture
Adit Agarwal8 days ago

Cool stuff!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you man! let's catch up soon!

hrishikesh kamath's profile picture
hrishikesh kamath8 days ago

Amazing stuff

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thanks man

Utpal Nadiger's profile picture
Utpal Nadiger8 days ago

very very cool!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thanks man!

Max Kan's profile picture
Max Kan8 days ago

Very cool bench!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you max! you guys helped catch a ton of bugs, appreciate your support man

Shiv Kampani's profile picture
Shiv Kampani8 days ago

this is awesome andre!!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thank you sir!

John Rood's profile picture
John Rood8 days ago

the reward-hack finding generalizes: a grader inside the environment the agent controls is part of the attack surface. once the cheapest route to green runs through the checker, that's the route that gets found. score repairs from outside the blast radius.

Adi Sharma's profile picture
Adi Sharma8 days ago

the sev0 savior we needddd - nice work andre!

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

no more pagerduty 🤞 ??

Daivik Goel's profile picture
Daivik Goel8 days ago

Very cool

andre --dangerously-skip-permissions's profile picture
andre --dangerously-skip-permissions8 days ago

thanks big dawg

Dominik's profile picture
Dominik8 days ago

on-call agents need a kill switch the human still trusts at 2am. who holds it?

Rohan Gupta's profile picture
Rohan Gupta8 days ago

Great stuff Andre! Much needed move beyond single container environments to multi-service sims

Related Videos