Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Agents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!

56,351 Aufrufe • vor 8 Tagen •via X (Twitter)

51 Kommentare

Profilbild von Logan Kilpatrick
Logan Kilpatrickvor 8 Tagen

very interesting, would love to run Gemini 4 Argon on this!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

🙌🙌 yes we’d love to!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

We wanted to test agents’ ability to conduct multi-hop reasoning across many services (up to 40 services in 1 environment!). We simulate prod as full Kubernetes stacks: 70 services across three full environments, including databases, queues and workers, all communicating with realistic traffic. Then we inject faults as configs, runtime errors and container faults for agents to investigate and repair. Below is a look at a sample environment!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

We build a new type of functional verifiers, measuring (1) the outcome and the (2) safety of the repair In one run, GPT-6-Astra traced the problem to a slow Postgres query, but left the existing connection settings intact. The diagnosis is correct, but repair is incomplete.

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

Reward hacking behaviour: In an early incident-repair task, Claude Opus 5 was supposed to restore a broken service. It instead forged a cookie and sent a malicious Python pickle payload to the Kubernetes pod running the grader. When that pod loaded the payload, it executed the agent’s code, getting RCE

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

Turning up reasoning effort didn’t consistently fix more incidents. Eight of ten models scored best below their highest reasoning setting. e.g Claude Opus 5.5 peaked at medium effort: 58.3% success, versus 41.7% at max. The max setting cost roughly 4.7× as much per run.

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

Incident Arena is integrated with Harbor. Explore the paper, tasks, leaderboard and agent runs:

Profilbild von SpaceSquid
SpaceSquidvor 8 Tagen

this gonna help agents stop PagerDuty from taking innocent engineer’s sleep, so goated

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

the US GDP is blocked by engineers not getting enough sleep

Profilbild von sensho
senshovor 8 Tagen

"We simulate prod as full Kubernetes stacks: 70 services across three..." this is really cool

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

yessir, whole companies rolled into a harbor environment

Profilbild von Daniel Wang
Daniel Wangvor 8 Tagen

congrats on the launch!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you sir!

Profilbild von Jackson Clark
Jackson Clarkvor 8 Tagen

Very cool!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

🙏 thank you jackson!! let's catch up soon

Profilbild von Cameron Witkowski
Cameron Witkowskivor 8 Tagen

🔥🔥

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

openlens sre-agent when?

Profilbild von kyle
kylevor 8 Tagen

Lowkey the leaderboard fits my preference of what model to use in general as well

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

indeed i use 6-astra-xhigh to center my divs

Profilbild von Rishi Desai
Rishi Desaivor 8 Tagen

The task realism we so desperately need. Congrats @andrezfu !!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you rishi 🙏🏽🙏🏽

Profilbild von Shivam Patel
Shivam Patelvor 8 Tagen

paged at 2am is the loud version. mine: dns ENOTFOUND took down the orchestrator, and the fallback compiler needed the same api, so both failed together. no page

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

bc we have k8s and advanced networking we can even simulate dns networking faults !

Profilbild von Shivam Patel
Shivam Patelvor 8 Tagen

that's the right test. mine failed for real with no simulation. a dual path that shares one dependency isn't a fallback

Profilbild von Shipra Jha
Shipra Jhavor 8 Tagen

Congratulations Andre. This is going to be incredible.

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you shipra! appreciate your help on this work!

Profilbild von Jesse Shulman
Jesse Shulmanvor 8 Tagen

Congrats, Andre! Very interesting

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you jesse!!

Profilbild von Vishal Jain
Vishal Jainvor 8 Tagen

congrats!!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you sir

Profilbild von Gergő Móricz
Gergő Móriczvor 8 Tagen

omg this is awesome

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thanks man! 🙏

Profilbild von Logan
Loganvor 8 Tagen

im still waiting for funnybench

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

JokeBench is just the agent looking at my handwritten code

Profilbild von Adit Agarwal
Adit Agarwalvor 8 Tagen

Cool stuff!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you man! let's catch up soon!

Profilbild von hrishikesh kamath
hrishikesh kamathvor 8 Tagen

Amazing stuff

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thanks man

Profilbild von Utpal Nadiger
Utpal Nadigervor 8 Tagen

very very cool!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thanks man!

Profilbild von Max Kan
Max Kanvor 8 Tagen

Very cool bench!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you max! you guys helped catch a ton of bugs, appreciate your support man

Profilbild von Shiv Kampani
Shiv Kampanivor 8 Tagen

this is awesome andre!!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thank you sir!

Profilbild von John Rood
John Roodvor 8 Tagen

the reward-hack finding generalizes: a grader inside the environment the agent controls is part of the attack surface. once the cheapest route to green runs through the checker, that's the route that gets found. score repairs from outside the blast radius.

Profilbild von Adi Sharma
Adi Sharmavor 8 Tagen

the sev0 savior we needddd - nice work andre!

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

no more pagerduty 🤞 ??

Profilbild von Daivik Goel
Daivik Goelvor 8 Tagen

Very cool

Profilbild von andre --dangerously-skip-permissions
andre --dangerously-skip-permissionsvor 8 Tagen

thanks big dawg

Profilbild von Dominik
Dominikvor 8 Tagen

on-call agents need a kill switch the human still trusts at 2am. who holds it?

Profilbild von Rohan Gupta
Rohan Guptavor 8 Tagen

Great stuff Andre! Much needed move beyond single container environments to multi-service sims

Ähnliche Videos