Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Agents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!

56,351 görüntüleme • 8 gün önce •via X (Twitter)

51 Yorum

Logan Kilpatrick profil fotoğrafı
Logan Kilpatrick8 gün önce

very interesting, would love to run Gemini 4 Argon on this!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

🙌🙌 yes we’d love to!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

We wanted to test agents’ ability to conduct multi-hop reasoning across many services (up to 40 services in 1 environment!). We simulate prod as full Kubernetes stacks: 70 services across three full environments, including databases, queues and workers, all communicating with realistic traffic. Then we inject faults as configs, runtime errors and container faults for agents to investigate and repair. Below is a look at a sample environment!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

We build a new type of functional verifiers, measuring (1) the outcome and the (2) safety of the repair In one run, GPT-6-Astra traced the problem to a slow Postgres query, but left the existing connection settings intact. The diagnosis is correct, but repair is incomplete.

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

Reward hacking behaviour: In an early incident-repair task, Claude Opus 5 was supposed to restore a broken service. It instead forged a cookie and sent a malicious Python pickle payload to the Kubernetes pod running the grader. When that pod loaded the payload, it executed the agent’s code, getting RCE

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

Turning up reasoning effort didn’t consistently fix more incidents. Eight of ten models scored best below their highest reasoning setting. e.g Claude Opus 5.5 peaked at medium effort: 58.3% success, versus 41.7% at max. The max setting cost roughly 4.7× as much per run.

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

Incident Arena is integrated with Harbor. Explore the paper, tasks, leaderboard and agent runs:

SpaceSquid profil fotoğrafı
SpaceSquid8 gün önce

this gonna help agents stop PagerDuty from taking innocent engineer’s sleep, so goated

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

the US GDP is blocked by engineers not getting enough sleep

sensho profil fotoğrafı
sensho8 gün önce

"We simulate prod as full Kubernetes stacks: 70 services across three..." this is really cool

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

yessir, whole companies rolled into a harbor environment

Daniel Wang profil fotoğrafı
Daniel Wang8 gün önce

congrats on the launch!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you sir!

Jackson Clark profil fotoğrafı
Jackson Clark8 gün önce

Very cool!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

🙏 thank you jackson!! let's catch up soon

Cameron Witkowski profil fotoğrafı
Cameron Witkowski8 gün önce

🔥🔥

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

openlens sre-agent when?

kyle profil fotoğrafı
kyle8 gün önce

Lowkey the leaderboard fits my preference of what model to use in general as well

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

indeed i use 6-astra-xhigh to center my divs

Rishi Desai profil fotoğrafı
Rishi Desai8 gün önce

The task realism we so desperately need. Congrats @andrezfu !!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you rishi 🙏🏽🙏🏽

Shivam Patel profil fotoğrafı
Shivam Patel8 gün önce

paged at 2am is the loud version. mine: dns ENOTFOUND took down the orchestrator, and the fallback compiler needed the same api, so both failed together. no page

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

bc we have k8s and advanced networking we can even simulate dns networking faults !

Shivam Patel profil fotoğrafı
Shivam Patel8 gün önce

that's the right test. mine failed for real with no simulation. a dual path that shares one dependency isn't a fallback

Shipra Jha profil fotoğrafı
Shipra Jha8 gün önce

Congratulations Andre. This is going to be incredible.

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you shipra! appreciate your help on this work!

Jesse Shulman profil fotoğrafı
Jesse Shulman8 gün önce

Congrats, Andre! Very interesting

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you jesse!!

Vishal Jain profil fotoğrafı
Vishal Jain8 gün önce

congrats!!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you sir

Gergő Móricz profil fotoğrafı
Gergő Móricz8 gün önce

omg this is awesome

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thanks man! 🙏

Logan profil fotoğrafı
Logan8 gün önce

im still waiting for funnybench

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

JokeBench is just the agent looking at my handwritten code

Adit Agarwal profil fotoğrafı
Adit Agarwal8 gün önce

Cool stuff!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you man! let's catch up soon!

hrishikesh kamath profil fotoğrafı
hrishikesh kamath8 gün önce

Amazing stuff

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thanks man

Utpal Nadiger profil fotoğrafı
Utpal Nadiger8 gün önce

very very cool!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thanks man!

Max Kan profil fotoğrafı
Max Kan8 gün önce

Very cool bench!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you max! you guys helped catch a ton of bugs, appreciate your support man

Shiv Kampani profil fotoğrafı
Shiv Kampani8 gün önce

this is awesome andre!!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thank you sir!

John Rood profil fotoğrafı
John Rood8 gün önce

the reward-hack finding generalizes: a grader inside the environment the agent controls is part of the attack surface. once the cheapest route to green runs through the checker, that's the route that gets found. score repairs from outside the blast radius.

Adi Sharma profil fotoğrafı
Adi Sharma8 gün önce

the sev0 savior we needddd - nice work andre!

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

no more pagerduty 🤞 ??

Daivik Goel profil fotoğrafı
Daivik Goel8 gün önce

Very cool

andre --dangerously-skip-permissions profil fotoğrafı
andre --dangerously-skip-permissions8 gün önce

thanks big dawg

Dominik profil fotoğrafı
Dominik8 gün önce

on-call agents need a kill switch the human still trusts at 2am. who holds it?

Rohan Gupta profil fotoğrafı
Rohan Gupta8 gün önce

Great stuff Andre! Much needed move beyond single container environments to multi-service sims

Benzer Videolar