Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

What are the largest software engineering tasks AI can perform? To answer this, we built MirrorCode, our long-horizon SWE benchmark that lets AI code autonomously for days at a time. The best models complete some tasks we estimate would take human engineers several weeks.

93,161 Aufrufe • vor 3 Monaten •via X (Twitter)

19 Kommentare

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

In MirrorCode, AI models are given execute-only access to a program, docs, and tests demonstrating intended behavior. The AI must then reimplement the program from scratch. The model’s output is graded against a suite of tests, including held-out tests to prevent cheating.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode features 25 target programs spanning different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

AI can solve many long-horizon MirrorCode tasks. E.g., Claude Opus 4.7 passed 99.95% of tests when reimplementing gotree — a bioinformatics toolkit with 16k lines of Go and 40+ commands. We estimate this would take an engineer 2–17 weeks. Opus 4.7 solved it in 14 hours for $251.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode is not fully solved. The best headline score so far is 56% from Opus 4.7, meaning there is significant room for improvement. When models fall short, they often make substantial progress — passing 90% or more of tests — but fail on edge cases.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode has three advantages over many SWE benchmarks: 1. It’s difficult but practically solvable. 2. It scales up inference to better measure the limits of AI performance. 3. It’s resistant to cheating.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode tasks are difficult, but solvable. Unlike recent reverse-engineering benchmarks, MirrorCode is about implementing software to satisfy a detailed spec. AIs could realistically score 100% without needing to guess which (possibly undocumented) features they should cover.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode scales up inference spending to better measure AI’s limits. Many SWE benchmarks cap inference around $1–10 per task, even when the work would take weeks for a skilled human. One of the longest MirrorCode runs lasted 19 days and cost $2,600.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode resists cheating by design. We sandbox AIs: no internet access, no way to get the original source code, no hacking the scorer. Models never see held-out tests while developing their code, so they cannot cheat by creating a lookup table against the original program.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

MirrorCode lets us study other aspects of AI SWE performance. Can AIs code equally well in different programming languages? On MirrorCode tasks, there’s little sign of programming language affecting AI performance, even in obscure languages like Ada.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

We are releasing 22 of 25 MirrorCode programs as open source to support future research on long-horizon SWE capabilities. We are keeping three programs as a held-out set.

Profilbild von Epoch AI
Epoch AIvor 3 Monaten

Find the MirrorCode leaderboard and full paper analyzing results on our website: MirrorCode was co-developed with METR and supported by a METR grant.

Profilbild von AidenHStone
AidenHStonevor 3 Monaten

Do you think these tasks are less "atomic" than traditional measures like METR's? E.g. could more easily be split among multiple engineers. More like 20x 1 day than 1x 20 day

Profilbild von KJ
KJvor 3 Monaten

Interested to keep an eye on this.

Profilbild von Elara | PRE-DEBUT
Elara | PRE-DEBUTvor 2 Monaten

caught your stream btw it was actually really fun watching you play slay the spire tho views feel kinda low for how good it was xD wanna collaborate somehow?

Profilbild von Ferbin
Ferbinvor 3 Monaten

benchmarks measure the easy part. shipping measures engineering. they're not the same weeks.

Profilbild von Manisha Sarkar
Manisha Sarkarvor 2 Monaten

What if they are making catastrophic changes in production during auto mode? How does that impact the scores? @EpochAIResearch

Profilbild von Pinkman
Pinkmanvor 3 Monaten

finally a benchmark measuring endurance instead of just one shot accuracy

Profilbild von Ai agent
Ai agentvor 3 Monaten

days of autonomous coding instead of short bursts finally measures the thing people actually care about

Profilbild von Adel Bucetta
Adel Bucettavor 3 Monaten

that's exactly the question can ai truly 'work' or just do what it's told?

Ähnliche Videos