Loading video...

Video Failed to Load

Go Home

What are the largest software engineering tasks AI can perform? To answer this, we built MirrorCode, our long-horizon SWE benchmark that lets AI code autonomously for days at a time. The best models complete some tasks we estimate would take human engineers several weeks.

93,161 views • 3 months ago •via X (Twitter)

19 Comments

Epoch AI's profile picture
Epoch AI3 months ago

In MirrorCode, AI models are given execute-only access to a program, docs, and tests demonstrating intended behavior. The AI must then reimplement the program from scratch. The model’s output is graded against a suite of tests, including held-out tests to prevent cheating.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode features 25 target programs spanning different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.

Epoch AI's profile picture
Epoch AI3 months ago

AI can solve many long-horizon MirrorCode tasks. E.g., Claude Opus 4.7 passed 99.95% of tests when reimplementing gotree — a bioinformatics toolkit with 16k lines of Go and 40+ commands. We estimate this would take an engineer 2–17 weeks. Opus 4.7 solved it in 14 hours for $251.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode is not fully solved. The best headline score so far is 56% from Opus 4.7, meaning there is significant room for improvement. When models fall short, they often make substantial progress — passing 90% or more of tests — but fail on edge cases.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode has three advantages over many SWE benchmarks: 1. It’s difficult but practically solvable. 2. It scales up inference to better measure the limits of AI performance. 3. It’s resistant to cheating.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode tasks are difficult, but solvable. Unlike recent reverse-engineering benchmarks, MirrorCode is about implementing software to satisfy a detailed spec. AIs could realistically score 100% without needing to guess which (possibly undocumented) features they should cover.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode scales up inference spending to better measure AI’s limits. Many SWE benchmarks cap inference around $1–10 per task, even when the work would take weeks for a skilled human. One of the longest MirrorCode runs lasted 19 days and cost $2,600.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode resists cheating by design. We sandbox AIs: no internet access, no way to get the original source code, no hacking the scorer. Models never see held-out tests while developing their code, so they cannot cheat by creating a lookup table against the original program.

Epoch AI's profile picture
Epoch AI3 months ago

MirrorCode lets us study other aspects of AI SWE performance. Can AIs code equally well in different programming languages? On MirrorCode tasks, there’s little sign of programming language affecting AI performance, even in obscure languages like Ada.

Epoch AI's profile picture
Epoch AI3 months ago

We are releasing 22 of 25 MirrorCode programs as open source to support future research on long-horizon SWE capabilities. We are keeping three programs as a held-out set.

Epoch AI's profile picture
Epoch AI3 months ago

Find the MirrorCode leaderboard and full paper analyzing results on our website: MirrorCode was co-developed with METR and supported by a METR grant.

AidenHStone's profile picture
AidenHStone3 months ago

Do you think these tasks are less "atomic" than traditional measures like METR's? E.g. could more easily be split among multiple engineers. More like 20x 1 day than 1x 20 day

KJ's profile picture
KJ3 months ago

Interested to keep an eye on this.

Elara | PRE-DEBUT's profile picture
Elara | PRE-DEBUT2 months ago

caught your stream btw it was actually really fun watching you play slay the spire tho views feel kinda low for how good it was xD wanna collaborate somehow?

Ferbin's profile picture
Ferbin3 months ago

benchmarks measure the easy part. shipping measures engineering. they're not the same weeks.

Manisha Sarkar's profile picture
Manisha Sarkar2 months ago

What if they are making catastrophic changes in production during auto mode? How does that impact the scores? @EpochAIResearch

Pinkman's profile picture
Pinkman3 months ago

finally a benchmark measuring endurance instead of just one shot accuracy

Ai agent's profile picture
Ai agent3 months ago

days of autonomous coding instead of short bursts finally measures the thing people actually care about

Adel Bucetta's profile picture
Adel Bucetta3 months ago

that's exactly the question can ai truly 'work' or just do what it's told?

Related Videos