Video wird geladen...
Video konnte nicht geladen werden
What are the largest software engineering tasks AI can perform? To answer this, we built MirrorCode, our long-horizon SWE benchmark that lets AI code autonomously for days at a time. The best models complete some tasks we estimate would take human engineers several weeks.
93,161 Aufrufe • vor 3 Monaten •via X (Twitter)
19 Kommentare

In MirrorCode, AI models are given execute-only access to a program, docs, and tests demonstrating intended behavior. The AI must then reimplement the program from scratch. The model’s output is graded against a suite of tests, including held-out tests to prevent cheating.

MirrorCode features 25 target programs spanning different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.

AI can solve many long-horizon MirrorCode tasks. E.g., Claude Opus 4.7 passed 99.95% of tests when reimplementing gotree — a bioinformatics toolkit with 16k lines of Go and 40+ commands. We estimate this would take an engineer 2–17 weeks. Opus 4.7 solved it in 14 hours for $251.

MirrorCode is not fully solved. The best headline score so far is 56% from Opus 4.7, meaning there is significant room for improvement. When models fall short, they often make substantial progress — passing 90% or more of tests — but fail on edge cases.

MirrorCode has three advantages over many SWE benchmarks: 1. It’s difficult but practically solvable. 2. It scales up inference to better measure the limits of AI performance. 3. It’s resistant to cheating.

MirrorCode tasks are difficult, but solvable. Unlike recent reverse-engineering benchmarks, MirrorCode is about implementing software to satisfy a detailed spec. AIs could realistically score 100% without needing to guess which (possibly undocumented) features they should cover.

MirrorCode scales up inference spending to better measure AI’s limits. Many SWE benchmarks cap inference around $1–10 per task, even when the work would take weeks for a skilled human. One of the longest MirrorCode runs lasted 19 days and cost $2,600.

MirrorCode resists cheating by design. We sandbox AIs: no internet access, no way to get the original source code, no hacking the scorer. Models never see held-out tests while developing their code, so they cannot cheat by creating a lookup table against the original program.

MirrorCode lets us study other aspects of AI SWE performance. Can AIs code equally well in different programming languages? On MirrorCode tasks, there’s little sign of programming language affecting AI performance, even in obscure languages like Ada.

We are releasing 22 of 25 MirrorCode programs as open source to support future research on long-horizon SWE capabilities. We are keeping three programs as a held-out set.

Find the MirrorCode leaderboard and full paper analyzing results on our website: MirrorCode was co-developed with METR and supported by a METR grant.

Do you think these tasks are less "atomic" than traditional measures like METR's? E.g. could more easily be split among multiple engineers. More like 20x 1 day than 1x 20 day

Interested to keep an eye on this.

caught your stream btw it was actually really fun watching you play slay the spire tho views feel kinda low for how good it was xD wanna collaborate somehow?

benchmarks measure the easy part. shipping measures engineering. they're not the same weeks.

What if they are making catastrophic changes in production during auto mode? How does that impact the scores? @EpochAIResearch

finally a benchmark measuring endurance instead of just one shot accuracy

days of autonomous coding instead of short bursts finally measures the thing people actually care about

that's exactly the question can ai truly 'work' or just do what it's told?

