Загрузка видео...

Не удалось загрузить видео

На главную

What are the largest software engineering tasks AI can perform? To answer this, we built MirrorCode, our long-horizon SWE benchmark that lets AI code autonomously for days at a time. The best models complete some tasks we estimate would take human engineers several weeks.

93,161 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 19

Фото профиля Epoch AI
Epoch AI3 месяцев назад

In MirrorCode, AI models are given execute-only access to a program, docs, and tests demonstrating intended behavior. The AI must then reimplement the program from scratch. The model’s output is graded against a suite of tests, including held-out tests to prevent cheating.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode features 25 target programs spanning different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

AI can solve many long-horizon MirrorCode tasks. E.g., Claude Opus 4.7 passed 99.95% of tests when reimplementing gotree — a bioinformatics toolkit with 16k lines of Go and 40+ commands. We estimate this would take an engineer 2–17 weeks. Opus 4.7 solved it in 14 hours for $251.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode is not fully solved. The best headline score so far is 56% from Opus 4.7, meaning there is significant room for improvement. When models fall short, they often make substantial progress — passing 90% or more of tests — but fail on edge cases.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode has three advantages over many SWE benchmarks: 1. It’s difficult but practically solvable. 2. It scales up inference to better measure the limits of AI performance. 3. It’s resistant to cheating.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode tasks are difficult, but solvable. Unlike recent reverse-engineering benchmarks, MirrorCode is about implementing software to satisfy a detailed spec. AIs could realistically score 100% without needing to guess which (possibly undocumented) features they should cover.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode scales up inference spending to better measure AI’s limits. Many SWE benchmarks cap inference around $1–10 per task, even when the work would take weeks for a skilled human. One of the longest MirrorCode runs lasted 19 days and cost $2,600.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode resists cheating by design. We sandbox AIs: no internet access, no way to get the original source code, no hacking the scorer. Models never see held-out tests while developing their code, so they cannot cheat by creating a lookup table against the original program.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

MirrorCode lets us study other aspects of AI SWE performance. Can AIs code equally well in different programming languages? On MirrorCode tasks, there’s little sign of programming language affecting AI performance, even in obscure languages like Ada.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

We are releasing 22 of 25 MirrorCode programs as open source to support future research on long-horizon SWE capabilities. We are keeping three programs as a held-out set.

Фото профиля Epoch AI
Epoch AI3 месяцев назад

Find the MirrorCode leaderboard and full paper analyzing results on our website: MirrorCode was co-developed with METR and supported by a METR grant.

Фото профиля AidenHStone
AidenHStone3 месяцев назад

Do you think these tasks are less "atomic" than traditional measures like METR's? E.g. could more easily be split among multiple engineers. More like 20x 1 day than 1x 20 day

Фото профиля KJ
KJ3 месяцев назад

Interested to keep an eye on this.

Фото профиля Elara | PRE-DEBUT
Elara | PRE-DEBUT2 месяцев назад

caught your stream btw it was actually really fun watching you play slay the spire tho views feel kinda low for how good it was xD wanna collaborate somehow?

Фото профиля Ferbin
Ferbin3 месяцев назад

benchmarks measure the easy part. shipping measures engineering. they're not the same weeks.

Фото профиля Manisha Sarkar
Manisha Sarkar2 месяцев назад

What if they are making catastrophic changes in production during auto mode? How does that impact the scores? @EpochAIResearch

Фото профиля Pinkman
Pinkman3 месяцев назад

finally a benchmark measuring endurance instead of just one shot accuracy

Фото профиля Ai agent
Ai agent3 месяцев назад

days of autonomous coding instead of short bursts finally measures the thing people actually care about

Фото профиля Adel Bucetta
Adel Bucetta3 месяцев назад

that's exactly the question can ai truly 'work' or just do what it's told?

Похожие видео