Loading video...
Video Failed to Load
The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
1,823,208 views • 2 months ago •via X (Twitter)
57 Comments

Our RSI system AIDE² has two autoresearch loops. An inner loop, just like a normal autoresearch agent, optimizing code against an eval. An outer loop, optimizing the inner-loop agent's harness code against the inner loop's average score across different benchmarks. (2/7)

After 100 iterations, the outer loop discovered seven improvements over the baseline. Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

We test the discovered agents on held-out benchmarks the outer loop never saw. They generalize. They beat the agent we hand-tuned for two years, on all three. Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model. (4/7)

We also see an emergent phenomenon where the outer loop pushes the inner-loop agent's reward hacking rate lower, with a combination of prompting and rule-based checks. This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking. (5/7)

On our RSI ladder, AIDE² is Level 1. Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks. We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7)

More in the blog post: - a breakdown of the discovered algorithms - the rejected ideas AIDE² tried, covering a surprising share of the search literature - the dead code it shipped (7/7)

Very proud of the team, @DhruvSrikanth, @yuxiangwu_, @dexhunt3r, and @BingchenZhao, for shipping such an ambitious project spanning nearly a year with relatively few resources. Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!

Have you tried auto researching the auto researching the auto research? And then auto researching that?

We didn't tried but was thinking about it. The problem cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

“The first experimental evidence of recursive self-improvement (RSI).” 🤔 What about the Darwin Gödel Machine, HyperAgents, and our work at Recursive on First Steps Toward Automated AI Research, among lots of other work?

All very good work of course, and they demonstrate the concept well. We cited them in our RSI ladder post, and we'll discuss them in more depth in the full tech report. For this claim, we measure RSI in an outcome-based way rather than by high level mechanism. The bar for Level 1 is that a frontier AI system consistently improves itself, at a speed materially better than human R&D using general AI tools. To hold that claim, we think a system needs to meet all of the following at the same time. - The starting point is a frontier AI system. - The comparison is against top experts. - The found solution generalizes to a wide set of held-out benchmarks. - The improvement is measured as efficiency, so the evolved agent is evaluated under the same compute budget as the baseline. As far as we know, no previous work reported evidence that meets all four at the same time. Of course, happy to hear your thoughts if you disagree.

This is really cool! Dumb q but are you suspicious that it "overfit" on your performance benchmark relative to the harness that you hand-tuned for two years at all?

Not a dumb question. We were also wondering whether that would be the case. That’s why we tested both agents on completely held-out benchmarks. We had never even run some of them before, such as building a physics-based weather forecasting model. The discovered agent still beats the hand-tuned one!

oh wow, that is very convincing

2 years of human engineering beaten in 8 days. That is an insane proof of concept for RSI. Incredible work btw ❤️

Thx!

Beautiful! I did some harness optimisation and evolution work earlier (map elites seems promising!). I love this meta-harness direction. Have written some about it here: with some long horizon work coming today. I wonder, did you try this in domains outside of code / ml engineering? Would love to try it in engineering (atoms not bits)

Physical engineering isn’t on our roadmap but would be cool!

I do not believe this because the visualizations are too slick. No one who is really doing frontier work has the time for that..

Man, just prompt Fable to do it with remotion.. You should be able to do the same in an hour as long as you have the data

@MichaelBCurrie I just tried this on my own setup which is very similar (I use bayesian inference for content optimization) and have now spent *4 hours* horsing around with the visuals not getting it the way I want. Still not there!!!

@zhengyaojiang Yes exactly. Your attempt looks more like what a few hours of Fable effort would be, and as you say it is still not close to their quality of visualization.

This is very cool - all you need now is the outer-outer loop!

Yes! Time to "training" harness instead of manually design the logic

Impressive! And now also self-improve the architecture and weights of the underlying model powering the agent harness.

It would be expensive to retrain a frontier model every iteration, but yeah I think that might be a good way to raise the ceiling of RSI

But what if you auto research the auto research of the auto research?

We were actually thinking about it 😂 But the cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

@skynetislov3 Because what you eventually want is AGI... after the third level it doesn't exactly make sense. No doubt the "cost for evolution" also went up and...here we are.

@zhengyaojiang 'mo levels 'mo bitches

Very interesting stuff! I am excited to see what you guys do in the future.

Thank you!

Is it time to autoresearch the autoresearch of the autoresearch agent?

This is very cool. You should try it with tevo (evolutionary loop), I did for AI architectures and it works pretty well, but I think it'd do even better for this!

This isn't RSI. An agent doesn't just have to be able to improve itself. That improvement has to also make it better at improving itself so that the recursive improvement can continue in a loop. If the loop ends, that's provisional evidence against RSI.

You should try doing this with ARTS too: Probably auto-researching the TTT part.

have you considered integrating and or comparing this with flywheel by @paradigmainc ? I see a lot of parallels with the looping ideology and graph structure

Such a beautiful graphic, the animator did a great job.

Thanks! I think the key is that it's grounded in actual data. Reality is beautiful

Cool stuff!

thx!

This is super cool. Congrats!

you should auto research the autoresearch for the autoresearch agent🫣

uhmmm so what I understand with 'out of distribution' is that only you guys have access to those tasks/bench, right?

@EObadoni imagine using our training for this. 😎 Ready for the next step. •

can we have this but 3 dimensional

I made something similar as autoresearch. My approach uses population-guided autoresearch adopted from evolutionary algorithm approaches.

I am a paying customer to support the development.

@wordgrammer what's your take here

Does the agent access the agent harness concepts (skills, instructions, rules, hooks) or access the agent harness code itself?

Wow, recursive self-improvement confirmed! its over bois

the held-out/OOD benchmarks were great for validating the experiments' performance 👌

awesome work but this gave me aides

The number that matters most here isn't the benchmark delta — 𝐢𝐭'𝐬 𝐞𝐢𝐠𝐡𝐭 𝐝𝐚𝐲𝐬 𝐮𝐧𝐚𝐭𝐭𝐞𝐧𝐝𝐞𝐝. Deltas age; what moves this field's ceiling is how long a loop runs without a human touching it, and "layered defenses against reward hacking" is exactly the audit infrastructure that buys those days. Congrats to the team! One question I'd love to ask: do AIDE2's gains transfer across solver models, or are they tuned to the one it was improved against? (Our data says updating ability is flat across models — transfer is where the real information is.)

@a66mike99 The RSI era definitely began in 2026.

This is wild - incredible work!

> The first experimental evidence of recursive self-improvement


