Loading video...

Video Failed to Load

Go Home

The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)

1,823,208 views • 2 months ago •via X (Twitter)

57 Comments

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Our RSI system AIDE² has two autoresearch loops. An inner loop, just like a normal autoresearch agent, optimizing code against an eval. An outer loop, optimizing the inner-loop agent's harness code against the inner loop's average score across different benchmarks. (2/7)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

After 100 iterations, the outer loop discovered seven improvements over the baseline. Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

We test the discovered agents on held-out benchmarks the outer loop never saw. They generalize. They beat the agent we hand-tuned for two years, on all three. Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model. (4/7)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

We also see an emergent phenomenon where the outer loop pushes the inner-loop agent's reward hacking rate lower, with a combination of prompting and rule-based checks. This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking. (5/7)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

On our RSI ladder, AIDE² is Level 1. Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks. We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

More in the blog post: - a breakdown of the discovered algorithms - the rejected ideas AIDE² tried, covering a surprising share of the search literature - the dead code it shipped (7/7)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Very proud of the team, @DhruvSrikanth, @yuxiangwu_, @dexhunt3r, and @BingchenZhao, for shipping such an ambitious project spanning nearly a year with relatively few resources. Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!

Shiny Gen Wizard's profile picture
Shiny Gen Wizard2 months ago

Have you tried auto researching the auto researching the auto research? And then auto researching that?

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

We didn't tried but was thinking about it. The problem cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Jeff Clune's profile picture
Jeff Clune2 months ago

“The first experimental evidence of recursive self-improvement (RSI).” 🤔 What about the Darwin Gödel Machine, HyperAgents, and our work at Recursive on First Steps Toward Automated AI Research, among lots of other work?

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

All very good work of course, and they demonstrate the concept well. We cited them in our RSI ladder post, and we'll discuss them in more depth in the full tech report. For this claim, we measure RSI in an outcome-based way rather than by high level mechanism. The bar for Level 1 is that a frontier AI system consistently improves itself, at a speed materially better than human R&D using general AI tools. To hold that claim, we think a system needs to meet all of the following at the same time. - The starting point is a frontier AI system. - The comparison is against top experts. - The found solution generalizes to a wide set of held-out benchmarks. - The improvement is measured as efficiency, so the evolved agent is evaluated under the same compute budget as the baseline. As far as we know, no previous work reported evidence that meets all four at the same time. Of course, happy to hear your thoughts if you disagree.

Henry Dowling's profile picture
Henry Dowling2 months ago

This is really cool! Dumb q but are you suspicious that it "overfit" on your performance benchmark relative to the harness that you hand-tuned for two years at all?

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Not a dumb question. We were also wondering whether that would be the case. That’s why we tested both agents on completely held-out benchmarks. We had never even run some of them before, such as building a physics-based weather forecasting model. The discovered agent still beats the hand-tuned one!

Henry Dowling's profile picture
Henry Dowling2 months ago

oh wow, that is very convincing

Obedience Adara's profile picture
Obedience Adara2 months ago

2 years of human engineering beaten in 8 days. That is an insane proof of concept for RSI. Incredible work btw ❤️

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Thx!

Theodore Galanos's profile picture
Theodore Galanos2 months ago

Beautiful! I did some harness optimisation and evolution work earlier (map elites seems promising!). I love this meta-harness direction. Have written some about it here: with some long horizon work coming today. I wonder, did you try this in domains outside of code / ml engineering? Would love to try it in engineering (atoms not bits)

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Physical engineering isn’t on our roadmap but would be cool!

Michael B. Currie's profile picture
Michael B. Currie2 months ago

I do not believe this because the visualizations are too slick. No one who is really doing frontier work has the time for that..

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Man, just prompt Fable to do it with remotion.. You should be able to do the same in an hour as long as you have the data

DEJAN's profile picture
DEJAN2 months ago

@MichaelBCurrie I just tried this on my own setup which is very similar (I use bayesian inference for content optimization) and have now spent *4 hours* horsing around with the visuals not getting it the way I want. Still not there!!!

Michael B. Currie's profile picture
Michael B. Currie2 months ago

@zhengyaojiang Yes exactly. Your attempt looks more like what a few hours of Fable effort would be, and as you say it is still not close to their quality of visualization.

Fred Jonsson's profile picture
Fred Jonsson2 months ago

This is very cool - all you need now is the outer-outer loop!

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Yes! Time to "training" harness instead of manually design the logic

Burny - Effective Curiosity's profile picture
Burny - Effective Curiosity2 months ago

Impressive! And now also self-improve the architecture and weights of the underlying model powering the agent harness.

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

It would be expensive to retrain a frontier model every iteration, but yeah I think that might be a good way to raise the ceiling of RSI

gpu go brr...'s profile picture
gpu go brr...2 months ago

But what if you auto research the auto research of the auto research?

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

We were actually thinking about it 😂 But the cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Pawel Pachniewski's profile picture
Pawel Pachniewski2 months ago

@skynetislov3 Because what you eventually want is AGI... after the third level it doesn't exactly make sense. No doubt the "cost for evolution" also went up and...here we are.

gpu go brr...'s profile picture
gpu go brr...2 months ago

@zhengyaojiang 'mo levels 'mo bitches

Nisan Chhetri's profile picture
Nisan Chhetri2 months ago

Very interesting stuff! I am excited to see what you guys do in the future.

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Thank you!

Oliver Wang's profile picture
Oliver Wang2 months ago

Is it time to autoresearch the autoresearch of the autoresearch agent?

rohit's profile picture
rohit2 months ago

This is very cool. You should try it with tevo (evolutionary loop), I did for AI architectures and it works pretty well, but I think it'd do even better for this!

Kevin's profile picture
Kevin2 months ago

This isn't RSI. An agent doesn't just have to be able to improve itself. That improvement has to also make it better at improving itself so that the recursive improvement can continue in a loop. If the loop ends, that's provisional evidence against RSI.

Gurusha Juneja (✈️ ICML'26)'s profile picture
Gurusha Juneja (✈️ ICML'26)2 months ago

You should try doing this with ARTS too: Probably auto-researching the TTT part.

Theo's profile picture
Theo2 months ago

have you considered integrating and or comparing this with flywheel by @paradigmainc ? I see a lot of parallels with the looping ideology and graph structure

420Trades's profile picture
420Trades2 months ago

Such a beautiful graphic, the animator did a great job.

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

Thanks! I think the key is that it's grounded in actual data. Reality is beautiful

Alok Bishoyi's profile picture
Alok Bishoyi2 months ago

Cool stuff!

Zhengyao Jiang's profile picture
Zhengyao Jiang2 months ago

thx!

Ramez Naam's profile picture
Ramez Naam2 months ago

This is super cool. Congrats!

Dev's profile picture
Dev2 months ago

you should auto research the autoresearch for the autoresearch agent🫣

ismaelvega's profile picture
ismaelvega2 months ago

uhmmm so what I understand with 'out of distribution' is that only you guys have access to those tasks/bench, right?

Kirk Patrick Miller's profile picture
Kirk Patrick Miller2 months ago

@EObadoni imagine using our training for this. 😎 Ready for the next step. •

forloop's profile picture
forloop2 months ago

can we have this but 3 dimensional

Tino Wening's profile picture
Tino Wening2 months ago

I made something similar as autoresearch. My approach uses population-guided autoresearch adopted from evolutionary algorithm approaches.

La View Claire's profile picture
La View Claire2 months ago

I am a paying customer to support the development.

tautologer's profile picture
tautologer2 months ago

@wordgrammer what's your take here

redgreenblue's profile picture
redgreenblue2 months ago

Does the agent access the agent harness concepts (skills, instructions, rules, hooks) or access the agent harness code itself?

chris's profile picture
chris2 months ago

Wow, recursive self-improvement confirmed! its over bois

Morgan McGuire's profile picture
Morgan McGuire2 months ago

the held-out/OOD benchmarks were great for validating the experiments' performance 👌

Shiv Kampani's profile picture
Shiv Kampani2 months ago

awesome work but this gave me aides

Henry Lu's profile picture
Henry Lu2 months ago

The number that matters most here isn't the benchmark delta — 𝐢𝐭'𝐬 𝐞𝐢𝐠𝐡𝐭 𝐝𝐚𝐲𝐬 𝐮𝐧𝐚𝐭𝐭𝐞𝐧𝐝𝐞𝐝. Deltas age; what moves this field's ceiling is how long a loop runs without a human touching it, and "layered defenses against reward hacking" is exactly the audit infrastructure that buys those days. Congrats to the team! One question I'd love to ask: do AIDE2's gains transfer across solver models, or are they tuned to the one it was improved against? (Our data says updating ability is flat across models — transfer is where the real information is.)

Larry Panozzo's profile picture
Larry Panozzo2 months ago

@a66mike99 The RSI era definitely began in 2026.

arya 🏴's profile picture
arya 🏴2 months ago

This is wild - incredible work!

jsd's profile picture
jsd2 months ago

> The first experimental evidence of recursive self-improvement

Related Videos