Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)

1,823,208 Aufrufe • vor 2 Monaten •via X (Twitter)

57 Kommentare

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Our RSI system AIDE² has two autoresearch loops. An inner loop, just like a normal autoresearch agent, optimizing code against an eval. An outer loop, optimizing the inner-loop agent's harness code against the inner loop's average score across different benchmarks. (2/7)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

After 100 iterations, the outer loop discovered seven improvements over the baseline. Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

We test the discovered agents on held-out benchmarks the outer loop never saw. They generalize. They beat the agent we hand-tuned for two years, on all three. Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model. (4/7)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

We also see an emergent phenomenon where the outer loop pushes the inner-loop agent's reward hacking rate lower, with a combination of prompting and rule-based checks. This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking. (5/7)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

On our RSI ladder, AIDE² is Level 1. Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks. We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

More in the blog post: - a breakdown of the discovered algorithms - the rejected ideas AIDE² tried, covering a surprising share of the search literature - the dead code it shipped (7/7)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Very proud of the team, @DhruvSrikanth, @yuxiangwu_, @dexhunt3r, and @BingchenZhao, for shipping such an ambitious project spanning nearly a year with relatively few resources. Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!

Profilbild von Shiny Gen Wizard
Shiny Gen Wizardvor 2 Monaten

Have you tried auto researching the auto researching the auto research? And then auto researching that?

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

We didn't tried but was thinking about it. The problem cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Profilbild von Jeff Clune
Jeff Clunevor 2 Monaten

“The first experimental evidence of recursive self-improvement (RSI).” 🤔 What about the Darwin Gödel Machine, HyperAgents, and our work at Recursive on First Steps Toward Automated AI Research, among lots of other work?

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

All very good work of course, and they demonstrate the concept well. We cited them in our RSI ladder post, and we'll discuss them in more depth in the full tech report. For this claim, we measure RSI in an outcome-based way rather than by high level mechanism. The bar for Level 1 is that a frontier AI system consistently improves itself, at a speed materially better than human R&D using general AI tools. To hold that claim, we think a system needs to meet all of the following at the same time. - The starting point is a frontier AI system. - The comparison is against top experts. - The found solution generalizes to a wide set of held-out benchmarks. - The improvement is measured as efficiency, so the evolved agent is evaluated under the same compute budget as the baseline. As far as we know, no previous work reported evidence that meets all four at the same time. Of course, happy to hear your thoughts if you disagree.

Profilbild von Henry Dowling
Henry Dowlingvor 2 Monaten

This is really cool! Dumb q but are you suspicious that it "overfit" on your performance benchmark relative to the harness that you hand-tuned for two years at all?

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Not a dumb question. We were also wondering whether that would be the case. That’s why we tested both agents on completely held-out benchmarks. We had never even run some of them before, such as building a physics-based weather forecasting model. The discovered agent still beats the hand-tuned one!

Profilbild von Henry Dowling
Henry Dowlingvor 2 Monaten

oh wow, that is very convincing

Profilbild von Obedience Adara
Obedience Adaravor 2 Monaten

2 years of human engineering beaten in 8 days. That is an insane proof of concept for RSI. Incredible work btw ❤️

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Thx!

Profilbild von Theodore Galanos
Theodore Galanosvor 2 Monaten

Beautiful! I did some harness optimisation and evolution work earlier (map elites seems promising!). I love this meta-harness direction. Have written some about it here: with some long horizon work coming today. I wonder, did you try this in domains outside of code / ml engineering? Would love to try it in engineering (atoms not bits)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Physical engineering isn’t on our roadmap but would be cool!

Profilbild von Michael B. Currie
Michael B. Currievor 2 Monaten

I do not believe this because the visualizations are too slick. No one who is really doing frontier work has the time for that..

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Man, just prompt Fable to do it with remotion.. You should be able to do the same in an hour as long as you have the data

Profilbild von DEJAN
DEJANvor 2 Monaten

@MichaelBCurrie I just tried this on my own setup which is very similar (I use bayesian inference for content optimization) and have now spent *4 hours* horsing around with the visuals not getting it the way I want. Still not there!!!

Profilbild von Michael B. Currie
Michael B. Currievor 2 Monaten

@zhengyaojiang Yes exactly. Your attempt looks more like what a few hours of Fable effort would be, and as you say it is still not close to their quality of visualization.

Profilbild von Fred Jonsson
Fred Jonssonvor 2 Monaten

This is very cool - all you need now is the outer-outer loop!

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Yes! Time to "training" harness instead of manually design the logic

Profilbild von Burny - Effective Curiosity
Burny - Effective Curiosityvor 2 Monaten

Impressive! And now also self-improve the architecture and weights of the underlying model powering the agent harness.

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

It would be expensive to retrain a frontier model every iteration, but yeah I think that might be a good way to raise the ceiling of RSI

Profilbild von gpu go brr...
gpu go brr...vor 2 Monaten

But what if you auto research the auto research of the auto research?

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

We were actually thinking about it 😂 But the cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Profilbild von Pawel Pachniewski
Pawel Pachniewskivor 2 Monaten

@skynetislov3 Because what you eventually want is AGI... after the third level it doesn't exactly make sense. No doubt the "cost for evolution" also went up and...here we are.

Profilbild von gpu go brr...
gpu go brr...vor 2 Monaten

@zhengyaojiang 'mo levels 'mo bitches

Profilbild von Nisan Chhetri
Nisan Chhetrivor 2 Monaten

Very interesting stuff! I am excited to see what you guys do in the future.

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Thank you!

Profilbild von Oliver Wang
Oliver Wangvor 2 Monaten

Is it time to autoresearch the autoresearch of the autoresearch agent?

Profilbild von rohit
rohitvor 2 Monaten

This is very cool. You should try it with tevo (evolutionary loop), I did for AI architectures and it works pretty well, but I think it'd do even better for this!

Profilbild von Kevin
Kevinvor 2 Monaten

This isn't RSI. An agent doesn't just have to be able to improve itself. That improvement has to also make it better at improving itself so that the recursive improvement can continue in a loop. If the loop ends, that's provisional evidence against RSI.

Profilbild von Gurusha Juneja (✈️ ICML'26)
Gurusha Juneja (✈️ ICML'26)vor 2 Monaten

You should try doing this with ARTS too: Probably auto-researching the TTT part.

Profilbild von Theo
Theovor 2 Monaten

have you considered integrating and or comparing this with flywheel by @paradigmainc ? I see a lot of parallels with the looping ideology and graph structure

Profilbild von 420Trades
420Tradesvor 2 Monaten

Such a beautiful graphic, the animator did a great job.

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Thanks! I think the key is that it's grounded in actual data. Reality is beautiful

Profilbild von Alok Bishoyi
Alok Bishoyivor 2 Monaten

Cool stuff!

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

thx!

Profilbild von Ramez Naam
Ramez Naamvor 2 Monaten

This is super cool. Congrats!

Profilbild von Dev
Devvor 2 Monaten

you should auto research the autoresearch for the autoresearch agent🫣

Profilbild von ismaelvega
ismaelvegavor 2 Monaten

uhmmm so what I understand with 'out of distribution' is that only you guys have access to those tasks/bench, right?

Profilbild von Kirk Patrick Miller
Kirk Patrick Millervor 2 Monaten

@EObadoni imagine using our training for this. 😎 Ready for the next step. •

Profilbild von forloop
forloopvor 2 Monaten

can we have this but 3 dimensional

Profilbild von Tino Wening
Tino Weningvor 2 Monaten

I made something similar as autoresearch. My approach uses population-guided autoresearch adopted from evolutionary algorithm approaches.

Profilbild von La View Claire
La View Clairevor 2 Monaten

I am a paying customer to support the development.

Profilbild von tautologer
tautologervor 2 Monaten

@wordgrammer what's your take here

Profilbild von redgreenblue
redgreenbluevor 2 Monaten

Does the agent access the agent harness concepts (skills, instructions, rules, hooks) or access the agent harness code itself?

Profilbild von chris
chrisvor 2 Monaten

Wow, recursive self-improvement confirmed! its over bois

Profilbild von Morgan McGuire
Morgan McGuirevor 2 Monaten

the held-out/OOD benchmarks were great for validating the experiments' performance 👌

Profilbild von Shiv Kampani
Shiv Kampanivor 2 Monaten

awesome work but this gave me aides

Profilbild von Henry Lu
Henry Luvor 2 Monaten

The number that matters most here isn't the benchmark delta — 𝐢𝐭'𝐬 𝐞𝐢𝐠𝐡𝐭 𝐝𝐚𝐲𝐬 𝐮𝐧𝐚𝐭𝐭𝐞𝐧𝐝𝐞𝐝. Deltas age; what moves this field's ceiling is how long a loop runs without a human touching it, and "layered defenses against reward hacking" is exactly the audit infrastructure that buys those days. Congrats to the team! One question I'd love to ask: do AIDE2's gains transfer across solver models, or are they tuned to the one it was improved against? (Our data says updating ability is flat across models — transfer is where the real information is.)

Profilbild von Larry Panozzo
Larry Panozzovor 2 Monaten

@a66mike99 The RSI era definitely began in 2026.

Profilbild von arya 🏴
arya 🏴vor 2 Monaten

This is wild - incredible work!

Profilbild von jsd
jsdvor 2 Monaten

> The first experimental evidence of recursive self-improvement

Ähnliche Videos