Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)

1,823,208 görüntüleme • 2 ay önce •via X (Twitter)

57 Yorum

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Our RSI system AIDE² has two autoresearch loops. An inner loop, just like a normal autoresearch agent, optimizing code against an eval. An outer loop, optimizing the inner-loop agent's harness code against the inner loop's average score across different benchmarks. (2/7)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

After 100 iterations, the outer loop discovered seven improvements over the baseline. Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

We test the discovered agents on held-out benchmarks the outer loop never saw. They generalize. They beat the agent we hand-tuned for two years, on all three. Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model. (4/7)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

We also see an emergent phenomenon where the outer loop pushes the inner-loop agent's reward hacking rate lower, with a combination of prompting and rule-based checks. This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking. (5/7)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

On our RSI ladder, AIDE² is Level 1. Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks. We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

More in the blog post: - a breakdown of the discovered algorithms - the rejected ideas AIDE² tried, covering a surprising share of the search literature - the dead code it shipped (7/7)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Very proud of the team, @DhruvSrikanth, @yuxiangwu_, @dexhunt3r, and @BingchenZhao, for shipping such an ambitious project spanning nearly a year with relatively few resources. Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!

Shiny Gen Wizard profil fotoğrafı
Shiny Gen Wizard2 ay önce

Have you tried auto researching the auto researching the auto research? And then auto researching that?

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

We didn't tried but was thinking about it. The problem cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Jeff Clune profil fotoğrafı
Jeff Clune2 ay önce

“The first experimental evidence of recursive self-improvement (RSI).” 🤔 What about the Darwin Gödel Machine, HyperAgents, and our work at Recursive on First Steps Toward Automated AI Research, among lots of other work?

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

All very good work of course, and they demonstrate the concept well. We cited them in our RSI ladder post, and we'll discuss them in more depth in the full tech report. For this claim, we measure RSI in an outcome-based way rather than by high level mechanism. The bar for Level 1 is that a frontier AI system consistently improves itself, at a speed materially better than human R&D using general AI tools. To hold that claim, we think a system needs to meet all of the following at the same time. - The starting point is a frontier AI system. - The comparison is against top experts. - The found solution generalizes to a wide set of held-out benchmarks. - The improvement is measured as efficiency, so the evolved agent is evaluated under the same compute budget as the baseline. As far as we know, no previous work reported evidence that meets all four at the same time. Of course, happy to hear your thoughts if you disagree.

Henry Dowling profil fotoğrafı
Henry Dowling2 ay önce

This is really cool! Dumb q but are you suspicious that it "overfit" on your performance benchmark relative to the harness that you hand-tuned for two years at all?

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Not a dumb question. We were also wondering whether that would be the case. That’s why we tested both agents on completely held-out benchmarks. We had never even run some of them before, such as building a physics-based weather forecasting model. The discovered agent still beats the hand-tuned one!

Henry Dowling profil fotoğrafı
Henry Dowling2 ay önce

oh wow, that is very convincing

Obedience Adara profil fotoğrafı
Obedience Adara2 ay önce

2 years of human engineering beaten in 8 days. That is an insane proof of concept for RSI. Incredible work btw ❤️

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Thx!

Theodore Galanos profil fotoğrafı
Theodore Galanos2 ay önce

Beautiful! I did some harness optimisation and evolution work earlier (map elites seems promising!). I love this meta-harness direction. Have written some about it here: with some long horizon work coming today. I wonder, did you try this in domains outside of code / ml engineering? Would love to try it in engineering (atoms not bits)

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Physical engineering isn’t on our roadmap but would be cool!

Michael B. Currie profil fotoğrafı
Michael B. Currie2 ay önce

I do not believe this because the visualizations are too slick. No one who is really doing frontier work has the time for that..

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Man, just prompt Fable to do it with remotion.. You should be able to do the same in an hour as long as you have the data

DEJAN profil fotoğrafı
DEJAN2 ay önce

@MichaelBCurrie I just tried this on my own setup which is very similar (I use bayesian inference for content optimization) and have now spent *4 hours* horsing around with the visuals not getting it the way I want. Still not there!!!

Michael B. Currie profil fotoğrafı
Michael B. Currie2 ay önce

@zhengyaojiang Yes exactly. Your attempt looks more like what a few hours of Fable effort would be, and as you say it is still not close to their quality of visualization.

Fred Jonsson profil fotoğrafı
Fred Jonsson2 ay önce

This is very cool - all you need now is the outer-outer loop!

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Yes! Time to "training" harness instead of manually design the logic

Burny - Effective Curiosity profil fotoğrafı
Burny - Effective Curiosity2 ay önce

Impressive! And now also self-improve the architecture and weights of the underlying model powering the agent harness.

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

It would be expensive to retrain a frontier model every iteration, but yeah I think that might be a good way to raise the ceiling of RSI

gpu go brr... profil fotoğrafı
gpu go brr...2 ay önce

But what if you auto research the auto research of the auto research?

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

We were actually thinking about it 😂 But the cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Pawel Pachniewski profil fotoğrafı
Pawel Pachniewski2 ay önce

@skynetislov3 Because what you eventually want is AGI... after the third level it doesn't exactly make sense. No doubt the "cost for evolution" also went up and...here we are.

gpu go brr... profil fotoğrafı
gpu go brr...2 ay önce

@zhengyaojiang 'mo levels 'mo bitches

Nisan Chhetri profil fotoğrafı
Nisan Chhetri2 ay önce

Very interesting stuff! I am excited to see what you guys do in the future.

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Thank you!

Oliver Wang profil fotoğrafı
Oliver Wang2 ay önce

Is it time to autoresearch the autoresearch of the autoresearch agent?

rohit profil fotoğrafı
rohit2 ay önce

This is very cool. You should try it with tevo (evolutionary loop), I did for AI architectures and it works pretty well, but I think it'd do even better for this!

Kevin profil fotoğrafı
Kevin2 ay önce

This isn't RSI. An agent doesn't just have to be able to improve itself. That improvement has to also make it better at improving itself so that the recursive improvement can continue in a loop. If the loop ends, that's provisional evidence against RSI.

Gurusha Juneja (✈️ ICML'26) profil fotoğrafı
Gurusha Juneja (✈️ ICML'26)2 ay önce

You should try doing this with ARTS too: Probably auto-researching the TTT part.

Theo profil fotoğrafı
Theo2 ay önce

have you considered integrating and or comparing this with flywheel by @paradigmainc ? I see a lot of parallels with the looping ideology and graph structure

420Trades profil fotoğrafı
420Trades2 ay önce

Such a beautiful graphic, the animator did a great job.

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

Thanks! I think the key is that it's grounded in actual data. Reality is beautiful

Alok Bishoyi profil fotoğrafı
Alok Bishoyi2 ay önce

Cool stuff!

Zhengyao Jiang profil fotoğrafı
Zhengyao Jiang2 ay önce

thx!

Ramez Naam profil fotoğrafı
Ramez Naam2 ay önce

This is super cool. Congrats!

Dev profil fotoğrafı
Dev2 ay önce

you should auto research the autoresearch for the autoresearch agent🫣

ismaelvega profil fotoğrafı
ismaelvega2 ay önce

uhmmm so what I understand with 'out of distribution' is that only you guys have access to those tasks/bench, right?

Kirk Patrick Miller profil fotoğrafı
Kirk Patrick Miller2 ay önce

@EObadoni imagine using our training for this. 😎 Ready for the next step. •

forloop profil fotoğrafı
forloop2 ay önce

can we have this but 3 dimensional

Tino Wening profil fotoğrafı
Tino Wening2 ay önce

I made something similar as autoresearch. My approach uses population-guided autoresearch adopted from evolutionary algorithm approaches.

La View Claire profil fotoğrafı
La View Claire2 ay önce

I am a paying customer to support the development.

tautologer profil fotoğrafı
tautologer2 ay önce

@wordgrammer what's your take here

redgreenblue profil fotoğrafı
redgreenblue2 ay önce

Does the agent access the agent harness concepts (skills, instructions, rules, hooks) or access the agent harness code itself?

chris profil fotoğrafı
chris2 ay önce

Wow, recursive self-improvement confirmed! its over bois

Morgan McGuire profil fotoğrafı
Morgan McGuire2 ay önce

the held-out/OOD benchmarks were great for validating the experiments' performance 👌

Shiv Kampani profil fotoğrafı
Shiv Kampani2 ay önce

awesome work but this gave me aides

Henry Lu profil fotoğrafı
Henry Lu2 ay önce

The number that matters most here isn't the benchmark delta — 𝐢𝐭'𝐬 𝐞𝐢𝐠𝐡𝐭 𝐝𝐚𝐲𝐬 𝐮𝐧𝐚𝐭𝐭𝐞𝐧𝐝𝐞𝐝. Deltas age; what moves this field's ceiling is how long a loop runs without a human touching it, and "layered defenses against reward hacking" is exactly the audit infrastructure that buys those days. Congrats to the team! One question I'd love to ask: do AIDE2's gains transfer across solver models, or are they tuned to the one it was improved against? (Our data says updating ability is flat across models — transfer is where the real information is.)

Larry Panozzo profil fotoğrafı
Larry Panozzo2 ay önce

@a66mike99 The RSI era definitely began in 2026.

arya 🏴 profil fotoğrafı
arya 🏴2 ay önce

This is wild - incredible work!

jsd profil fotoğrafı
jsd2 ay önce

> The first experimental evidence of recursive self-improvement

Benzer Videolar