正在加载视频...

视频加载失败

The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)

1,823,208 次观看 • 2 个月前 •via X (Twitter)

57 条评论

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Our RSI system AIDE² has two autoresearch loops. An inner loop, just like a normal autoresearch agent, optimizing code against an eval. An outer loop, optimizing the inner-loop agent's harness code against the inner loop's average score across different benchmarks. (2/7)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

After 100 iterations, the outer loop discovered seven improvements over the baseline. Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

We test the discovered agents on held-out benchmarks the outer loop never saw. They generalize. They beat the agent we hand-tuned for two years, on all three. Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model. (4/7)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

We also see an emergent phenomenon where the outer loop pushes the inner-loop agent's reward hacking rate lower, with a combination of prompting and rule-based checks. This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking. (5/7)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

On our RSI ladder, AIDE² is Level 1. Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks. We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

More in the blog post: - a breakdown of the discovered algorithms - the rejected ideas AIDE² tried, covering a surprising share of the search literature - the dead code it shipped (7/7)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Very proud of the team, @DhruvSrikanth, @yuxiangwu_, @dexhunt3r, and @BingchenZhao, for shipping such an ambitious project spanning nearly a year with relatively few resources. Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!

Shiny Gen Wizard 的头像
Shiny Gen Wizard2 个月前

Have you tried auto researching the auto researching the auto research? And then auto researching that?

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

We didn't tried but was thinking about it. The problem cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Jeff Clune 的头像
Jeff Clune2 个月前

“The first experimental evidence of recursive self-improvement (RSI).” 🤔 What about the Darwin Gödel Machine, HyperAgents, and our work at Recursive on First Steps Toward Automated AI Research, among lots of other work?

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

All very good work of course, and they demonstrate the concept well. We cited them in our RSI ladder post, and we'll discuss them in more depth in the full tech report. For this claim, we measure RSI in an outcome-based way rather than by high level mechanism. The bar for Level 1 is that a frontier AI system consistently improves itself, at a speed materially better than human R&D using general AI tools. To hold that claim, we think a system needs to meet all of the following at the same time. - The starting point is a frontier AI system. - The comparison is against top experts. - The found solution generalizes to a wide set of held-out benchmarks. - The improvement is measured as efficiency, so the evolved agent is evaluated under the same compute budget as the baseline. As far as we know, no previous work reported evidence that meets all four at the same time. Of course, happy to hear your thoughts if you disagree.

Henry Dowling 的头像
Henry Dowling2 个月前

This is really cool! Dumb q but are you suspicious that it "overfit" on your performance benchmark relative to the harness that you hand-tuned for two years at all?

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Not a dumb question. We were also wondering whether that would be the case. That’s why we tested both agents on completely held-out benchmarks. We had never even run some of them before, such as building a physics-based weather forecasting model. The discovered agent still beats the hand-tuned one!

Henry Dowling 的头像
Henry Dowling2 个月前

oh wow, that is very convincing

Obedience Adara 的头像
Obedience Adara2 个月前

2 years of human engineering beaten in 8 days. That is an insane proof of concept for RSI. Incredible work btw ❤️

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Thx!

Theodore Galanos 的头像
Theodore Galanos2 个月前

Beautiful! I did some harness optimisation and evolution work earlier (map elites seems promising!). I love this meta-harness direction. Have written some about it here: with some long horizon work coming today. I wonder, did you try this in domains outside of code / ml engineering? Would love to try it in engineering (atoms not bits)

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Physical engineering isn’t on our roadmap but would be cool!

Michael B. Currie 的头像
Michael B. Currie2 个月前

I do not believe this because the visualizations are too slick. No one who is really doing frontier work has the time for that..

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Man, just prompt Fable to do it with remotion.. You should be able to do the same in an hour as long as you have the data

DEJAN 的头像
DEJAN2 个月前

@MichaelBCurrie I just tried this on my own setup which is very similar (I use bayesian inference for content optimization) and have now spent *4 hours* horsing around with the visuals not getting it the way I want. Still not there!!!

Michael B. Currie 的头像
Michael B. Currie2 个月前

@zhengyaojiang Yes exactly. Your attempt looks more like what a few hours of Fable effort would be, and as you say it is still not close to their quality of visualization.

Fred Jonsson 的头像
Fred Jonsson2 个月前

This is very cool - all you need now is the outer-outer loop!

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Yes! Time to "training" harness instead of manually design the logic

Burny - Effective Curiosity 的头像
Burny - Effective Curiosity2 个月前

Impressive! And now also self-improve the architecture and weights of the underlying model powering the agent harness.

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

It would be expensive to retrain a frontier model every iteration, but yeah I think that might be a good way to raise the ceiling of RSI

gpu go brr... 的头像
gpu go brr...2 个月前

But what if you auto research the auto research of the auto research?

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

We were actually thinking about it 😂 But the cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Pawel Pachniewski 的头像
Pawel Pachniewski2 个月前

@skynetislov3 Because what you eventually want is AGI... after the third level it doesn't exactly make sense. No doubt the "cost for evolution" also went up and...here we are.

gpu go brr... 的头像
gpu go brr...2 个月前

@zhengyaojiang 'mo levels 'mo bitches

Nisan Chhetri 的头像
Nisan Chhetri2 个月前

Very interesting stuff! I am excited to see what you guys do in the future.

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Thank you!

Oliver Wang 的头像
Oliver Wang2 个月前

Is it time to autoresearch the autoresearch of the autoresearch agent?

rohit 的头像
rohit2 个月前

This is very cool. You should try it with tevo (evolutionary loop), I did for AI architectures and it works pretty well, but I think it'd do even better for this!

Kevin 的头像
Kevin2 个月前

This isn't RSI. An agent doesn't just have to be able to improve itself. That improvement has to also make it better at improving itself so that the recursive improvement can continue in a loop. If the loop ends, that's provisional evidence against RSI.

Gurusha Juneja (✈️ ICML'26) 的头像
Gurusha Juneja (✈️ ICML'26)2 个月前

You should try doing this with ARTS too: Probably auto-researching the TTT part.

Theo 的头像
Theo2 个月前

have you considered integrating and or comparing this with flywheel by @paradigmainc ? I see a lot of parallels with the looping ideology and graph structure

420Trades 的头像
420Trades2 个月前

Such a beautiful graphic, the animator did a great job.

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

Thanks! I think the key is that it's grounded in actual data. Reality is beautiful

Alok Bishoyi 的头像
Alok Bishoyi2 个月前

Cool stuff!

Zhengyao Jiang 的头像
Zhengyao Jiang2 个月前

thx!

Ramez Naam 的头像
Ramez Naam2 个月前

This is super cool. Congrats!

Dev 的头像
Dev2 个月前

you should auto research the autoresearch for the autoresearch agent🫣

ismaelvega 的头像
ismaelvega2 个月前

uhmmm so what I understand with 'out of distribution' is that only you guys have access to those tasks/bench, right?

Kirk Patrick Miller 的头像
Kirk Patrick Miller2 个月前

@EObadoni imagine using our training for this. 😎 Ready for the next step. •

forloop 的头像
forloop2 个月前

can we have this but 3 dimensional

Tino Wening 的头像
Tino Wening2 个月前

I made something similar as autoresearch. My approach uses population-guided autoresearch adopted from evolutionary algorithm approaches.

La View Claire 的头像
La View Claire2 个月前

I am a paying customer to support the development.

tautologer 的头像
tautologer2 个月前

@wordgrammer what's your take here

redgreenblue 的头像
redgreenblue2 个月前

Does the agent access the agent harness concepts (skills, instructions, rules, hooks) or access the agent harness code itself?

chris 的头像
chris2 个月前

Wow, recursive self-improvement confirmed! its over bois

Morgan McGuire 的头像
Morgan McGuire2 个月前

the held-out/OOD benchmarks were great for validating the experiments' performance 👌

Shiv Kampani 的头像
Shiv Kampani2 个月前

awesome work but this gave me aides

Henry Lu 的头像
Henry Lu2 个月前

The number that matters most here isn't the benchmark delta — 𝐢𝐭'𝐬 𝐞𝐢𝐠𝐡𝐭 𝐝𝐚𝐲𝐬 𝐮𝐧𝐚𝐭𝐭𝐞𝐧𝐝𝐞𝐝. Deltas age; what moves this field's ceiling is how long a loop runs without a human touching it, and "layered defenses against reward hacking" is exactly the audit infrastructure that buys those days. Congrats to the team! One question I'd love to ask: do AIDE2's gains transfer across solver models, or are they tuned to the one it was improved against? (Our data says updating ability is flat across models — transfer is where the real information is.)

Larry Panozzo 的头像
Larry Panozzo2 个月前

@a66mike99 The RSI era definitely began in 2026.

arya 🏴 的头像
arya 🏴2 个月前

This is wild - incredible work!

jsd 的头像
jsd2 个月前

> The first experimental evidence of recursive self-improvement

相关视频