Загрузка видео...

Не удалось загрузить видео

На главную

The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)

1,823,208 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 57

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Our RSI system AIDE² has two autoresearch loops. An inner loop, just like a normal autoresearch agent, optimizing code against an eval. An outer loop, optimizing the inner-loop agent's harness code against the inner loop's average score across different benchmarks. (2/7)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

After 100 iterations, the outer loop discovered seven improvements over the baseline. Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

We test the discovered agents on held-out benchmarks the outer loop never saw. They generalize. They beat the agent we hand-tuned for two years, on all three. Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model. (4/7)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

We also see an emergent phenomenon where the outer loop pushes the inner-loop agent's reward hacking rate lower, with a combination of prompting and rule-based checks. This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking. (5/7)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

On our RSI ladder, AIDE² is Level 1. Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks. We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

More in the blog post: - a breakdown of the discovered algorithms - the rejected ideas AIDE² tried, covering a surprising share of the search literature - the dead code it shipped (7/7)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Very proud of the team, @DhruvSrikanth, @yuxiangwu_, @dexhunt3r, and @BingchenZhao, for shipping such an ambitious project spanning nearly a year with relatively few resources. Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!

Фото профиля Shiny Gen Wizard
Shiny Gen Wizard2 месяцев назад

Have you tried auto researching the auto researching the auto research? And then auto researching that?

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

We didn't tried but was thinking about it. The problem cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Фото профиля Jeff Clune
Jeff Clune2 месяцев назад

“The first experimental evidence of recursive self-improvement (RSI).” 🤔 What about the Darwin Gödel Machine, HyperAgents, and our work at Recursive on First Steps Toward Automated AI Research, among lots of other work?

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

All very good work of course, and they demonstrate the concept well. We cited them in our RSI ladder post, and we'll discuss them in more depth in the full tech report. For this claim, we measure RSI in an outcome-based way rather than by high level mechanism. The bar for Level 1 is that a frontier AI system consistently improves itself, at a speed materially better than human R&D using general AI tools. To hold that claim, we think a system needs to meet all of the following at the same time. - The starting point is a frontier AI system. - The comparison is against top experts. - The found solution generalizes to a wide set of held-out benchmarks. - The improvement is measured as efficiency, so the evolved agent is evaluated under the same compute budget as the baseline. As far as we know, no previous work reported evidence that meets all four at the same time. Of course, happy to hear your thoughts if you disagree.

Фото профиля Henry Dowling
Henry Dowling2 месяцев назад

This is really cool! Dumb q but are you suspicious that it "overfit" on your performance benchmark relative to the harness that you hand-tuned for two years at all?

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Not a dumb question. We were also wondering whether that would be the case. That’s why we tested both agents on completely held-out benchmarks. We had never even run some of them before, such as building a physics-based weather forecasting model. The discovered agent still beats the hand-tuned one!

Фото профиля Henry Dowling
Henry Dowling2 месяцев назад

oh wow, that is very convincing

Фото профиля Obedience Adara
Obedience Adara2 месяцев назад

2 years of human engineering beaten in 8 days. That is an insane proof of concept for RSI. Incredible work btw ❤️

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Thx!

Фото профиля Theodore Galanos
Theodore Galanos2 месяцев назад

Beautiful! I did some harness optimisation and evolution work earlier (map elites seems promising!). I love this meta-harness direction. Have written some about it here: with some long horizon work coming today. I wonder, did you try this in domains outside of code / ml engineering? Would love to try it in engineering (atoms not bits)

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Physical engineering isn’t on our roadmap but would be cool!

Фото профиля Michael B. Currie
Michael B. Currie2 месяцев назад

I do not believe this because the visualizations are too slick. No one who is really doing frontier work has the time for that..

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Man, just prompt Fable to do it with remotion.. You should be able to do the same in an hour as long as you have the data

Фото профиля DEJAN
DEJAN2 месяцев назад

@MichaelBCurrie I just tried this on my own setup which is very similar (I use bayesian inference for content optimization) and have now spent *4 hours* horsing around with the visuals not getting it the way I want. Still not there!!!

Фото профиля Michael B. Currie
Michael B. Currie2 месяцев назад

@zhengyaojiang Yes exactly. Your attempt looks more like what a few hours of Fable effort would be, and as you say it is still not close to their quality of visualization.

Фото профиля Fred Jonsson
Fred Jonsson2 месяцев назад

This is very cool - all you need now is the outer-outer loop!

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Yes! Time to "training" harness instead of manually design the logic

Фото профиля Burny - Effective Curiosity
Burny - Effective Curiosity2 месяцев назад

Impressive! And now also self-improve the architecture and weights of the underlying model powering the agent harness.

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

It would be expensive to retrain a frontier model every iteration, but yeah I think that might be a good way to raise the ceiling of RSI

Фото профиля gpu go brr...
gpu go brr...2 месяцев назад

But what if you auto research the auto research of the auto research?

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

We were actually thinking about it 😂 But the cost increases exponentially with each order you go up. I think the more realistic path is to make the inner-loop improvements generalize to the outer loop, so you effectively get infinite-order improvement by regularly syncing the inner and outer loops.

Фото профиля Pawel Pachniewski
Pawel Pachniewski2 месяцев назад

@skynetislov3 Because what you eventually want is AGI... after the third level it doesn't exactly make sense. No doubt the "cost for evolution" also went up and...here we are.

Фото профиля gpu go brr...
gpu go brr...2 месяцев назад

@zhengyaojiang 'mo levels 'mo bitches

Фото профиля Nisan Chhetri
Nisan Chhetri2 месяцев назад

Very interesting stuff! I am excited to see what you guys do in the future.

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Thank you!

Фото профиля Oliver Wang
Oliver Wang2 месяцев назад

Is it time to autoresearch the autoresearch of the autoresearch agent?

Фото профиля rohit
rohit2 месяцев назад

This is very cool. You should try it with tevo (evolutionary loop), I did for AI architectures and it works pretty well, but I think it'd do even better for this!

Фото профиля Kevin
Kevin2 месяцев назад

This isn't RSI. An agent doesn't just have to be able to improve itself. That improvement has to also make it better at improving itself so that the recursive improvement can continue in a loop. If the loop ends, that's provisional evidence against RSI.

Фото профиля Gurusha Juneja (✈️ ICML'26)
Gurusha Juneja (✈️ ICML'26)2 месяцев назад

You should try doing this with ARTS too: Probably auto-researching the TTT part.

Фото профиля Theo
Theo2 месяцев назад

have you considered integrating and or comparing this with flywheel by @paradigmainc ? I see a lot of parallels with the looping ideology and graph structure

Фото профиля 420Trades
420Trades2 месяцев назад

Such a beautiful graphic, the animator did a great job.

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

Thanks! I think the key is that it's grounded in actual data. Reality is beautiful

Фото профиля Alok Bishoyi
Alok Bishoyi2 месяцев назад

Cool stuff!

Фото профиля Zhengyao Jiang
Zhengyao Jiang2 месяцев назад

thx!

Фото профиля Ramez Naam
Ramez Naam2 месяцев назад

This is super cool. Congrats!

Фото профиля Dev
Dev2 месяцев назад

you should auto research the autoresearch for the autoresearch agent🫣

Фото профиля ismaelvega
ismaelvega2 месяцев назад

uhmmm so what I understand with 'out of distribution' is that only you guys have access to those tasks/bench, right?

Фото профиля Kirk Patrick Miller
Kirk Patrick Miller2 месяцев назад

@EObadoni imagine using our training for this. 😎 Ready for the next step. •

Фото профиля forloop
forloop2 месяцев назад

can we have this but 3 dimensional

Фото профиля Tino Wening
Tino Wening2 месяцев назад

I made something similar as autoresearch. My approach uses population-guided autoresearch adopted from evolutionary algorithm approaches.

Фото профиля La View Claire
La View Claire2 месяцев назад

I am a paying customer to support the development.

Фото профиля tautologer
tautologer2 месяцев назад

@wordgrammer what's your take here

Фото профиля redgreenblue
redgreenblue2 месяцев назад

Does the agent access the agent harness concepts (skills, instructions, rules, hooks) or access the agent harness code itself?

Фото профиля chris
chris2 месяцев назад

Wow, recursive self-improvement confirmed! its over bois

Фото профиля Morgan McGuire
Morgan McGuire2 месяцев назад

the held-out/OOD benchmarks were great for validating the experiments' performance 👌

Фото профиля Shiv Kampani
Shiv Kampani2 месяцев назад

awesome work but this gave me aides

Фото профиля Henry Lu
Henry Lu2 месяцев назад

The number that matters most here isn't the benchmark delta — 𝐢𝐭'𝐬 𝐞𝐢𝐠𝐡𝐭 𝐝𝐚𝐲𝐬 𝐮𝐧𝐚𝐭𝐭𝐞𝐧𝐝𝐞𝐝. Deltas age; what moves this field's ceiling is how long a loop runs without a human touching it, and "layered defenses against reward hacking" is exactly the audit infrastructure that buys those days. Congrats to the team! One question I'd love to ask: do AIDE2's gains transfer across solver models, or are they tuned to the one it was improved against? (Our data says updating ability is flat across models — transfer is where the real information is.)

Фото профиля Larry Panozzo
Larry Panozzo2 месяцев назад

@a66mike99 The RSI era definitely began in 2026.

Фото профиля arya 🏴
arya 🏴2 месяцев назад

This is wild - incredible work!

Фото профиля jsd
jsd2 месяцев назад

> The first experimental evidence of recursive self-improvement

Похожие видео