Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Superintelligence will be built on Self Improvement. Today Hexo Labs, we’re excited to release ‘SIA’ - an open-source Self-Improving AI, to achieve any goal through recursive self improvement. While trying to solve a problem, SIA doesn't just improve it's abilities by updating it's harness, it updates it's own weights...

529,123 görüntüleme • 4 ay önce •via X (Twitter)

65 Yorum

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

SIA is true recursive self improvement. Every Agent is composed of two main components - Model Weights & Harness The harness-update school of research has a meta-agent rewrites the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while the model weights are held fixed. The test-time training school uses hand-written RL pipelines to update the model's own weights on task feedback while the harness is held fixed. These two silos operate in isolation. We introduce SIA - a self-improving loop in which an agent updates both the harness and the weights of a task-specific agent. Read the paper: Github Repo:

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

SIA is the pioneering technique addressing both harness and weights improvement in one loop.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

Self Improvement is all you need. In line with the bitter lesson, SIA outperforms specialised agents at diverse tasks by self improving. Yes - SIA is a general purpose agent that outperforms agents trained to specialise on legal work on LawBench, as well as coding agents on CUDA Kernal Optimisation and biology focused agents on Denoising RNA samples respectively.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

SIA competes only with itself. SIA was tasked to tackle one of the hardest tasks from MLE Bench: Google Brain's Ventilator Pressure Prediction problem, a medical time-series problem where the model must predict airway pressure during mechanical ventilation from control-signal inputs. SIA did not just beat other agents but repeatedly beat it’s own performance to self improve on the task.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

Karpathy is the bottleneck We benchmarked SIA against @karpathy's hand-crafted Autoresearch agent on a task that predicts final validation R² from early ML run signals, hyperparameters, configs, and agent strategies. SIA - a general purpose agent, self improved itself to outperform a specialised agent built by an elite researcher in his field. The human is the final bottleneck to superintelligence.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

Weights update is the real breakthrough in continual learning Frontier coding agents are strong, but they're frozen. Point Claude Code or Codex at a task and they can't keep getting better at it. You can see it across all three benchmarks: they sit near the baseline and barely reach the prior SOTA line. With harness-only updates, we land in roughly the same neighbourhood. The breakaway happens once you update the model's weights, far past everything else. Self-improvement isn't a smarter scaffold. It's letting the model actually learn. In a harness update, the meta-agent only teaches the task-specific agent software engineering i.e. better parsers, retry logic, tool dispatch, search procedure. It never touches the domain itself. On LawBench it can build a cleaner classification pipeline, but it can't make the model understand Chinese criminal law. Weight updates do exactly that: gradient pressure pushes the model into latent reasoning about the problem - disambiguating 191 charge categories, internalizing H100 kernel patterns, learning that imputed RNA counts must be non-negative integers. The harness shapes how the agent searches; the weights make it a domain expert.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

Read the paper: Github Repo:

Filecoin profil fotoğrafı
Filecoin4 ay önce

@hexoai A model that updates its own weights needs verifiable checkpoints, cryptographic proof of what version existed at what point. Without that, you can't audit what it learned or when. That's a storage problem as much as a safety one.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Fair

Filecoin profil fotoğrafı
Filecoin4 ay önce

@hexoai glad it resonated

Piyush Garg profil fotoğrafı
Piyush Garg4 ay önce

@hexoai Recursive compound improvement agent is actually an interesting idea.

Alok Bishoyi profil fotoğrafı
Alok Bishoyi4 ay önce

first, congrats @kunalbhatia91 & @hexoai team on the launch i went through the repo- the harness loop update loop seems to exist. but i cant find the weight update half in the code. does sia not have any specific guidace towards carrying out rollouts / SFT etc or it is picked by free-form LLM judgment ? cause there's no specific skills or prompts related to weight updates or any integrations to do so Thanks

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai We will be co-releasing it with a partner and requires coordination with multiple stakeholders. Humans are still involved 😄

Alok Bishoyi profil fotoğrafı
Alok Bishoyi4 ay önce

@hexoai sir, "weight updates" are in half the name but cool! looking forward to it and rooting for you folks

Rohan Paul profil fotoğrafı
Rohan Paul4 ay önce

@hexoai wow. these numbers are wild "The gains are 56.6% on LawBench, 91.9% runtime reduction on GPU kernels, and 502% on denoising over the initial baseline"

Yev Marusenko profil fotoğrafı
Yev Marusenko4 ay önce

@hexoai In my own work finding that the next growth, optimization, and scale always comes from improving a process or the layer. Self improvement is an untapped area now. Big findings! Congrats on the publication and launch! Will see if @karpathy has any thoughts on it.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai @karpathy Recursive self improvement is in line with the bitter lesson, it’s the next primitive

Wesley profil fotoğrafı
Wesley4 ay önce

@hexoai This is next-level 🔥 SIA’s recursive loop blending harness evolution + live weight updates feels like the missing piece for true agentic growth. Those 500%+ jumps on RNA denoising? Insane. Open-source self-improvers just won. Diving in now! 🚀

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Weights is the real breakthrough!

Krishna Mehra profil fotoğrafı
Krishna Mehra4 ay önce

@hexoai Amazing to see where Sia will take civilization!

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai To infinity ∞

Neo Kim profil fotoğrafı
Neo Kim4 ay önce

@hexoai this is pretty wild! we now have self improving agents. taking a closer read of the paper.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai The future is here

Dylan Knox profil fotoğrafı
Dylan Knox4 ay önce

@hexoai Recursive self-improvement is the unlock that changes everything.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai It’s what will take us to ∞

Nawi profil fotoğrafı
Nawi4 ay önce

@hexoai Huge moment for AI agents. Hexo Labs’ SIA just topped MLE-Bench, outperformed research agents, and improved beyond its own prior version. The self-improving era is starting.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Not just came out on top, but kept improving to beat itself again and again

Samuel Verboomen profil fotoğrafı
Samuel Verboomen4 ay önce

@hexoai 🚀 Super excited and proud to be part of this journey. Happy to connect and exchange on the topic 💬 Excited for what’s ahead ✨

MONICA THUKKARAM profil fotoğrafı
MONICA THUKKARAM4 ay önce

@hexoai Congratulations to the team @hexoai

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Thanks @MonicaThukkaram! Wonderful to be collaborating with you at @Stanford

MONICA THUKKARAM profil fotoğrafı
MONICA THUKKARAM4 ay önce

@hexoai @Stanford Looking forward to our collaboration. This is Massive!

Rathin Shah profil fotoğrafı
Rathin Shah4 ay önce

@hexoai Good stuff ser! 🫡 Is this the last company we ever need?

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai It will humanity’s last invention!

Rebecca profil fotoğrafı
Rebecca4 ay önce

@hexoai Incredible achievement! Congrats to the entire team!

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai The future is here

ZOYA ✪ profil fotoğrafı
ZOYA ✪4 ay önce

@hexoai This feels like an important shift in AI agents.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai It is. Recursive self improvement will be the ultimate lever of progress

ZOYA ✪ profil fotoğrafı
ZOYA ✪4 ay önce

@hexoai awesome

Tyler Wayne profil fotoğrafı
Tyler Wayne4 ay önce

@hexoai Open-sourcing this is the most important part of the announcement.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai It shouldn’t be gatekept

Dharmik Harinkhede profil fotoğrafı
Dharmik Harinkhede4 ay önce

@hexoai Self-improving AI is the next major leap. An AI that upgrades both its workflow and its own weights while solving problems is a fascinating direction for the future of intelligence. 🚀

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@Dharmikpawar31 @hexoai It is the ultimate lever

Csaba Kissi profil fotoğrafı
Csaba Kissi4 ay önce

Most agent frameworks are still basically static: plan, call tools, execute. SIA is interesting because it focuses on the feedback loop itself - the harness, model, and memory layer improving from previous runs. That’s a much more important direction than just making agents “do more steps.”

Aiswarya Sankar profil fotoğrafı
Aiswarya Sankar4 ay önce

@hexoai Huge!!! So exciting to see how far this has come congrats @kunalbhatia91 and @hexoai team

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Thanks @Aiswarya_Sankar

Linus ✦ Ekenstam profil fotoğrafı
Linus ✦ Ekenstam4 ay önce

@hexoai guys, what’s to most impressive external use-case you’ve seen during early testing?

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai 15x speedup in GPU Kernel Optimisation

Chanukya Patnaik profil fotoğrafı
Chanukya Patnaik4 ay önce

@hexoai Kudos Vignesh, Kunal and Hexo team

Farlane profil fotoğrafı
Farlane4 ay önce

@hexoai Self-improving agents feel like the next major leap for AI infrastructure. Curious to watch how SIA evolves from here.

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai SIA will be a living, breathing organism by itself

Julius Ritter profil fotoğrafı
Julius Ritter3 ay önce

@hexoai Congrats guys, this is so big!

Subodh Kolhe profil fotoğrafı
Subodh Kolhe4 ay önce

@hexoai good stuff Bhatia🫡

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Thanks!

Saumya Saxena profil fotoğrafı
Saumya Saxena4 ay önce

@hexoai Incredible work man! Excited to see it come alive

Supriya Sharma profil fotoğrafı
Supriya Sharma4 ay önce

@hexoai Congratulations @hexoai This is phenomenal

Abhinav Gupta profil fotoğrafı
Abhinav Gupta4 ay önce

@hexoai Machao bhai

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Thanks!

Rahil Bhansali profil fotoğrafı
Rahil Bhansali4 ay önce

@hexoai This is awesome! Excited to try it

Kunal Bhatia profil fotoğrafı
Kunal Bhatia4 ay önce

@hexoai Thanks @rahilbhansali!

Tamaz Gadaev profil fotoğrafı
Tamaz Gadaev4 ay önce

@hexoai great thread! congrats with the launch

usman profil fotoğrafı
usman4 ay önce

@BenHolfeld @hexoai excited to try this!

Jay profil fotoğrafı
Jay4 ay önce

@hexoai The future belongs to AI that learns how to get better while working. 💡

arc. profil fotoğrafı
arc.4 ay önce

@beckfastattiffs @hexoai congrats on the launch! love to see it

Lily yang profil fotoğrafı
Lily yang4 ay önce

@hexoai Wonderful

Harish Uthayakumar profil fotoğrafı
Harish Uthayakumar4 ay önce

@hexoai Let’s go bro!

Benzer Videolar

New Paper! Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents A longstanding goal of AI research has been the creation of AI that can learn indefinitely. One path toward that goal is an AI that improves itself by rewriting its own code, including any code responsible for learning. That idea, known as a Gödel Machine, proposed by Jürgen Schmidhuber over two decades ago, is a hypothetical self-improving AI. It optimally solves problems by recursively rewriting its own code when it can mathematically prove a better strategy, making it a key concept in meta-learning or “learning to learn.” While the theoretical Gödel Machine promised provably beneficial self-modifications, its realization relied on an impractical assumption: that the AI could mathematically prove that a proposed change in its own code would yield a net improvement before adopting it. Sakana AI, in collaboration with Jeff Clune’s lab at UBC, proposes something more feasible: a system that harnesses the principles of open-ended algorithms like Darwinian evolution to search for improvements that empirically improve performance. We call the result the Darwin Gödel Machine. DGMs leverage foundation models to propose code improvements, and use recent innovations in open-ended algorithms to search for a growing library of diverse, high-quality AI agents. Applied to practical tasks, we implemented Darwin Gödel Machine as a self-improving coding agent that rewrites its own code to improve performance on programming tasks. It creates various self-improvements, such as a patch validation step, better file viewing, enhanced editing tools, generating and ranking multiple solutions to choose the best one, and adding a history of what has been tried before (and why it failed) when making new changes (see the attached video). We believe that Darwin Gödel Machines represent a concrete step towards AI systems that can autonomously gather their own stepping stones to learn and innovate forever!

hardmaru

105,033 görüntüleme • 1 yıl önce