Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Superintelligence will be built on Self Improvement. Today Hexo Labs, we’re excited to release ‘SIA’ - an open-source Self-Improving AI, to achieve any goal through recursive self improvement. While trying to solve a problem, SIA doesn't just improve it's abilities by updating it's harness, it updates it's own weights...

529,123 Aufrufe • vor 4 Monaten •via X (Twitter)

65 Kommentare

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

SIA is true recursive self improvement. Every Agent is composed of two main components - Model Weights & Harness The harness-update school of research has a meta-agent rewrites the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while the model weights are held fixed. The test-time training school uses hand-written RL pipelines to update the model's own weights on task feedback while the harness is held fixed. These two silos operate in isolation. We introduce SIA - a self-improving loop in which an agent updates both the harness and the weights of a task-specific agent. Read the paper: Github Repo:

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

SIA is the pioneering technique addressing both harness and weights improvement in one loop.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

Self Improvement is all you need. In line with the bitter lesson, SIA outperforms specialised agents at diverse tasks by self improving. Yes - SIA is a general purpose agent that outperforms agents trained to specialise on legal work on LawBench, as well as coding agents on CUDA Kernal Optimisation and biology focused agents on Denoising RNA samples respectively.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

SIA competes only with itself. SIA was tasked to tackle one of the hardest tasks from MLE Bench: Google Brain's Ventilator Pressure Prediction problem, a medical time-series problem where the model must predict airway pressure during mechanical ventilation from control-signal inputs. SIA did not just beat other agents but repeatedly beat it’s own performance to self improve on the task.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

Karpathy is the bottleneck We benchmarked SIA against @karpathy's hand-crafted Autoresearch agent on a task that predicts final validation R² from early ML run signals, hyperparameters, configs, and agent strategies. SIA - a general purpose agent, self improved itself to outperform a specialised agent built by an elite researcher in his field. The human is the final bottleneck to superintelligence.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

Weights update is the real breakthrough in continual learning Frontier coding agents are strong, but they're frozen. Point Claude Code or Codex at a task and they can't keep getting better at it. You can see it across all three benchmarks: they sit near the baseline and barely reach the prior SOTA line. With harness-only updates, we land in roughly the same neighbourhood. The breakaway happens once you update the model's weights, far past everything else. Self-improvement isn't a smarter scaffold. It's letting the model actually learn. In a harness update, the meta-agent only teaches the task-specific agent software engineering i.e. better parsers, retry logic, tool dispatch, search procedure. It never touches the domain itself. On LawBench it can build a cleaner classification pipeline, but it can't make the model understand Chinese criminal law. Weight updates do exactly that: gradient pressure pushes the model into latent reasoning about the problem - disambiguating 191 charge categories, internalizing H100 kernel patterns, learning that imputed RNA counts must be non-negative integers. The harness shapes how the agent searches; the weights make it a domain expert.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

Read the paper: Github Repo:

Profilbild von Filecoin
Filecoinvor 4 Monaten

@hexoai A model that updates its own weights needs verifiable checkpoints, cryptographic proof of what version existed at what point. Without that, you can't audit what it learned or when. That's a storage problem as much as a safety one.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Fair

Profilbild von Filecoin
Filecoinvor 4 Monaten

@hexoai glad it resonated

Profilbild von Piyush Garg
Piyush Gargvor 4 Monaten

@hexoai Recursive compound improvement agent is actually an interesting idea.

Profilbild von Alok Bishoyi
Alok Bishoyivor 4 Monaten

first, congrats @kunalbhatia91 & @hexoai team on the launch i went through the repo- the harness loop update loop seems to exist. but i cant find the weight update half in the code. does sia not have any specific guidace towards carrying out rollouts / SFT etc or it is picked by free-form LLM judgment ? cause there's no specific skills or prompts related to weight updates or any integrations to do so Thanks

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai We will be co-releasing it with a partner and requires coordination with multiple stakeholders. Humans are still involved 😄

Profilbild von Alok Bishoyi
Alok Bishoyivor 4 Monaten

@hexoai sir, "weight updates" are in half the name but cool! looking forward to it and rooting for you folks

Profilbild von Rohan Paul
Rohan Paulvor 4 Monaten

@hexoai wow. these numbers are wild "The gains are 56.6% on LawBench, 91.9% runtime reduction on GPU kernels, and 502% on denoising over the initial baseline"

Profilbild von Yev Marusenko
Yev Marusenkovor 4 Monaten

@hexoai In my own work finding that the next growth, optimization, and scale always comes from improving a process or the layer. Self improvement is an untapped area now. Big findings! Congrats on the publication and launch! Will see if @karpathy has any thoughts on it.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai @karpathy Recursive self improvement is in line with the bitter lesson, it’s the next primitive

Profilbild von Wesley
Wesleyvor 4 Monaten

@hexoai This is next-level 🔥 SIA’s recursive loop blending harness evolution + live weight updates feels like the missing piece for true agentic growth. Those 500%+ jumps on RNA denoising? Insane. Open-source self-improvers just won. Diving in now! 🚀

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Weights is the real breakthrough!

Profilbild von Krishna Mehra
Krishna Mehravor 4 Monaten

@hexoai Amazing to see where Sia will take civilization!

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai To infinity ∞

Profilbild von Neo Kim
Neo Kimvor 4 Monaten

@hexoai this is pretty wild! we now have self improving agents. taking a closer read of the paper.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai The future is here

Profilbild von Dylan Knox
Dylan Knoxvor 4 Monaten

@hexoai Recursive self-improvement is the unlock that changes everything.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai It’s what will take us to ∞

Profilbild von Nawi
Nawivor 4 Monaten

@hexoai Huge moment for AI agents. Hexo Labs’ SIA just topped MLE-Bench, outperformed research agents, and improved beyond its own prior version. The self-improving era is starting.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Not just came out on top, but kept improving to beat itself again and again

Profilbild von Samuel Verboomen
Samuel Verboomenvor 4 Monaten

@hexoai 🚀 Super excited and proud to be part of this journey. Happy to connect and exchange on the topic 💬 Excited for what’s ahead ✨

Profilbild von MONICA THUKKARAM
MONICA THUKKARAMvor 4 Monaten

@hexoai Congratulations to the team @hexoai

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Thanks @MonicaThukkaram! Wonderful to be collaborating with you at @Stanford

Profilbild von MONICA THUKKARAM
MONICA THUKKARAMvor 4 Monaten

@hexoai @Stanford Looking forward to our collaboration. This is Massive!

Profilbild von Rathin Shah
Rathin Shahvor 4 Monaten

@hexoai Good stuff ser! 🫡 Is this the last company we ever need?

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai It will humanity’s last invention!

Profilbild von Rebecca
Rebeccavor 4 Monaten

@hexoai Incredible achievement! Congrats to the entire team!

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai The future is here

Profilbild von ZOYA ✪
ZOYA ✪vor 4 Monaten

@hexoai This feels like an important shift in AI agents.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai It is. Recursive self improvement will be the ultimate lever of progress

Profilbild von ZOYA ✪
ZOYA ✪vor 4 Monaten

@hexoai awesome

Profilbild von Tyler Wayne
Tyler Waynevor 4 Monaten

@hexoai Open-sourcing this is the most important part of the announcement.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai It shouldn’t be gatekept

Profilbild von Dharmik Harinkhede
Dharmik Harinkhedevor 4 Monaten

@hexoai Self-improving AI is the next major leap. An AI that upgrades both its workflow and its own weights while solving problems is a fascinating direction for the future of intelligence. 🚀

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@Dharmikpawar31 @hexoai It is the ultimate lever

Profilbild von Csaba Kissi
Csaba Kissivor 4 Monaten

Most agent frameworks are still basically static: plan, call tools, execute. SIA is interesting because it focuses on the feedback loop itself - the harness, model, and memory layer improving from previous runs. That’s a much more important direction than just making agents “do more steps.”

Profilbild von Aiswarya Sankar
Aiswarya Sankarvor 4 Monaten

@hexoai Huge!!! So exciting to see how far this has come congrats @kunalbhatia91 and @hexoai team

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Thanks @Aiswarya_Sankar

Profilbild von Linus ✦ Ekenstam
Linus ✦ Ekenstamvor 4 Monaten

@hexoai guys, what’s to most impressive external use-case you’ve seen during early testing?

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai 15x speedup in GPU Kernel Optimisation

Profilbild von Chanukya Patnaik
Chanukya Patnaikvor 4 Monaten

@hexoai Kudos Vignesh, Kunal and Hexo team

Profilbild von Farlane
Farlanevor 4 Monaten

@hexoai Self-improving agents feel like the next major leap for AI infrastructure. Curious to watch how SIA evolves from here.

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai SIA will be a living, breathing organism by itself

Profilbild von Julius Ritter
Julius Rittervor 3 Monaten

@hexoai Congrats guys, this is so big!

Profilbild von Subodh Kolhe
Subodh Kolhevor 4 Monaten

@hexoai good stuff Bhatia🫡

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Thanks!

Profilbild von Saumya Saxena
Saumya Saxenavor 4 Monaten

@hexoai Incredible work man! Excited to see it come alive

Profilbild von Supriya Sharma
Supriya Sharmavor 4 Monaten

@hexoai Congratulations @hexoai This is phenomenal

Profilbild von Abhinav Gupta
Abhinav Guptavor 4 Monaten

@hexoai Machao bhai

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Thanks!

Profilbild von Rahil Bhansali
Rahil Bhansalivor 4 Monaten

@hexoai This is awesome! Excited to try it

Profilbild von Kunal Bhatia
Kunal Bhatiavor 4 Monaten

@hexoai Thanks @rahilbhansali!

Profilbild von Tamaz Gadaev
Tamaz Gadaevvor 4 Monaten

@hexoai great thread! congrats with the launch

Profilbild von usman
usmanvor 4 Monaten

@BenHolfeld @hexoai excited to try this!

Profilbild von Jay
Jayvor 4 Monaten

@hexoai The future belongs to AI that learns how to get better while working. 💡

Profilbild von arc.
arc.vor 4 Monaten

@beckfastattiffs @hexoai congrats on the launch! love to see it

Profilbild von Lily yang
Lily yangvor 4 Monaten

@hexoai Wonderful

Profilbild von Harish Uthayakumar
Harish Uthayakumarvor 4 Monaten

@hexoai Let’s go bro!

Ähnliche Videos

New Paper! Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents A longstanding goal of AI research has been the creation of AI that can learn indefinitely. One path toward that goal is an AI that improves itself by rewriting its own code, including any code responsible for learning. That idea, known as a Gödel Machine, proposed by Jürgen Schmidhuber over two decades ago, is a hypothetical self-improving AI. It optimally solves problems by recursively rewriting its own code when it can mathematically prove a better strategy, making it a key concept in meta-learning or “learning to learn.” While the theoretical Gödel Machine promised provably beneficial self-modifications, its realization relied on an impractical assumption: that the AI could mathematically prove that a proposed change in its own code would yield a net improvement before adopting it. Sakana AI, in collaboration with Jeff Clune’s lab at UBC, proposes something more feasible: a system that harnesses the principles of open-ended algorithms like Darwinian evolution to search for improvements that empirically improve performance. We call the result the Darwin Gödel Machine. DGMs leverage foundation models to propose code improvements, and use recent innovations in open-ended algorithms to search for a growing library of diverse, high-quality AI agents. Applied to practical tasks, we implemented Darwin Gödel Machine as a self-improving coding agent that rewrites its own code to improve performance on programming tasks. It creates various self-improvements, such as a patch validation step, better file viewing, enhanced editing tools, generating and ranking multiple solutions to choose the best one, and adding a history of what has been tried before (and why it failed) when making new changes (see the attached video). We believe that Darwin Gödel Machines represent a concrete step towards AI systems that can autonomously gather their own stepping stones to learn and innovate forever!

hardmaru

105,033 Aufrufe • vor 1 Jahr