Video wird geladen...
Video konnte nicht geladen werden
Superintelligence will be built on Self Improvement. Today Hexo Labs, we’re excited to release ‘SIA’ - an open-source Self-Improving AI, to achieve any goal through recursive self improvement. While trying to solve a problem, SIA doesn't just improve it's abilities by updating it's harness, it updates it's own weights... show more
529,123 Aufrufe • vor 4 Monaten •via X (Twitter)
65 Kommentare

SIA is true recursive self improvement. Every Agent is composed of two main components - Model Weights & Harness The harness-update school of research has a meta-agent rewrites the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while the model weights are held fixed. The test-time training school uses hand-written RL pipelines to update the model's own weights on task feedback while the harness is held fixed. These two silos operate in isolation. We introduce SIA - a self-improving loop in which an agent updates both the harness and the weights of a task-specific agent. Read the paper: Github Repo:

SIA is the pioneering technique addressing both harness and weights improvement in one loop.

Self Improvement is all you need. In line with the bitter lesson, SIA outperforms specialised agents at diverse tasks by self improving. Yes - SIA is a general purpose agent that outperforms agents trained to specialise on legal work on LawBench, as well as coding agents on CUDA Kernal Optimisation and biology focused agents on Denoising RNA samples respectively.

SIA competes only with itself. SIA was tasked to tackle one of the hardest tasks from MLE Bench: Google Brain's Ventilator Pressure Prediction problem, a medical time-series problem where the model must predict airway pressure during mechanical ventilation from control-signal inputs. SIA did not just beat other agents but repeatedly beat it’s own performance to self improve on the task.

Karpathy is the bottleneck We benchmarked SIA against @karpathy's hand-crafted Autoresearch agent on a task that predicts final validation R² from early ML run signals, hyperparameters, configs, and agent strategies. SIA - a general purpose agent, self improved itself to outperform a specialised agent built by an elite researcher in his field. The human is the final bottleneck to superintelligence.

Weights update is the real breakthrough in continual learning Frontier coding agents are strong, but they're frozen. Point Claude Code or Codex at a task and they can't keep getting better at it. You can see it across all three benchmarks: they sit near the baseline and barely reach the prior SOTA line. With harness-only updates, we land in roughly the same neighbourhood. The breakaway happens once you update the model's weights, far past everything else. Self-improvement isn't a smarter scaffold. It's letting the model actually learn. In a harness update, the meta-agent only teaches the task-specific agent software engineering i.e. better parsers, retry logic, tool dispatch, search procedure. It never touches the domain itself. On LawBench it can build a cleaner classification pipeline, but it can't make the model understand Chinese criminal law. Weight updates do exactly that: gradient pressure pushes the model into latent reasoning about the problem - disambiguating 191 charge categories, internalizing H100 kernel patterns, learning that imputed RNA counts must be non-negative integers. The harness shapes how the agent searches; the weights make it a domain expert.

Read the paper: Github Repo:

@hexoai A model that updates its own weights needs verifiable checkpoints, cryptographic proof of what version existed at what point. Without that, you can't audit what it learned or when. That's a storage problem as much as a safety one.

@hexoai Fair

@hexoai glad it resonated

@hexoai Recursive compound improvement agent is actually an interesting idea.

first, congrats @kunalbhatia91 & @hexoai team on the launch i went through the repo- the harness loop update loop seems to exist. but i cant find the weight update half in the code. does sia not have any specific guidace towards carrying out rollouts / SFT etc or it is picked by free-form LLM judgment ? cause there's no specific skills or prompts related to weight updates or any integrations to do so Thanks

@hexoai We will be co-releasing it with a partner and requires coordination with multiple stakeholders. Humans are still involved 😄

@hexoai sir, "weight updates" are in half the name but cool! looking forward to it and rooting for you folks

@hexoai wow. these numbers are wild "The gains are 56.6% on LawBench, 91.9% runtime reduction on GPU kernels, and 502% on denoising over the initial baseline"

@hexoai In my own work finding that the next growth, optimization, and scale always comes from improving a process or the layer. Self improvement is an untapped area now. Big findings! Congrats on the publication and launch! Will see if @karpathy has any thoughts on it.

@hexoai @karpathy Recursive self improvement is in line with the bitter lesson, it’s the next primitive

@hexoai This is next-level 🔥 SIA’s recursive loop blending harness evolution + live weight updates feels like the missing piece for true agentic growth. Those 500%+ jumps on RNA denoising? Insane. Open-source self-improvers just won. Diving in now! 🚀

@hexoai Weights is the real breakthrough!

@hexoai Amazing to see where Sia will take civilization!

@hexoai To infinity ∞

@hexoai this is pretty wild! we now have self improving agents. taking a closer read of the paper.

@hexoai The future is here

@hexoai Recursive self-improvement is the unlock that changes everything.

@hexoai It’s what will take us to ∞

@hexoai Huge moment for AI agents. Hexo Labs’ SIA just topped MLE-Bench, outperformed research agents, and improved beyond its own prior version. The self-improving era is starting.

@hexoai Not just came out on top, but kept improving to beat itself again and again

@hexoai 🚀 Super excited and proud to be part of this journey. Happy to connect and exchange on the topic 💬 Excited for what’s ahead ✨

@hexoai Congratulations to the team @hexoai

@hexoai Thanks @MonicaThukkaram! Wonderful to be collaborating with you at @Stanford

@hexoai @Stanford Looking forward to our collaboration. This is Massive!

@hexoai Good stuff ser! 🫡 Is this the last company we ever need?

@hexoai It will humanity’s last invention!

@hexoai Incredible achievement! Congrats to the entire team!

@hexoai The future is here

@hexoai This feels like an important shift in AI agents.

@hexoai It is. Recursive self improvement will be the ultimate lever of progress

@hexoai awesome

@hexoai Open-sourcing this is the most important part of the announcement.

@hexoai It shouldn’t be gatekept

@hexoai Self-improving AI is the next major leap. An AI that upgrades both its workflow and its own weights while solving problems is a fascinating direction for the future of intelligence. 🚀

@Dharmikpawar31 @hexoai It is the ultimate lever

Most agent frameworks are still basically static: plan, call tools, execute. SIA is interesting because it focuses on the feedback loop itself - the harness, model, and memory layer improving from previous runs. That’s a much more important direction than just making agents “do more steps.”

@hexoai Huge!!! So exciting to see how far this has come congrats @kunalbhatia91 and @hexoai team

@hexoai Thanks @Aiswarya_Sankar

@hexoai guys, what’s to most impressive external use-case you’ve seen during early testing?

@hexoai 15x speedup in GPU Kernel Optimisation

@hexoai Kudos Vignesh, Kunal and Hexo team

@hexoai Self-improving agents feel like the next major leap for AI infrastructure. Curious to watch how SIA evolves from here.

@hexoai SIA will be a living, breathing organism by itself

@hexoai Congrats guys, this is so big!

@hexoai good stuff Bhatia🫡

@hexoai Thanks!

@hexoai Incredible work man! Excited to see it come alive

@hexoai Congratulations @hexoai This is phenomenal

@hexoai Machao bhai

@hexoai Thanks!

@hexoai This is awesome! Excited to try it

@hexoai Thanks @rahilbhansali!

@hexoai great thread! congrats with the launch

@BenHolfeld @hexoai excited to try this!

@hexoai The future belongs to AI that learns how to get better while working. 💡

@beckfastattiffs @hexoai congrats on the launch! love to see it

@hexoai Wonderful

@hexoai Let’s go bro!

