Загрузка видео...

Не удалось загрузить видео

На главную

Superintelligence will be built on Self Improvement. Today Hexo Labs, we’re excited to release ‘SIA’ - an open-source Self-Improving AI, to achieve any goal through recursive self improvement. While trying to solve a problem, SIA doesn't just improve it's abilities by updating it's harness, it updates it's own weights...

529,123 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 65

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

SIA is true recursive self improvement. Every Agent is composed of two main components - Model Weights & Harness The harness-update school of research has a meta-agent rewrites the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while the model weights are held fixed. The test-time training school uses hand-written RL pipelines to update the model's own weights on task feedback while the harness is held fixed. These two silos operate in isolation. We introduce SIA - a self-improving loop in which an agent updates both the harness and the weights of a task-specific agent. Read the paper: Github Repo:

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

SIA is the pioneering technique addressing both harness and weights improvement in one loop.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

Self Improvement is all you need. In line with the bitter lesson, SIA outperforms specialised agents at diverse tasks by self improving. Yes - SIA is a general purpose agent that outperforms agents trained to specialise on legal work on LawBench, as well as coding agents on CUDA Kernal Optimisation and biology focused agents on Denoising RNA samples respectively.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

SIA competes only with itself. SIA was tasked to tackle one of the hardest tasks from MLE Bench: Google Brain's Ventilator Pressure Prediction problem, a medical time-series problem where the model must predict airway pressure during mechanical ventilation from control-signal inputs. SIA did not just beat other agents but repeatedly beat it’s own performance to self improve on the task.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

Karpathy is the bottleneck We benchmarked SIA against @karpathy's hand-crafted Autoresearch agent on a task that predicts final validation R² from early ML run signals, hyperparameters, configs, and agent strategies. SIA - a general purpose agent, self improved itself to outperform a specialised agent built by an elite researcher in his field. The human is the final bottleneck to superintelligence.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

Weights update is the real breakthrough in continual learning Frontier coding agents are strong, but they're frozen. Point Claude Code or Codex at a task and they can't keep getting better at it. You can see it across all three benchmarks: they sit near the baseline and barely reach the prior SOTA line. With harness-only updates, we land in roughly the same neighbourhood. The breakaway happens once you update the model's weights, far past everything else. Self-improvement isn't a smarter scaffold. It's letting the model actually learn. In a harness update, the meta-agent only teaches the task-specific agent software engineering i.e. better parsers, retry logic, tool dispatch, search procedure. It never touches the domain itself. On LawBench it can build a cleaner classification pipeline, but it can't make the model understand Chinese criminal law. Weight updates do exactly that: gradient pressure pushes the model into latent reasoning about the problem - disambiguating 191 charge categories, internalizing H100 kernel patterns, learning that imputed RNA counts must be non-negative integers. The harness shapes how the agent searches; the weights make it a domain expert.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

Read the paper: Github Repo:

Фото профиля Filecoin
Filecoin4 месяцев назад

@hexoai A model that updates its own weights needs verifiable checkpoints, cryptographic proof of what version existed at what point. Without that, you can't audit what it learned or when. That's a storage problem as much as a safety one.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Fair

Фото профиля Filecoin
Filecoin4 месяцев назад

@hexoai glad it resonated

Фото профиля Piyush Garg
Piyush Garg4 месяцев назад

@hexoai Recursive compound improvement agent is actually an interesting idea.

Фото профиля Alok Bishoyi
Alok Bishoyi4 месяцев назад

first, congrats @kunalbhatia91 & @hexoai team on the launch i went through the repo- the harness loop update loop seems to exist. but i cant find the weight update half in the code. does sia not have any specific guidace towards carrying out rollouts / SFT etc or it is picked by free-form LLM judgment ? cause there's no specific skills or prompts related to weight updates or any integrations to do so Thanks

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai We will be co-releasing it with a partner and requires coordination with multiple stakeholders. Humans are still involved 😄

Фото профиля Alok Bishoyi
Alok Bishoyi4 месяцев назад

@hexoai sir, "weight updates" are in half the name but cool! looking forward to it and rooting for you folks

Фото профиля Rohan Paul
Rohan Paul4 месяцев назад

@hexoai wow. these numbers are wild "The gains are 56.6% on LawBench, 91.9% runtime reduction on GPU kernels, and 502% on denoising over the initial baseline"

Фото профиля Yev Marusenko
Yev Marusenko4 месяцев назад

@hexoai In my own work finding that the next growth, optimization, and scale always comes from improving a process or the layer. Self improvement is an untapped area now. Big findings! Congrats on the publication and launch! Will see if @karpathy has any thoughts on it.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai @karpathy Recursive self improvement is in line with the bitter lesson, it’s the next primitive

Фото профиля Wesley
Wesley4 месяцев назад

@hexoai This is next-level 🔥 SIA’s recursive loop blending harness evolution + live weight updates feels like the missing piece for true agentic growth. Those 500%+ jumps on RNA denoising? Insane. Open-source self-improvers just won. Diving in now! 🚀

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Weights is the real breakthrough!

Фото профиля Krishna Mehra
Krishna Mehra4 месяцев назад

@hexoai Amazing to see where Sia will take civilization!

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai To infinity ∞

Фото профиля Neo Kim
Neo Kim4 месяцев назад

@hexoai this is pretty wild! we now have self improving agents. taking a closer read of the paper.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai The future is here

Фото профиля Dylan Knox
Dylan Knox4 месяцев назад

@hexoai Recursive self-improvement is the unlock that changes everything.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai It’s what will take us to ∞

Фото профиля Nawi
Nawi4 месяцев назад

@hexoai Huge moment for AI agents. Hexo Labs’ SIA just topped MLE-Bench, outperformed research agents, and improved beyond its own prior version. The self-improving era is starting.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Not just came out on top, but kept improving to beat itself again and again

Фото профиля Samuel Verboomen
Samuel Verboomen4 месяцев назад

@hexoai 🚀 Super excited and proud to be part of this journey. Happy to connect and exchange on the topic 💬 Excited for what’s ahead ✨

Фото профиля MONICA THUKKARAM
MONICA THUKKARAM4 месяцев назад

@hexoai Congratulations to the team @hexoai

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Thanks @MonicaThukkaram! Wonderful to be collaborating with you at @Stanford

Фото профиля MONICA THUKKARAM
MONICA THUKKARAM4 месяцев назад

@hexoai @Stanford Looking forward to our collaboration. This is Massive!

Фото профиля Rathin Shah
Rathin Shah4 месяцев назад

@hexoai Good stuff ser! 🫡 Is this the last company we ever need?

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai It will humanity’s last invention!

Фото профиля Rebecca
Rebecca4 месяцев назад

@hexoai Incredible achievement! Congrats to the entire team!

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai The future is here

Фото профиля ZOYA ✪
ZOYA ✪4 месяцев назад

@hexoai This feels like an important shift in AI agents.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai It is. Recursive self improvement will be the ultimate lever of progress

Фото профиля ZOYA ✪
ZOYA ✪4 месяцев назад

@hexoai awesome

Фото профиля Tyler Wayne
Tyler Wayne4 месяцев назад

@hexoai Open-sourcing this is the most important part of the announcement.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai It shouldn’t be gatekept

Фото профиля Dharmik Harinkhede
Dharmik Harinkhede4 месяцев назад

@hexoai Self-improving AI is the next major leap. An AI that upgrades both its workflow and its own weights while solving problems is a fascinating direction for the future of intelligence. 🚀

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@Dharmikpawar31 @hexoai It is the ultimate lever

Фото профиля Csaba Kissi
Csaba Kissi4 месяцев назад

Most agent frameworks are still basically static: plan, call tools, execute. SIA is interesting because it focuses on the feedback loop itself - the harness, model, and memory layer improving from previous runs. That’s a much more important direction than just making agents “do more steps.”

Фото профиля Aiswarya Sankar
Aiswarya Sankar4 месяцев назад

@hexoai Huge!!! So exciting to see how far this has come congrats @kunalbhatia91 and @hexoai team

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Thanks @Aiswarya_Sankar

Фото профиля Linus ✦ Ekenstam
Linus ✦ Ekenstam4 месяцев назад

@hexoai guys, what’s to most impressive external use-case you’ve seen during early testing?

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai 15x speedup in GPU Kernel Optimisation

Фото профиля Chanukya Patnaik
Chanukya Patnaik4 месяцев назад

@hexoai Kudos Vignesh, Kunal and Hexo team

Фото профиля Farlane
Farlane4 месяцев назад

@hexoai Self-improving agents feel like the next major leap for AI infrastructure. Curious to watch how SIA evolves from here.

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai SIA will be a living, breathing organism by itself

Фото профиля Julius Ritter
Julius Ritter3 месяцев назад

@hexoai Congrats guys, this is so big!

Фото профиля Subodh Kolhe
Subodh Kolhe4 месяцев назад

@hexoai good stuff Bhatia🫡

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Thanks!

Фото профиля Saumya Saxena
Saumya Saxena4 месяцев назад

@hexoai Incredible work man! Excited to see it come alive

Фото профиля Supriya Sharma
Supriya Sharma4 месяцев назад

@hexoai Congratulations @hexoai This is phenomenal

Фото профиля Abhinav Gupta
Abhinav Gupta4 месяцев назад

@hexoai Machao bhai

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Thanks!

Фото профиля Rahil Bhansali
Rahil Bhansali4 месяцев назад

@hexoai This is awesome! Excited to try it

Фото профиля Kunal Bhatia
Kunal Bhatia4 месяцев назад

@hexoai Thanks @rahilbhansali!

Фото профиля Tamaz Gadaev
Tamaz Gadaev4 месяцев назад

@hexoai great thread! congrats with the launch

Фото профиля usman
usman4 месяцев назад

@BenHolfeld @hexoai excited to try this!

Фото профиля Jay
Jay4 месяцев назад

@hexoai The future belongs to AI that learns how to get better while working. 💡

Фото профиля arc.
arc.4 месяцев назад

@beckfastattiffs @hexoai congrats on the launch! love to see it

Фото профиля Lily yang
Lily yang4 месяцев назад

@hexoai Wonderful

Фото профиля Harish Uthayakumar
Harish Uthayakumar4 месяцев назад

@hexoai Let’s go bro!

Похожие видео

New Paper! Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents A longstanding goal of AI research has been the creation of AI that can learn indefinitely. One path toward that goal is an AI that improves itself by rewriting its own code, including any code responsible for learning. That idea, known as a Gödel Machine, proposed by Jürgen Schmidhuber over two decades ago, is a hypothetical self-improving AI. It optimally solves problems by recursively rewriting its own code when it can mathematically prove a better strategy, making it a key concept in meta-learning or “learning to learn.” While the theoretical Gödel Machine promised provably beneficial self-modifications, its realization relied on an impractical assumption: that the AI could mathematically prove that a proposed change in its own code would yield a net improvement before adopting it. Sakana AI, in collaboration with Jeff Clune’s lab at UBC, proposes something more feasible: a system that harnesses the principles of open-ended algorithms like Darwinian evolution to search for improvements that empirically improve performance. We call the result the Darwin Gödel Machine. DGMs leverage foundation models to propose code improvements, and use recent innovations in open-ended algorithms to search for a growing library of diverse, high-quality AI agents. Applied to practical tasks, we implemented Darwin Gödel Machine as a self-improving coding agent that rewrites its own code to improve performance on programming tasks. It creates various self-improvements, such as a patch validation step, better file viewing, enhanced editing tools, generating and ranking multiple solutions to choose the best one, and adding a history of what has been tried before (and why it failed) when making new changes (see the attached video). We believe that Darwin Gödel Machines represent a concrete step towards AI systems that can autonomously gather their own stepping stones to learn and innovate forever!

hardmaru

105,033 просмотров • 1 год назад