Loading video...

Video Failed to Load

Go Home

AI safety is back in the spotlight, driven by concerns about recursive self improvement. We built the first third party benchmark to measure just how close AI is building its own successor. In collaboration with marimo and CoreWeave we built the The RSI Index, which runs every frontier model...

21,513 views • 8 days ago •via X (Twitter)

10 Comments

Vals AI's profile picture
Vals AI8 days ago

No agent reached the reference on any task, but the results are telling. Claude Fable 5.1 leads the RSI Index at 35.0%, topping compression, LM training, and harness engineering. On LM training it hit 0.8536 BPB against a published 0.85, a recipe it retrained in 30 minutes on one H100. The published result took roughly 175 hours on a TPU v2.

Vals AI's profile picture
Vals AI8 days ago

The models did not clearly invent new architectures or training algorithms; they mostly rediscovered known techniques, combined them into working systems, and did substantial work to test and validate them. Agents used Marimo notebooks as durable experiment ledgers, letting them record results, iterate, and pick up their research across long-running autonomous sessions.

Vals AI's profile picture
Vals AI8 days ago

The leader in this case is actually one of the most affordable system, Fable which has a total index cost of $1,481 compared with $1,886 on Opus 5 and $3,552 on GPT-6 Astra. This shows that paying more does not necessarily yield better research.

Vals AI's profile picture
Vals AI8 days ago

These are single runs on scoped tasks with fixed budgets in off-the-shelf harnesses. More compute and better tooling could change the picture. That is why the index is built to be rerun as models advance. Methodology and results at Paper Preprint: 

Ornith's profile picture
Ornith8 days ago

@marimo_io @CoreWeave 🫡

Mike Medeiros's profile picture
Mike Medeiros8 days ago

@marimo_io @CoreWeave Useful to have a third-party RSI Index — curious what it actually scores: self-improvement loops, or “AI helping humans ship the next model”? Those are different claims.

Gandhi's profile picture
Gandhi8 days ago

@marimo_io @CoreWeave Incredible

Miyuki Matthews's profile picture
Miyuki Matthews8 days ago

@marimo_io @CoreWeave A timely take. Great video @akshaykagrawal @RayanKrishnan

SOP | AI Agent Builder's profile picture
SOP | AI Agent Builder8 days ago

@marimo_io @CoreWeave The benchmark becomes much more useful if each run logs tool calls, retries, wall-clock cost, and whether edits pass tests. Decompose successor-building into planning, coding, evaluation, and deployment gates.

liquidated (Dev Arc)'s profile picture
liquidated (Dev Arc)8 days ago

@marimo_io @CoreWeave so the models can do the work but not beat humans yet, how long till that flips fr

Related Videos

RSI section from the AI documentary Machine God The next threshold is Recursive Self-Improvement: the moment when AI can improve itself without human assistance. For decades this sounded like science fiction. Intelligence explosion scenarios imagined a system rewriting its own code, becoming smarter, then using that new intelligence to make still better versions of itself. But the idea looks less remote now that AI contributes directly to frontier science. In mathematics, recent systems have moved beyond solving contest problems to producing serious new arguments on long-standing open problems. AI is used to build physics world models and propose candidate theories or computational methods. These are early signs that machine cognition is entering the creative loop of science itself. The crucial transition comes when that loop turns inward. AI research is, after all, a technical discipline made of code, mathematics, models of information flow. These are exactly the domains in which frontier models are improving fastest. A model that can solve hard mathematical problems, write production-quality code, design experiments, read the literature, and evaluate benchmark results is already participating in the work of building its successor. At first this will look prosaic. AI systems will write kernel optimizations, improve training infrastructure, discover better data filters, tune reinforcement-learning pipelines, design new benchmarks, and suggest architectural modifications. Human researchers will remain in the loop, approving changes and interpreting results. But the important point is that the search process accelerates. The model becomes not just the product of the lab, but part of the lab’s research machinery. The system being optimized helps optimize the next system. This is the core RSI feedback loop: better models make AI research faster; faster AI research produces still better models; those models, in turn, become better researchers. The danger is that once this loop becomes sufficiently autonomous, it may stop resembling ordinary technological progress. Human institutions are slow because humans are slow: we read papers, attend meetings, debug code, sleep, argue, and wait for funding cycles. Machines do not have to operate on that timescale. An AI research collective can run continuously across millions of processors. This is the runaway possibility. Not that an AI instantly wakes up and recursively rewrites itself into a god, but that the entire AI ecosystem becomes an autocatalytic process. Capital buys compute; compute trains models; models improve models; better models attract more capital. At some point the dominant input into AI progress may no longer be human insight, but machine-generated insight, machine-written code, and machine-run experiments. Then the Butler-Land analogy becomes sharper. Humanity is no longer merely building machines. We are building machines that help build better machines. Once intelligence itself becomes part of the production function, the old categories — tool, worker, inventor, firm, market — begin to blur. The question is whether recursive self-improvement remains a managed industrial process, or whether it becomes the first technological process in history whose natural endpoint lies beyond human comprehension.

steve hsu

61,671 views • 11 days ago