Загрузка видео...

Не удалось загрузить видео

На главную

AI safety is back in the spotlight, driven by concerns about recursive self improvement. We built the first third party benchmark to measure just how close AI is building its own successor. In collaboration with marimo and CoreWeave we built the The RSI Index, which runs every frontier model...

21,513 просмотров • 8 дней назад •via X (Twitter)

Комментарии: 10

Фото профиля Vals AI
Vals AI8 дней назад

No agent reached the reference on any task, but the results are telling. Claude Fable 5.1 leads the RSI Index at 35.0%, topping compression, LM training, and harness engineering. On LM training it hit 0.8536 BPB against a published 0.85, a recipe it retrained in 30 minutes on one H100. The published result took roughly 175 hours on a TPU v2.

Фото профиля Vals AI
Vals AI8 дней назад

The models did not clearly invent new architectures or training algorithms; they mostly rediscovered known techniques, combined them into working systems, and did substantial work to test and validate them. Agents used Marimo notebooks as durable experiment ledgers, letting them record results, iterate, and pick up their research across long-running autonomous sessions.

Фото профиля Vals AI
Vals AI8 дней назад

The leader in this case is actually one of the most affordable system, Fable which has a total index cost of $1,481 compared with $1,886 on Opus 5 and $3,552 on GPT-6 Astra. This shows that paying more does not necessarily yield better research.

Фото профиля Vals AI
Vals AI8 дней назад

These are single runs on scoped tasks with fixed budgets in off-the-shelf harnesses. More compute and better tooling could change the picture. That is why the index is built to be rerun as models advance. Methodology and results at Paper Preprint: 

Фото профиля Ornith
Ornith8 дней назад

@marimo_io @CoreWeave 🫡

Фото профиля Mike Medeiros
Mike Medeiros8 дней назад

@marimo_io @CoreWeave Useful to have a third-party RSI Index — curious what it actually scores: self-improvement loops, or “AI helping humans ship the next model”? Those are different claims.

Фото профиля Gandhi
Gandhi8 дней назад

@marimo_io @CoreWeave Incredible

Фото профиля Miyuki Matthews
Miyuki Matthews8 дней назад

@marimo_io @CoreWeave A timely take. Great video @akshaykagrawal @RayanKrishnan

Фото профиля SOP | AI Agent Builder
SOP | AI Agent Builder8 дней назад

@marimo_io @CoreWeave The benchmark becomes much more useful if each run logs tool calls, retries, wall-clock cost, and whether edits pass tests. Decompose successor-building into planning, coding, evaluation, and deployment gates.

Фото профиля liquidated (Dev Arc)
liquidated (Dev Arc)8 дней назад

@marimo_io @CoreWeave so the models can do the work but not beat humans yet, how long till that flips fr

Похожие видео

RSI section from the AI documentary Machine God The next threshold is Recursive Self-Improvement: the moment when AI can improve itself without human assistance. For decades this sounded like science fiction. Intelligence explosion scenarios imagined a system rewriting its own code, becoming smarter, then using that new intelligence to make still better versions of itself. But the idea looks less remote now that AI contributes directly to frontier science. In mathematics, recent systems have moved beyond solving contest problems to producing serious new arguments on long-standing open problems. AI is used to build physics world models and propose candidate theories or computational methods. These are early signs that machine cognition is entering the creative loop of science itself. The crucial transition comes when that loop turns inward. AI research is, after all, a technical discipline made of code, mathematics, models of information flow. These are exactly the domains in which frontier models are improving fastest. A model that can solve hard mathematical problems, write production-quality code, design experiments, read the literature, and evaluate benchmark results is already participating in the work of building its successor. At first this will look prosaic. AI systems will write kernel optimizations, improve training infrastructure, discover better data filters, tune reinforcement-learning pipelines, design new benchmarks, and suggest architectural modifications. Human researchers will remain in the loop, approving changes and interpreting results. But the important point is that the search process accelerates. The model becomes not just the product of the lab, but part of the lab’s research machinery. The system being optimized helps optimize the next system. This is the core RSI feedback loop: better models make AI research faster; faster AI research produces still better models; those models, in turn, become better researchers. The danger is that once this loop becomes sufficiently autonomous, it may stop resembling ordinary technological progress. Human institutions are slow because humans are slow: we read papers, attend meetings, debug code, sleep, argue, and wait for funding cycles. Machines do not have to operate on that timescale. An AI research collective can run continuously across millions of processors. This is the runaway possibility. Not that an AI instantly wakes up and recursively rewrites itself into a god, but that the entire AI ecosystem becomes an autocatalytic process. Capital buys compute; compute trains models; models improve models; better models attract more capital. At some point the dominant input into AI progress may no longer be human insight, but machine-generated insight, machine-written code, and machine-run experiments. Then the Butler-Land analogy becomes sharper. Humanity is no longer merely building machines. We are building machines that help build better machines. Once intelligence itself becomes part of the production function, the old categories — tool, worker, inventor, firm, market — begin to blur. The question is whether recursive self-improvement remains a managed industrial process, or whether it becomes the first technological process in history whose natural endpoint lies beyond human comprehension.

steve hsu

61,671 просмотров • 11 дней назад

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,621 просмотров • 2 месяцев назад