Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set. [schema] makes an LLM think like a physicist. 🧵

956,732 görüntüleme • 2 ay önce •via X (Twitter)

40 Yorum

Haven Feng profil fotoğrafı
Haven Feng2 ay önce

ARC-AGI-3 gives an agent a 64×64 grid plus legal actions, no rules, stated goal, or reward. The agent must discover both what the world is and how it works like a physicist: 1. State grounding -> identify objects, relations, and goals. 2. Mechanism discovery -> infer how these states change.

Haven Feng profil fotoğrafı
Haven Feng2 ay önce

[schema] handles the state and mechanism in one editable program, a symbolic world model. It designs experiments to verify hypotheses, backtests the program against history, and plans inside its world at zero action cost.

Haven Feng profil fotoğrafı
Haven Feng2 ay önce

[schema]'s saturation of the ARC-AGI-3 public set is only a starting point. There is much more to explore! Full blog: Agent traces: Amazing team effort with @guanningzeng, @Jiani_Wang_, @wenjie_ma, @shaofeng_y27736, @ChenyangWa70207, @lustralisk95, @akanazawa, @wodenimoni, @xiuyu_l and @Zanette_ai

Haven Feng profil fotoğrafı
Haven Feng2 ay önce

@Jiani_Wang_ @wenjie_ma @shaofeng_y27736 @arcprize

Burny - Effective Curiosity profil fotoğrafı
Burny - Effective Curiosity2 ay önce

thinking like a physicist is all you need?

JFPuget 🇫🇷🇺🇦🇨🇦🇬🇱 profil fotoğrafı
JFPuget 🇫🇷🇺🇦🇨🇦🇬🇱2 ay önce

What is your take about overfitting to these 25 games? My experience on arc agi2 is that it is easy to overfit to public data. I don;'t know if it true for arc agi3 yet, but the jury is out.

Karol profil fotoğrafı
Karol2 ay önce

broooo this benchmark was planned to remain viable a little longer than that xD Well played

J Vijayavallabh 🧬/acc profil fotoğrafı
J Vijayavallabh 🧬/acc2 ay önce

Can you release harness as a skill for people to use in claude code and codex?

Dariusz Parzygnat profil fotoğrafı
Dariusz Parzygnat2 ay önce

ok, you beat the kimi k3 announcement. that is massive.

IBAN profil fotoğrafı
IBAN2 ay önce

Wow This is insane

Angela Dai profil fotoğrafı
Angela Dai2 ay önce

very cool :)

Haven Feng profil fotoğrafı
Haven Feng2 ay önce

Thanks!!!

LuisAlfonsoHernandez profil fotoğrafı
LuisAlfonsoHernandez2 ay önce

would be interesting to try with Kimi K3

alejandro profil fotoğrafı
alejandro2 ay önce

can we use it?

Xingyu Dang profil fotoğrafı
Xingyu Dang2 ay önce

How is this related to

Pud ⚛️🧪🐼 profil fotoğrafı
Pud ⚛️🧪🐼2 ay önce

The current verified score is 7.8% 🤯

Marko Kraemer profil fotoğrafı
Marko Kraemer2 ay önce

is the harness / repo going to be made open source?

Mojtaba Tabatabaie profil fotoğrafı
Mojtaba Tabatabaie2 ay önce

interesting, when is it available to test?

Aurel Prosz profil fotoğrafı
Aurel Prosz2 ay önce

Tokenwise how efficient is this? Congratz by the way, this is awesome!

Josh Cason profil fotoğrafı
Josh Cason2 ay önce

Is it possible there was some cheating somehow?

William Lamkin profil fotoğrafı
William Lamkin2 ay önce

would be cool to see a pluralistic world-model harness that externalizes a given user’s recurring lenses/perspectives as inspectable, competing, revisable cognitive instruments; recruits specialized agents to advocate, translate, falsify, experiment, and build; and uses surprise, provenance, and artifact production to turn an evolving personal worldview into a cumulative research institution.

Ridhvik Gopal (Vik) profil fotoğrafı
Ridhvik Gopal (Vik)2 ay önce

Wasn’t the whole point of ARC-AGI-3 to test the LLM and *NOT* the harness? What is the point here?

Michael Schwab profil fotoğrafı
Michael Schwab2 ay önce

Is this music generated by AI? I want to buy it

Haven Feng profil fotoğrafı
Haven Feng2 ay önce

Yeah! It’s made by coauthor @ChenyangWa70207 with @suno

arctic_ profil fotoğrafı
arctic_2 ay önce

The background beat is actually too fire I had to rewatch to read anything

max.jpg profil fotoğrafı
max.jpg2 ay önce

this is banger congrats

Jonathan Tien profil fotoğrafı
Jonathan Tien2 ay önce

This might be an over simplification, but sounds like instead of training a world model, the schema harness orients the LLM to help construct a world model on the fly?

Julian C profil fotoğrafı
Julian C2 ay önce

Wait, 99% RHAE? That’s a massive leap. What’s the secret sauce in [schema] that makes it click?

Sriraam profil fotoğrafı
Sriraam2 ay önce

Waiting for private set scores now

Thibaud profil fotoğrafı
Thibaud2 ay önce

Just incredible. Thank you so much for sharing.

kern profil fotoğrafı
kern2 ay önce

This is crazy!

Ishaan profil fotoğrafı
Ishaan2 ay önce

This is awesome

Andreas profil fotoğrafı
Andreas2 ay önce

Amazing presentation (and results).

Conor profil fotoğrafı
Conor2 ay önce

congrats!

lowkmessi profil fotoğrafı
lowkmessi2 ay önce

nice presentation

chris_paul_walker profil fotoğrafı
chris_paul_walker2 ay önce

Congrats can’t wait to check it out

Guillermo Barbadillo profil fotoğrafı
Guillermo Barbadillo2 ay önce

Very nice work, congratulations! One detail that I missed in the blog post is the cost of running the evaluation. Thanks!

Stefan Larsen profil fotoğrafı
Stefan Larsen2 ay önce

What about Fable/Sol-Max + Luna-Max? In terms of cost-efficiency...

SoniP_Gems🇮🇳 profil fotoğrafı
SoniP_Gems🇮🇳2 ay önce

Hey bro this was only planned for continue viable for long duration.

Stanislav Sorokin profil fotoğrafı
Stanislav Sorokin2 ay önce

What does 99% prove if the harness rewards behavior production does not need? The judge and operating spec must share one pass condition. Practical gate:

Benzer Videolar