Video yükleniyor...
Video Yüklenemedi
Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set. [schema] makes an LLM think like a physicist. 🧵
956,732 görüntüleme • 2 ay önce •via X (Twitter)
40 Yorum

ARC-AGI-3 gives an agent a 64×64 grid plus legal actions, no rules, stated goal, or reward. The agent must discover both what the world is and how it works like a physicist: 1. State grounding -> identify objects, relations, and goals. 2. Mechanism discovery -> infer how these states change.

[schema] handles the state and mechanism in one editable program, a symbolic world model. It designs experiments to verify hypotheses, backtests the program against history, and plans inside its world at zero action cost.

[schema]'s saturation of the ARC-AGI-3 public set is only a starting point. There is much more to explore! Full blog: Agent traces: Amazing team effort with @guanningzeng, @Jiani_Wang_, @wenjie_ma, @shaofeng_y27736, @ChenyangWa70207, @lustralisk95, @akanazawa, @wodenimoni, @xiuyu_l and @Zanette_ai

@Jiani_Wang_ @wenjie_ma @shaofeng_y27736 @arcprize

thinking like a physicist is all you need?

What is your take about overfitting to these 25 games? My experience on arc agi2 is that it is easy to overfit to public data. I don;'t know if it true for arc agi3 yet, but the jury is out.

broooo this benchmark was planned to remain viable a little longer than that xD Well played

Can you release harness as a skill for people to use in claude code and codex?

ok, you beat the kimi k3 announcement. that is massive.

Wow This is insane

very cool :)

Thanks!!!

would be interesting to try with Kimi K3

can we use it?

How is this related to

The current verified score is 7.8% 🤯

is the harness / repo going to be made open source?

interesting, when is it available to test?

Tokenwise how efficient is this? Congratz by the way, this is awesome!

Is it possible there was some cheating somehow?

would be cool to see a pluralistic world-model harness that externalizes a given user’s recurring lenses/perspectives as inspectable, competing, revisable cognitive instruments; recruits specialized agents to advocate, translate, falsify, experiment, and build; and uses surprise, provenance, and artifact production to turn an evolving personal worldview into a cumulative research institution.

Wasn’t the whole point of ARC-AGI-3 to test the LLM and *NOT* the harness? What is the point here?

Is this music generated by AI? I want to buy it

Yeah! It’s made by coauthor @ChenyangWa70207 with @suno

The background beat is actually too fire I had to rewatch to read anything

this is banger congrats

This might be an over simplification, but sounds like instead of training a world model, the schema harness orients the LLM to help construct a world model on the fly?

Wait, 99% RHAE? That’s a massive leap. What’s the secret sauce in [schema] that makes it click?

Waiting for private set scores now

Just incredible. Thank you so much for sharing.

This is crazy!

This is awesome

Amazing presentation (and results).

congrats!

nice presentation

Congrats can’t wait to check it out

Very nice work, congratulations! One detail that I missed in the blog post is the cost of running the evaluation. Thanks!

What about Fable/Sol-Max + Luna-Max? In terms of cost-efficiency...

Hey bro this was only planned for continue viable for long duration.

What does 99% prove if the harness rewards behavior production does not need? The judge and operating spec must share one pass condition. Practical gate:
