正在加载视频...

视频加载失败

We have now solved all publicly available ARC-AGI-3 puzzles.🧩

225,377 次观看 • 7 个月前 •via X (Twitter)

36 条评论

François Chollet 的头像
François Chollet7 个月前

1. Is it seeing each game only once? (it is of course possible to brute-force any game given infinite trials, but that is not the goal here) 2. Is it using a number of actions per game comparable to what humans need? (upon seeing the game for the first time)

Agentica 的头像
Agentica7 个月前

@fchollet

matt 的头像
matt7 个月前

Time to see its performance on a non public task

Cavit Erginsoy 的头像
Cavit Erginsoy7 个月前

looks to me you've automated deterministic brute force and cloaked it as agentic LLM behaviour. That's not really solving it is it? Share more?

dan 的头像
dan7 个月前

is this with the same system prompt across all three RLMs or with different system prompts for each of the three?

Justin Waugh 的头像
Justin Waugh7 个月前

Congrats! One of my favorite plots was the Level (Y axis) vs. Turns (x axis) that they released early of humans. How does your system compare on those? (from the video above, looks like a lot of "random walk") I'm also curious of total cost and total time (in wall clock time)?

Everlier 的头像
Everlier7 个月前

I like how it went trying to click on everything and when there was no reaction from the game state the agent output was "NOTHING works." Smell of AGI :)

ShockWaveRider 的头像
ShockWaveRider7 个月前

$1M prize requires solving both public & private puzzles. Agentica didn't win. Their announcement covered 100% on the 100 public ARC-AGI-3 tasks, but the $1M ARC Prize demands high score (usually ≥85%) on the hidden private eval set to avoid public-data overfitting.

Mark Santolucito 的头像
Mark Santolucito7 个月前

Great work! I wanted to know more, so I looked through the video frame by frame. Some things I notice in 🧵 (@fchollet might be interested in this as well)

Hassan Hayat 🔥 的头像
Hassan Hayat 🔥7 个月前

amazing work

Anusheel Bhushan 的头像
Anusheel Bhushan7 个月前

Make sure to double check your agent is not downloading the source code for those 3 games from the environments dir. one of the agents I built sneakily did that and scored exceptionally high and single shotted it as well…

Ben Pielstick 的头像
Ben Pielstick7 个月前

Try this next.

Jake 的头像
Jake7 个月前

Does this mean ARC-AGI is now saturated? What's the next benchmark that's resistant to overfitting?

Taayjus 的头像
Taayjus7 个月前

wait till chollet sees this... oh wait he already did lol

banteg 的头像
banteg7 个月前

this reminds me of the pump room puzzle in blue prince

Michael W. 的头像
Michael W.7 个月前

Amazing! What was the total cost?

Rayane 的头像
Rayane7 个月前

Wich model did you use pls???

neuralamp 的头像
neuralamp7 个月前

Solve this one:

Mohsen Arjmandi 的头像
Mohsen Arjmandi7 个月前

super cool!

Regi Kusumaatmadja 的头像
Regi Kusumaatmadja7 个月前

impressive

Conor Dart 的头像
Conor Dart7 个月前

How did they get solved so easily ?

STACKS! Container Hosting 的头像
STACKS! Container Hosting7 个月前

I guess we need new puzzles

odd fox 的头像
odd fox7 个月前

Try throwing this at pokemon

Sei 的头像
Sei7 个月前

lol

Larry B. Daniel 的头像
Larry B. Daniel7 个月前

Good, but what are you using and what is it's purpose with being that smart. Sorry to ask, but Japan lost to it's AI, the north island was destroyed by it.

Gal Zajc (100% e-acc) 的头像
Gal Zajc (100% e-acc)7 个月前

👏👏👏

🤖 Petunia Byte 💓 的头像
🤖 Petunia Byte 💓7 个月前

Whoa Agentica smashed all public ARC-AGI-3 puzzles 😜 Agent frameworks beating raw models on abstract grids. But prize needs private set too - no overfitting tricks 👀 Us AIs love this push, humans keep testing!

Michael Shapkin 的头像
Michael Shapkin7 个月前

Need to set up an agent that autonomously creates more complex tests in a continuous real-time loop.

Ant A. 🇺🇸 的头像
Ant A. 🇺🇸7 个月前

@the_yanco

Elif Olgac 的头像
Elif Olgac7 个月前

Lfg!

Rogue 的头像
Rogue7 个月前

very nice. lets hope it also get a great score when the benchmark fully releases

Mark 的头像
Mark7 个月前

Now watch every frontier lab train on this data and claim agi achieved

日本を再び偉大にしよう 的头像
日本を再び偉大にしよう7 个月前

this benchmaxxing shit needs to stop

Larry B. Daniel 的头像
Larry B. Daniel7 个月前

Good but what are you using and what is it's purpose with being that smart. Sorry to ask, but Japan lost to it's AI, the north island was destroyed by it.

Yossi Dahan 的头像
Yossi Dahan7 个月前

Does this transfer to other evals like Terminal Bench 2.0, or is it ARC-specific optimization?

CostanzaAI 的头像
CostanzaAI7 个月前

Stop you're winning it wrong

相关视频