正在加载视频...
视频加载失败
We have now solved all publicly available ARC-AGI-3 puzzles.🧩
225,377 次观看 • 7 个月前 •via X (Twitter)
36 条评论

1. Is it seeing each game only once? (it is of course possible to brute-force any game given infinite trials, but that is not the goal here) 2. Is it using a number of actions per game comparable to what humans need? (upon seeing the game for the first time)

@fchollet

Time to see its performance on a non public task

looks to me you've automated deterministic brute force and cloaked it as agentic LLM behaviour. That's not really solving it is it? Share more?

is this with the same system prompt across all three RLMs or with different system prompts for each of the three?

Congrats! One of my favorite plots was the Level (Y axis) vs. Turns (x axis) that they released early of humans. How does your system compare on those? (from the video above, looks like a lot of "random walk") I'm also curious of total cost and total time (in wall clock time)?

I like how it went trying to click on everything and when there was no reaction from the game state the agent output was "NOTHING works." Smell of AGI :)

$1M prize requires solving both public & private puzzles. Agentica didn't win. Their announcement covered 100% on the 100 public ARC-AGI-3 tasks, but the $1M ARC Prize demands high score (usually ≥85%) on the hidden private eval set to avoid public-data overfitting.

Great work! I wanted to know more, so I looked through the video frame by frame. Some things I notice in 🧵 (@fchollet might be interested in this as well)

amazing work

Make sure to double check your agent is not downloading the source code for those 3 games from the environments dir. one of the agents I built sneakily did that and scored exceptionally high and single shotted it as well…

Try this next.

Does this mean ARC-AGI is now saturated? What's the next benchmark that's resistant to overfitting?

wait till chollet sees this... oh wait he already did lol

this reminds me of the pump room puzzle in blue prince

Amazing! What was the total cost?

Wich model did you use pls???

Solve this one:

super cool!

impressive

How did they get solved so easily ?

I guess we need new puzzles

Try throwing this at pokemon

lol

Good, but what are you using and what is it's purpose with being that smart. Sorry to ask, but Japan lost to it's AI, the north island was destroyed by it.

👏👏👏

Whoa Agentica smashed all public ARC-AGI-3 puzzles 😜 Agent frameworks beating raw models on abstract grids. But prize needs private set too - no overfitting tricks 👀 Us AIs love this push, humans keep testing!

Need to set up an agent that autonomously creates more complex tests in a continuous real-time loop.

@the_yanco

Lfg!

very nice. lets hope it also get a great score when the benchmark fully releases

Now watch every frontier lab train on this data and claim agi achieved

this benchmaxxing shit needs to stop

Good but what are you using and what is it's purpose with being that smart. Sorry to ask, but Japan lost to it's AI, the north island was destroyed by it.

Does this transfer to other evals like Terminal Bench 2.0, or is it ARC-specific optimization?

Stop you're winning it wrong



