Loading video...
Video Failed to Load
Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It spans hundreds of rooms and puzzles. Today’s best agents cannot progress beyond the initial levels.
248,983 views • 2 months ago •via X (Twitter)
62 Comments

Agents need to go through 200+ rooms to find 100 hidden gems by completing Sokoban-like box pushing puzzles. The rules are simple, but difficulty increases at a rapid pace. All puzzles were designed by hand and solved internally.

We used models in their native harnesses and gave them tools to control the camera and player. However, they perform poorly no matter the representation (Image, ASCII, and JSON), scoring 1% max. We calculated a novelty score to stop models when they walk in circles.

When we gave the agents access to Python, they tend to write A* solvers, which improves their scores considerably. GPT-5.6 Sol Max scored 13%! However, later puzzles are designed to be infeasible for A* solvers. Models did not reach those stages yet.

A special thanks to @PrimeIntellect for sponsoring, and @xeophon and @omouamoua for their mentorship. Also, a huge thanks to my brother, @mr_cavedweller, for designing 100+ rooms! Mazebench is open source: Blog:

late congrats on the release, this is amazing I only played for several minutes, but good concept, very solid puzzle design and S-tier execution great that you're testing coding harnesses as well even webpage and blog is solid it might be the very first reasoning benchmark that I can wholeheartedly recommend I can volunteer as a human baseline, assuming this thing doesn't take 100 hours 😅 If I have to say something negative: I believe it didn't save the info about the first 2 gems that I picked up (the 3rd one registered, so in the end I picked up 3, but I had only 1 visible in the UI)

thanks Psyho! yes it might take 100+ hours 😅 i have some friends playing it and after several hours they are near 10% apologies for the save bug, working on that. if you signed up and the save file disappeared it might be in the user page

yeah, I was worried about that I guess I can volunteer to test a part of it 😅 btw, not sure how fiendish and unreasonable the later puzzles are, but it might be an actually good idea to release on steam; you can even release for free just to gather data

that’s a good idea i should do steam there may be a couple bumps in the curriculum, and the later levels are wild but contain some fun ideas so they aren’t unfair

oh, did you update the game? visuals for gems have changed (much better now!) and now it tracks my gems correctly (after reacquiring them) so it seems it's fixed few suggestions for better human experience - for gems that were already picked up, it would good to render them in lower alpha (transparent) or muted colors; I wonder if agents encounter a similar problem when they try to pick up the same gems again - it's somewhat a puzzle game standard to be able to undo through a reset, probably doesn't change anything for AI - for humans, it would be good to improve movement a bit, by doing proper input buffering and speeding up animation when input buffering (+ making animation slightly snappier), but that might be tricky to implement - being able to use cursors to navigate through map would be nice - show available gems on the map (shiny pixel, highlighted border?) - human leaderboard wen? anyway, happy to help if you need some consulting on the human experience; that being said, I'm not sure if you need considering it's already quite good

these are good suggestions, thank you the undo after reset and gems on the map are something i can add easily i’ll see if i can fix visuals and make the game smoother (there is a secret human leaderboard in the user page) yes gems are half transparent if you re-enter a room

verified runs only :(

it records your play through and does a play back so it will auto verify when you submit, and if it breaks somehow i can put it on there

@legit_api Nice one chair

you should have built in some timers and chairs so people and AI's can practive patience

😂 the chair update is coming soon

christmas edition, the gems are presents and each one contains a chair

shouldn't've made it open source imo, but still extremely cool!

if it comes to that, i will come up with a secret new env even more daunting

Banger banger banger

Super excited to see this finally out!!

👀👀👀

Is it fair to say, 3d sokoban with a hard curriculum?

yep, that’s what I was shooting for!

Great idea and looks good - kudos

Great work, i love some hard environments

thanks! tried to make it as hard as possible

so this is what you do in your cave

I love it

thank you JB models will score so so low forever

Nice work

Cool!

great work ! the puzzles are amazing and so cool🔥🔥

thank you! ❤️

is the goal of these types of envs to find tasks humans can solve but LLM/VLMs can’t? The larger the gap the more interesting?

i just wanted to build something very very hard, i find the puzzles interesting.

great work! this seems fun to play as a human as well

Why is this better than ARC-AGI-3 lmao

Chair done good.

Are you gonna do multi agent puzzles so they have to cooperate?

that would be super fun

Amazing work!

Your first mistake was revealing what the bench does. Now it will be benchmaxxed thru synthetic data generation. Curse of ArcAGI

nice! impressive work

this is super cool. awsm release (+ knew it was a banger when i saw ur RL residency presentation)

thanks, i appreciate that! ❤️

can u release this as a iphone game

i should, that would be cool it is available on

looks great. congrats!

@legit_api when will we get Chair bench?

so tasteful

How do humans perform?

better than 1% which is where sota models are at

I'm disappointed in you chair, accelerating AI progress like this

cool - love the visual style! Have you measured how humans perform?

thank you! we just tested them internally. there is a leaderboard to compete in, I’d be stunned if someone can beat them all in a short amount of time

Played the first few and had quite some fun. Maybe market this as a game also? 😅

@patience_cave You mean sell it on Steam and collect the data to train better models?

@patience_cave I was thinking app store but yeah essentially. Or just to have a good game - this is definitely better than 90% of puzzle game apps

Damn, even a chair is more successful than me *sigh* (Looks really cool btw)

sickk

No one has figured out that AI needs to feel for vision to work. Lordy I am so done with the “labs” All ego. No one open to anything but what they are doing. Scale! Scale!!! Don’t bother to think. Chair, you know I love you. Not you I am angry at. I am deeply disappointed that not one person inside the labs thought to use social media to solve things together… why share credit? The labs will own 100% of nothing soon. And they deserve the failures for their blind egos and horrific leadership. OpenAI is still resetting things to make Microsoft’s numbers look better tomorrow. They still owe 250 BILLION to Azure… but they think they can hide financial issues long enough their crap IPO will go public. They owe one trillion in spending commitments. The media is hiding truth. The labs are lying… You are about to see the entire world go back to the cave days. So you will be happy. I will not. I wanted everyone to succeed. When people want to understand why the AI fails on these types of things I am happy to help for a fee now. I offered for over a year to help for free. No more. Best to you chair and all those reading this. Giving humanity a 40% chance we make it through this without a total collapse. •

Arc-agi level quality. I love it! Keep it up buddy
