Loading video...
Video Failed to Load
Excited to share ESI-BENCH, a benchmark for Embodied Spatial Intelligence! Most spatial reasoning benchmarks assume an oracle observer: the agent is given the right image, view, or 3D scene. But in the real world, the observer is also an actor. To understand space, agents must decide where to look,... show more
53,182 views • 4 months ago •via X (Twitter)
14 Comments

The benchmark spans 10 task categories and 29 subcategories, grounded in Spelke’s core knowledge systems. the 10 categories are: physical capacity, physical dynamics, specular reflection, perceptual grounding, metric comparison, enumerative perception, spatial relations, cognitive mapping, temporal scene understanding, and procedural sequencing.

ESI-BENCH departs from prior spatial benchmarks in three ways: 1) From sensing to competence — agents are evaluated not only on what they perceive, but on whether they know how to act to perceive it. 2) Selective sensing — agents must decide which observations are worth acquiring, instead of passively consuming redundant or uninformative inputs. 3) Resolving perceptual mirages — agents must reason through incomplete or misleading observations to uncover hidden spatial structures and physical constraints

Our experiments show that active exploration substantially improves performance over passive single-view and random multi-view settings. More importantly, agents begin to exhibit emergent spatial strategies: moving to disambiguate occlusions, seeking better viewpoints, and interacting with objects to reveal hidden evidence. Yet the main bottleneck is often not weak perception. It is action blindness: poor action choices produce poor observations, which lead to cascading reasoning errors.

Interestingly, stronger 3D grounding can help on depth-sensitive tasks, but imperfect 3D reconstruction can also hurt fine-grained spatial reasoning by corrupting object counts, relations, and occlusion structure. Spatial intelligence is not solved by “more views” or “more 3D” alone.

A human-model comparison reveals the core gap: humans actively seek the evidence needed to resolve ambiguity, while models often overcommit to early observations and fail to revise their beliefs. We hope ESI-BENCH helps push AI systems beyond passive visual understanding toward agents that can actively reason, explore, and interact in 3D environments.

Very cool benchmark! Active perception and interaction are such important pieces of real-world spatial intelligence.

It indeed is!

Nmomic

Very cool work! A key missing piece in many embodied benchmarks is exactly this: the agent must decide where & how to look before reasoning. Active perception matters.

Thank you Jianwen!

Good to see that there's someone at least noticing the problem. I mean, other than @abramdemski and @ScottGarrabrant, back in 2019... (But this is AI, so I assume you're not allowed to cite papers more than 3 years old, regardless of relevance.)

@stanfordnlp emg this benchmark hype feels overdone

@arankomatsuzaki wait this actually looks clean

@harnessengr The fastest way to improve AI output? Better prompts. Try this: “Design a production-ready AI harness for [idea]. Include workflows, memory, tools, approvals, and scaling.” 20 more prompts for ChatGPT, Claude & Gemini:
