Loading video...

Video Failed to Load

Go Home

Excited to share ESI-BENCH, a benchmark for Embodied Spatial Intelligence! Most spatial reasoning benchmarks assume an oracle observer: the agent is given the right image, view, or 3D scene. But in the real world, the observer is also an actor. To understand space, agents must decide where to look,...

53,182 views • 4 months ago •via X (Twitter)

14 Comments

Yining Hong's profile picture
Yining Hong4 months ago

The benchmark spans 10 task categories and 29 subcategories, grounded in Spelke’s core knowledge systems. the 10 categories are: physical capacity, physical dynamics, specular reflection, perceptual grounding, metric comparison, enumerative perception, spatial relations, cognitive mapping, temporal scene understanding, and procedural sequencing.

Yining Hong's profile picture
Yining Hong4 months ago

ESI-BENCH departs from prior spatial benchmarks in three ways: 1) From sensing to competence — agents are evaluated not only on what they perceive, but on whether they know how to act to perceive it. 2) Selective sensing — agents must decide which observations are worth acquiring, instead of passively consuming redundant or uninformative inputs. 3) Resolving perceptual mirages — agents must reason through incomplete or misleading observations to uncover hidden spatial structures and physical constraints

Yining Hong's profile picture
Yining Hong4 months ago

Our experiments show that active exploration substantially improves performance over passive single-view and random multi-view settings. More importantly, agents begin to exhibit emergent spatial strategies: moving to disambiguate occlusions, seeking better viewpoints, and interacting with objects to reveal hidden evidence. Yet the main bottleneck is often not weak perception. It is action blindness: poor action choices produce poor observations, which lead to cascading reasoning errors.

Yining Hong's profile picture
Yining Hong4 months ago

Interestingly, stronger 3D grounding can help on depth-sensitive tasks, but imperfect 3D reconstruction can also hurt fine-grained spatial reasoning by corrupting object counts, relations, and occlusion structure. Spatial intelligence is not solved by “more views” or “more 3D” alone.

Yining Hong's profile picture
Yining Hong4 months ago

A human-model comparison reveals the core gap: humans actively seek the evidence needed to resolve ambiguity, while models often overcommit to early observations and fail to revise their beliefs. We hope ESI-BENCH helps push AI systems beyond passive visual understanding toward agents that can actively reason, explore, and interact in 3D environments.

EB1A Experts's profile picture
EB1A Experts4 months ago

Very cool benchmark! Active perception and interaction are such important pieces of real-world spatial intelligence.

Yining Hong's profile picture
Yining Hong4 months ago

It indeed is!

Zavian Pokharkar's profile picture
Zavian Pokharkar2 months ago

Nmomic

Jianwen Xie's profile picture
Jianwen Xie4 months ago

Very cool work! A key missing piece in many embodied benchmarks is exactly this: the agent must decide where & how to look before reasoning. Active perception matters.

Yining Hong's profile picture
Yining Hong4 months ago

Thank you Jianwen!

David Manheim's profile picture
David Manheim4 months ago

Good to see that there's someone at least noticing the problem. I mean, other than @abramdemski and @ScottGarrabrant, back in 2019... (But this is AI, so I assume you're not allowed to cite papers more than 3 years old, regardless of relevance.)

Strata's profile picture
Strata4 months ago

@stanfordnlp emg this benchmark hype feels overdone

Strata's profile picture
Strata4 months ago

@arankomatsuzaki wait this actually looks clean

Dhruv lunawath's profile picture
Dhruv lunawath4 months ago

@harnessengr The fastest way to improve AI output? Better prompts. Try this: “Design a production-ready AI harness for [idea]. Include workflows, memory, tools, approvals, and scaling.” 20 more prompts for ChatGPT, Claude & Gemini:

Related Videos

Dr. Fei-Fei Li just called out the biggest blind spot in the entire AI industry. We have been building half of human intelligence. And calling it the finish line. Li: “If you look at human intelligence, it pretty much boils down to two buckets.” The first bucket is language. Symbolic reasoning. Communication. The ability to think in words and abstractions. That’s what every major AI lab has spent the last decade building. The second bucket is the one the industry has almost entirely ignored. Li: “We call that in AI spatial intelligence.” How humans and animals perceive, navigate, and interact with the three-dimensional physical world. How we reach for objects. How we move through space. How we build and manipulate physical reality. From painting masterpieces to constructing the pyramids, non-verbal spatial intelligence is what actually shapes the world. Language describes reality. Spatial intelligence acts on it. And the gap between those two things is the gap between a chatbot and a robot. Li: “When this technology is ready, the robotic revolution is gonna start. We’re already seeing that trend.” Every robot is a moving agent. Every moving agent requires spatial intelligence to function in the real world. The humanoid robots being deployed in factories right now are hitting the ceiling of what language models alone can power. Spatial intelligence is the unlock. But Li didn’t stop at robotics. Li: “From a geopolitics point of view, this is part of the technology that goes straight into weapons.” Autonomous drone swarms. Battlefield navigation. Physical target acquisition without human oversight. Every military application of AI that operates in the real world runs on spatial intelligence. The nation that masters the transition from static text to dynamic three-dimensional perception doesn’t just win the software race. It commands the physical battlefield. The AI arms race just broke out of the data center. It’s operating in three dimensions now.

Dustin

122,861 views • 7 months ago

Today at Stanford, Fei-Fei Li (Fei-Fei Li),Cofounder/CEO World Labs, gave one of the clearest explanations I’ve heard of what a World Model really is. She broke it down into three layers: 1️⃣ Rendering — What does the world look like? This is where most of today’s video generation models operate: generating increasingly realistic and beautiful pixels. The question is: Can AI generate what the world looks like? The primary consumer is humans. 2️⃣ Simulation — How does the world actually work? Fei-Fei gave a simple example: “How will this bottle move? If I pour the water out, how will the water flow?” This goes far beyond generating something that looks realistic. The model needs to understand physics, spatial relationships, cause and effect, and how the world changes over time. The consumers are both humans and machines. 3️⃣ Planning — What should happen next? This is where things get really interesting. AI doesn't just render the world or simulate what might happen. It uses its understanding of the world to decide: What should I do next? At this layer, the primary consumer is the machine itself. And this connects directly to two enormous opportunities: Autonomous driving and robotics. The progression is powerful: Rendering → Simulation → Planning The real promise of World Models isn't simply generating better videos. It's building AI that can understand the world, predict what happens next, and ultimately take intelligent action in the physical world.

PaulFang

11,774 views • 1 month ago

Dr Fei-Fei-Li explains with a simple example how everyday household chores are so extremely difficult for Robots. "If you tell a robot to open the top drawer and watch out for the vase, this is actually a really hard task for robots." because the robot must ground language into the real world. Words like "top", "drawer", and "vase" are abstract. The system has to map them to 3D locations, objects, and relations in a noisy scene. This requires robust perception, object recognition, and spatial reasoning under uncertainty. The robot also lacks human commonsense. "Watch out" implies predicting consequences, estimating clearances, and understanding that vases are fragile. Encoding such priors, like how heavy a drawer is or how a vase might tip, is very complex and difficult without rich world knowledge. Learning the behavior from rewards is tough. The success signal is very sparse here, so naive exploration almost never stumbles on a full success sequence. This makes policy learning sample inefficient and brittle, especially when the environment changes between training and deployment. A sparse reward situation is when the agent only gets a success signal at the very end, and gets little or no feedback along the way. If a robot must open a drawer without hitting a vase, it might get reward only if the drawer ends up open and the vase is intact. Every partial try before that looks the same to the learner, reward equals 0. --- From "DSAI by Dr. Osbert Tay" YT channel

Rohan Paul

342,627 views • 10 months ago

The teams shipping AI agents right now are bleeding money on the dumbest possible expense: teaching a 400B-parameter model to read a file name. Every time an AI agent needs to "see" something today, it routes an image through a frontier model. OCR, object detection, checking if a button exists on screen. You're paying GPT-4o or Claude pricing for tasks that require perception, not reasoning. One agent workflow processing a few thousand screenshots per day can burn through more on vision calls than on the actual thinking. Perceptron's Isaac is 2B parameters. Built by the team that created Meta's Chameleon multimodal models. On perceptive benchmarks, it matches or beats models 50x its size. The VQA, OCR, and object detection scores are competitive with models running on infrastructure that costs orders of magnitude more. The MCP wrapper is the distribution play. One install command and every Claude Code agent can offload vision tasks to a model that runs on a single consumer GPU. The agent keeps its reasoning in the frontier model and routes perception to a specialist. That split is how you get vision-heavy agent workflows from "technically possible but expensive" to "cheap enough to run on everything." This is the same pattern that won in every other compute-intensive stack. General-purpose handles orchestration. Specialists handle the heavy lifting. Graphics went through it. Audio went through it. Video encoding went through it. Vision in AI agents is next. The teams building agents that see 10,000 images a day will care about this before anyone else does.

Aakash Gupta

55,978 views • 6 months ago