Loading video...

Video Failed to Load

Go Home

TL;DR: We bring coding agents to the physical world. Introducing Thea, a harness for embodied agents: one agentic loop drives a robot the way it drives a coding agent — every capability is a callable tool, and long-horizon behavior emerges from composing them. ▶️ For example: "Grab me a...

37,797 views • 1 month ago •via X (Twitter)

13 Comments

Wentao Zhu's profile picture
Wentao Zhu1 month ago

Project page: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. Each robot capability can likewise become a callable tool, and a language model orchestrates them through the same agentic loop. At first glance, the transfer should be a simple matter of redefining the tools. It is not. The physical world withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. (2/5)

Wentao Zhu's profile picture
Wentao Zhu1 month ago

Gap 1: reading the world. Source code is text: a model can read it, grep it, diff it. The physical world simply exists. A robot gets a 30 fps stream of RGB-D frames and sees only a momentary slice, never the whole. A coding agent's codebase is free; a physical agent's must be constructed. So Thea constructs it. 🗺 An object-centric scene graph turns fleeting pixels into persistent symbols: cup_2 keeps its position, confidence, and freshness across turns, and a compact brief enters the context before every decision. Scene Graph as Context. (3/5)

Wentao Zhu's profile picture
Wentao Zhu1 month ago

Gap 2: judging the outcome. In software, processes exit, commands return exit codes, tests print PASS or FAIL. In the physical world, the robot closes its gripper. Did it grasp the cup, or did the cup slip? Did it close on air, or on the rim? Nothing reports back. So Thea supplies the missing signal. ✅ After every manipulation, an independent evaluator checks the tool's post-condition and returns a verdict with the cause: "the robot stopped too far from the bottle to grasp it." The exit code the world never issues. Evaluation as Exit Codes. (4/5)

Wentao Zhu's profile picture
Wentao Zhu1 month ago

The rest is inherited from coding agents, modified as the physical world requires: context engineering assigns every input a lifetime (resident, refreshed, accumulated), skills load knowledge on demand, memory turns each task into per-tool experience, safety checks live in hooks below the model, and the user stays one tool call away. Through Thea, we demonstrate that with carefully adapted designs, the harness behind coding agents fits embodied agents well: 🧩 Extensible: adding a tool extends the agent, and complex behaviors emerge from free composition. 🔌 Portable: the model and the body are plug-and-play; one harness runs Astribot S1, AgileX Cobot Magic, and Unitree G1. 🌉 A bridge between user and world: instructions flow through dialogue, while the agent runs the perception-action loop in the physical world. 📄 💻 🎬 (5/5) Full demo:

Karolina Dubiel's profile picture
Karolina Dubiel1 month ago

So cool!

Anindyadeep's profile picture
Anindyadeep1 month ago

This is super cool, the good part of this type of system or harness is that you are converting latent understanding to more human readable form, which is reliable to debug, but at inference time it's kinda two step process, which is kinda producing an additional lag?

Clara Lafever Jane's profile picture
Clara Lafever Jane1 month ago

Pretty cool

Leo Lin's profile picture
Leo Lin1 month ago

The tool-call abstraction transfers cleanly right up until failure does. A coding agent retries a failed call for free; a robot has already moved something. Long-horizon composition in the physical world needs a way to mark the calls that cannot be replayed, and that has no equivalent in the software version.

RANA G's profile picture
RANA G1 month ago

How’s this different from VLA?

Wentao Zhu's profile picture
Wentao Zhu1 month ago

Orthogonal rather than competing. A VLA is a policy, Thea is the harness around policies. It runs an agentic loop that composes capabilities (a VLA can be one of its tools) into long-horizon behavior.

David Vainer MD PhD DSc's profile picture
David Vainer MD PhD DSc1 month ago

In the physical world, state must be inferred. That makes perception a core part of the agentic loop, not just an input.

Hosein R's profile picture
Hosein R1 month ago

but there should be at least 3 fingers to grasp most things

Domain Sale's profile picture
Domain Sale1 month ago

Ai-embodied .com interested?

Related Videos