Loading video...
Video Failed to Load
TL;DR: We bring coding agents to the physical world. Introducing Thea, a harness for embodied agents: one agentic loop drives a robot the way it drives a coding agent — every capability is a callable tool, and long-horizon behavior emerges from composing them. ▶️ For example: "Grab me a... show more
37,797 views • 1 month ago •via X (Twitter)
13 Comments

Project page: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. Each robot capability can likewise become a callable tool, and a language model orchestrates them through the same agentic loop. At first glance, the transfer should be a simple matter of redefining the tools. It is not. The physical world withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. (2/5)

Gap 1: reading the world. Source code is text: a model can read it, grep it, diff it. The physical world simply exists. A robot gets a 30 fps stream of RGB-D frames and sees only a momentary slice, never the whole. A coding agent's codebase is free; a physical agent's must be constructed. So Thea constructs it. 🗺 An object-centric scene graph turns fleeting pixels into persistent symbols: cup_2 keeps its position, confidence, and freshness across turns, and a compact brief enters the context before every decision. Scene Graph as Context. (3/5)

Gap 2: judging the outcome. In software, processes exit, commands return exit codes, tests print PASS or FAIL. In the physical world, the robot closes its gripper. Did it grasp the cup, or did the cup slip? Did it close on air, or on the rim? Nothing reports back. So Thea supplies the missing signal. ✅ After every manipulation, an independent evaluator checks the tool's post-condition and returns a verdict with the cause: "the robot stopped too far from the bottle to grasp it." The exit code the world never issues. Evaluation as Exit Codes. (4/5)

The rest is inherited from coding agents, modified as the physical world requires: context engineering assigns every input a lifetime (resident, refreshed, accumulated), skills load knowledge on demand, memory turns each task into per-tool experience, safety checks live in hooks below the model, and the user stays one tool call away. Through Thea, we demonstrate that with carefully adapted designs, the harness behind coding agents fits embodied agents well: 🧩 Extensible: adding a tool extends the agent, and complex behaviors emerge from free composition. 🔌 Portable: the model and the body are plug-and-play; one harness runs Astribot S1, AgileX Cobot Magic, and Unitree G1. 🌉 A bridge between user and world: instructions flow through dialogue, while the agent runs the perception-action loop in the physical world. 📄 💻 🎬 (5/5) Full demo:

So cool!

This is super cool, the good part of this type of system or harness is that you are converting latent understanding to more human readable form, which is reliable to debug, but at inference time it's kinda two step process, which is kinda producing an additional lag?

Pretty cool

The tool-call abstraction transfers cleanly right up until failure does. A coding agent retries a failed call for free; a robot has already moved something. Long-horizon composition in the physical world needs a way to mark the calls that cannot be replayed, and that has no equivalent in the software version.

How’s this different from VLA?

Orthogonal rather than competing. A VLA is a policy, Thea is the harness around policies. It runs an agentic loop that composes capabilities (a VLA can be one of its tools) into long-horizon behavior.

In the physical world, state must be inferred. That makes perception a core part of the agentic loop, not just an input.

but there should be at least 3 fingers to grasp most things

Ai-embodied .com interested?
