正在加载视频...
视频加载失败
We did a comparison study on where a robot's intelligence should live by using the same model (OpenAI GPT6-Astra), same bimanual YAM station, same LEGO pick-and-place task, all run on SE3' real-world remote eval stack: 1. Frozen code-as-policy (built on NVIDIA ENPIRE components): 6/20 placements 2. Agentic control +... show more
17,096 次观看 • 5 天前 •via X (Twitter)
6 条评论

Here’s how the two approaches worked: 1. Code-as-policy: Astra wrote and tested a Python policy, which we then froze for evaluation. The policy ran without further model calls. 2. Agentic control: Each episode started with a fresh Astra session that read camera snapshots and issued commands through a control harness.

Getting hold of the brick wasn’t enough. What mattered was robustness and recovery. The code-as-policy lifted the brick in 14/20 episodes but completed only 6 placements. In the other eight, it stopped before reaching the release pose or held the brick without releasing it. Agentic control completed 18/20 placements. Nine required grasp retries, including one that recovered from a drop.

The agent recovered from its own misses: 9 of its 18 completions needed more than one grasp. In episode 8, it missed five times, lifted on the sixth, dropped the brick, reacquired it and released at 419s. The code-as-policy never made a second grasp attempt.

Agentic control completed more placements, but successful runs took longer. Among successful episodes, median time to release was 131 seconds for code-as-policy versus 175 seconds for agentic control. Agentic control also had more time to recover: up to 600 seconds, compared with 180 seconds for code-as-policy. The timing gap reflects more than model reasoning alone.

Full report and all 40 episodes on video: Reach out if you have a policy or an idea you’d like to evaluate on real robots!

@OpenAI @se3labs @nvidia Keeping the model in the loop lets it adjust each movement from live feedback, rather than letting errors compound through a fixed sequence of actions.
