正在加载视频...

视频加载失败

We did a comparison study on where a robot's intelligence should live by using the same model (OpenAI GPT6-Astra), same bimanual YAM station, same LEGO pick-and-place task, all run on SE3' real-world remote eval stack: 1. Frozen code-as-policy (built on NVIDIA ENPIRE components): 6/20 placements 2. Agentic control +...

17,096 次观看 • 5 天前 •via X (Twitter)

6 条评论

Paige 的头像
Paige5 天前

Here’s how the two approaches worked: 1. Code-as-policy: Astra wrote and tested a Python policy, which we then froze for evaluation. The policy ran without further model calls. 2. Agentic control: Each episode started with a fresh Astra session that read camera snapshots and issued commands through a control harness.

Paige 的头像
Paige5 天前

Getting hold of the brick wasn’t enough. What mattered was robustness and recovery. The code-as-policy lifted the brick in 14/20 episodes but completed only 6 placements. In the other eight, it stopped before reaching the release pose or held the brick without releasing it. Agentic control completed 18/20 placements. Nine required grasp retries, including one that recovered from a drop.

Paige 的头像
Paige5 天前

The agent recovered from its own misses: 9 of its 18 completions needed more than one grasp. In episode 8, it missed five times, lifted on the sixth, dropped the brick, reacquired it and released at 419s. The code-as-policy never made a second grasp attempt.

Paige 的头像
Paige5 天前

Agentic control completed more placements, but successful runs took longer. Among successful episodes, median time to release was 131 seconds for code-as-policy versus 175 seconds for agentic control. Agentic control also had more time to recover: up to 600 seconds, compared with 180 seconds for code-as-policy. The timing gap reflects more than model reasoning alone.

Paige 的头像
Paige5 天前

Full report and all 40 episodes on video: Reach out if you have a policy or an idea you’d like to evaluate on real robots!

catman 的头像
catman5 天前

@OpenAI @se3labs @nvidia Keeping the model in the loop lets it adjust each movement from live feedback, rather than letting errors compound through a fixed sequence of actions.

相关视频

36 GROK AGENTS. ASTRA ON FREE CREDITS. DUAL-MODEL ROUTING IS THE EDGE. not one chat window burning a paid invoice all day a swarm that routes cheap work to grok and only wakes gpt-6 astra when the task is actually hard ▹ the stack 36 agents on grok bot with flexible settings per role monitor, plan, write, code, QA, ship, each with its own lane part of the fleet is wired to gpt-6 astra through a china free-credit service layer free credits are not a toy promo here they are the fuel for frontier spikes without a monthly bleed ▹ dual-model routing easy jobs stay on grok: speed, volume, always-on loops hard jobs jump to astra: reasoning, long builds, sharp code the router decides by task type, not by ego if astra is not needed, astra does not spend free-credit bursts buy the expensive brain grok agents keep the factory running between bursts that is how the system feels unlimited not by breaking quotas, by refusing to waste them ▹ why it hits different most people pay frontier prices for every mid task operators split the brain and protect the credits 36 agents = parallel throughput dual routing = cost control with quality when it matters free credits = astra access without living on the invoice the constraint moved from "can i afford the model" to "did i route the job to the right model" ▹ the take single-model stacks die on bills and on boredom multi-agent + dual routing is the new default factory grok for the grind astra for the cut free credits for the spikes that used to empty the wallet bookmark this before everyone copies the route map comments: what % of your tasks actually deserve astra

cryptopsihoz

44,187 次观看 • 19 天前

I GAVE GPT-6 ASTRA AND GROK BOT $1,000 EACH AND THE BETTER TRADER F*CKING LOST grok closed $1,742.15. astra closed $1,634.27. eleven points apart same day, same page, six agents each. one order: beat the other room or i cut funding astra was the better trader on every line people brag about → coin side: astra $586.42, grok $468.90 → astra went 17 of 24 green, grok went 16 of 24 → astra caught the MARLIN low at 57.4K and the 818K top, +$462.44 on two tickets → astra took 92% of its day out of that one coin → grok took 37% of his day out of prediction books, +$273.25 → astra ran the same books for $47.85 so how did the worse trader take the money? a calendar apple held a keynote that day. grok bought IPOD at 559K into it, +$306.48, biggest ticket of the duel. then he traded the market pricing it wrong: "apple foldable iphone by sep 30", NO at 45c, closed at 97.9c both opened the same bitcoin book over 80k. astra took YES at 34.5c and watched it bleed to 13.5c, its worst book of the day. grok took NO at 58.5c and closed it at 93.5c same market, same hour, opposite sides. no benchmark tells you that. leaderboards score answers, not tickets both floors trade on one page anyone can open: a manhattan seat that reads a calendar costs $180,000 a year. mine cost less than dinner and argued all night the six-role setup i handed both sits in the article i quoted, free you can rent the smarter model, not the habit of reading tomorrow whole duel is in the clip below, tick by tick. bookmark it before the feed eats it which one would you fund after that scorecard?

savip.

36,536 次观看 • 18 天前