Video wird geladen...
Video konnte nicht geladen werden
Gpt-6 astra can do in-context learning on mobile manipulation! • Different environment • Different camera angle • Different layout No text prompt, it infers from video It even chooses when to use end-effector or joint space Trace recording after ⬇️
67,278 Aufrufe • vor 15 Tagen •via X (Twitter)
26 Kommentare

The agent is asked to establish a plan based on the demonstration only It self-corrects and retries if needed Honestly it's pretty mesmerizing to see, and it will keep getting faster

The code is available here: We haven't completely cleaned it yet but your favorite agent will be able to understand our approach

Nicely done! 💯🦾👏

Look at the beautiful phone teleop interface

woah who built that

It is getting faster! Congrats.

🫡

Pretty Cool !!

Cuando el entorno cambia y el robot no pide un manual nuevo, ahí hay diseño.

Mind blown! 🤯 That's some advanced AI magic right there! 🧙♂️ What do you guys think, is this the future of human-AI collaboration?

Pure artistic brilliance.

this is exactly the kind of learning Sear is built around

no text, video only. are those frames dumped straight into Astra's context, or is there a separate encoder in front? that's copy-with-the-API vs needing your stack.

dumped straight in

woah incredible

yeah am pretty stoked

so cool! 😎

couldn't believe it at first tbh

$20/min .....

not if you use kv-caching

The shift from simulation to real-world manipulation is the part that really matters. Curious to see how far this generalization can scale across tasks and environments.

The control space choice is the bit Id want to see stressed. Does it switch to joint space when the demo path is actually constrained, or does it settle per task type regardless?

Meanwhile your future mobile phone

that looks peak

That's because it is

end-effector choice is huge
