Loading video...

Video Failed to Load

Go Home

Gpt-6 astra can do in-context learning on mobile manipulation! • Different environment • Different camera angle • Different layout No text prompt, it infers from video It even chooses when to use end-effector or joint space Trace recording after ⬇️

67,278 views • 15 days ago •via X (Twitter)

26 Comments

Axel's profile picture
Axel15 days ago

The agent is asked to establish a plan based on the demonstration only It self-corrects and retries if needed Honestly it's pretty mesmerizing to see, and it will keep getting faster

Axel's profile picture
Axel15 days ago

The code is available here: We haven't completely cleaned it yet but your favorite agent will be able to understand our approach

Dhruv Diddi's profile picture
Dhruv Diddi15 days ago

Nicely done! 💯🦾👏

David Dobáš's profile picture
David Dobáš15 days ago

Look at the beautiful phone teleop interface

Axel's profile picture
Axel15 days ago

woah who built that

Robert Scoble's profile picture
Robert Scoble15 days ago

It is getting faster! Congrats.

Axel's profile picture
Axel14 days ago

🫡

Andrew Lyubovsky's profile picture
Andrew Lyubovsky15 days ago

Pretty Cool !!

Miguel Gregori's profile picture
Miguel Gregori14 days ago

Cuando el entorno cambia y el robot no pide un manual nuevo, ahí hay diseño.

NAMAN RAJ's profile picture
NAMAN RAJ15 days ago

Mind blown! 🤯 That's some advanced AI magic right there! 🧙‍♂️ What do you guys think, is this the future of human-AI collaboration?

AIwithMinal's profile picture
AIwithMinal15 days ago

Pure artistic brilliance.

SEAR's profile picture
SEAR14 days ago

this is exactly the kind of learning Sear is built around

ethereagle · building's profile picture
ethereagle · building15 days ago

no text, video only. are those frames dumped straight into Astra's context, or is there a separate encoder in front? that's copy-with-the-API vs needing your stack.

Axel's profile picture
Axel15 days ago

dumped straight in

atharva ☆'s profile picture
atharva ☆15 days ago

woah incredible

Axel's profile picture
Axel15 days ago

yeah am pretty stoked

Aaron's profile picture
Aaron15 days ago

so cool! 😎

Axel's profile picture
Axel15 days ago

couldn't believe it at first tbh

Tepulous's profile picture
Tepulous15 days ago

$20/min .....

Axel's profile picture
Axel15 days ago

not if you use kv-caching

RealMan Robotics's profile picture
RealMan Robotics15 days ago

The shift from simulation to real-world manipulation is the part that really matters. Curious to see how far this generalization can scale across tasks and environments.

AI Quanting's profile picture
AI Quanting14 days ago

The control space choice is the bit Id want to see stressed. Does it switch to joint space when the demo path is actually constrained, or does it settle per task type regardless?

Hibrinix's profile picture
Hibrinix14 days ago

Meanwhile your future mobile phone

Amogh Shrivastava's profile picture
Amogh Shrivastava14 days ago

that looks peak

Axel's profile picture
Axel14 days ago

That's because it is

Degenpark's profile picture
Degenpark14 days ago

end-effector choice is huge

Related Videos