Video yükleniyor...
Video Yüklenemedi
"Does S1 exhibit physical prompt steerability: different prompts induce distinct behavior in the same environment?" Yes! Watch S1 follow 4 different video prompt recipes in the same kitchen. The first one is something far out of distribution -- "putting a plate in a toaster".
277,064 görüntüleme • 4 gün önce •via X (Twitter)
38 Yorum

Thanks @JagdeepBhatia8 for the question!

incredible stuff. long-horizon robotics and prompt steerability are really the same problem at 2 different scales.

That's true, well said!

Actually insane how fast it learns

next level robotics right here

totally obsessed with this robot 🫶

@deepakpathak, does this model just repeat the actions demonstrated in the example or it's capable of understanding the objective of those actions and execute that objective even in a different environment with different tools?

Thanks for the thoughtful question Nick - S1 understands the objective and can get there in another environment or with other tools. For example, here the demonstration uses a watering can. At deployment only a cup is available, so S1 uses the cup to water the plant:

Why the plate though?? 😭

We’ve seen many impressive results in robotics papers and controlled demos, but often under carefully curated conditions. The real test is: when will we see a robot chef at home that can reliably handle the mess, variability, and unpredictability of a real kitchen? @deepakpathak

coolest tech demo i've seen in a while

the chips and guac part was honestly impressive

it just learns by watching? that is crazy

me trying to make breakfast at 3am: puts plate in toaster

the side-by-side video is crazy, it literally just mimics perfectly

The key result isn’t merely one-shot imitation—it’s conditional policy modulation under a fixed scene. Distinct video prompts produce distinct closed-loop behaviors, suggesting S1 extracts task structure rather than replaying a memorized trajectory.

i need one of these ASAP

Putting a plate in a toaster is definitely out of distribution!

bro understood the assignment a little too well

Keeping the same kitchen across four prompts makes this a much clearer test. The unusual plate-and-toaster task helps separate following the video from guessing what usually happens in that room.

physical prompt steerability is exactly the right term for this

Plate in a toaster, weird test lol

so we can just train them using youtube videos now? wild

Not quite at the arbitrary video from YouTube level yet, but soon...

mind blowing stuff

Doing a different things back to back is seriously impressive

wow this is super cool

it used both arms so perfectly

Wow! crazy

honestly terrifying and amazing at the same time

Huge milestone for the team

le coup du toaster, franchement, c'est le detail qui devrait faire le buzz.une tâche jamais vue, zéro exemple dedans, et le robot trouve quand même une réponse qui a du sens!ça commence à ressembler à du bon sens, pas juste à de l'imitation

Skild AI is really cooking with this one 🔥

mapping human video directly to robot arms is the holy grail tbh

absolute wizardry

this is actually crazy

most robots would just freeze or break the plate!

physical AGI is arriving way faster than I thought
