Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

we won the embodied ai hackathon at last week this is how we did it 🧵

10,669 görüntüleme • 4 ay önce •via X (Twitter)

32 Yorum

Binh profil fotoğrafı
Binh4 ay önce

In just 2 days, we built a system of robots for assisting you on your desktop from scratch. You can do tasks such as desk cleaning or organization just by talking to a voice agent. For example, you can tell the robot to "put the screwdriver away" and it would pick up the screwdriver on one end of the table and put it to the bin at the other end of the table.

Binh profil fotoğrafı
Binh4 ay önce

Behind the scene, this is a group of multiple systems working together. First, the robots are orchestrated by a voice agent running on LiveKit. This agent has access to the overview of the table and can use such to make plans based on user-given commands. If the user gives a task like clean the table, the agent would first check what objects are on the table and utilize the tools it have to execute the goal. The agent has access to 3 tools.

Binh profil fotoğrafı
Binh4 ay önce

The first tool is move_to, which is a P loop for controlling the robot slider to any object on the table. The robot is localized using an April tag while the object is localized using a VLM.

Binh profil fotoğrafı
Binh4 ay önce

The second tool is run_policy, which allows us to run any of our trained ACT policy on demand. We collected data and trained 2 policy at SPC: > Pick up is an ACT policy trained on 200 episodes of picking random stuff up at SPC. It is used when the agent wants to pick stuff on the table. We also trained a sparse reward model so the agent knows when the pick up policy has finished and can stop it. > Put down is another policy trained on 50 episodes of putting down stuff. However, it didn't work well so we just sample trajectories from the dataset randomly. Just dropping stuff doesn’t require much intelligence anyway.

Binh profil fotoğrafı
Binh4 ay önce

The third tool is run_molmo, which utilizes MolmoACT2, a VLA that has proven great generalization capabilities on various robot embodiments. We fine-tuned this on all the dataset we gathered at SPC, so the model can better adapt to our embodiment. Results show that the model can reliably localize objects that were not in training dataset, showing its capabilities to generalize. It can occasionally pick up objects too, but it normally takes a long time to converge and is jittery. We suspect this is due to our naive method of remote inference and state sampling on the SO101 arm.

Binh profil fotoğrafı
Binh4 ay önce

Behind the scene, the entire system is powered by our own arbitration and network infrastructure. The robot is not controlled from a single computer. The voice agent lives on one laptop. The ACT policies live on one laptop. The slider control lives on one laptop. MolmoACT2 lives on a H200 instance in Finland. This is the power of our orchestration library: The datasets, models, and source code of this project can be found here:

South Park Commons profil fotoğrafı
South Park Commons4 ay önce

🙌

Binh profil fotoğrafı
Binh4 ay önce

tk u for hosting us 🙌

Dee profil fotoğrafı
Dee4 ay önce

@spc oh wow congrats!

Binh profil fotoğrafı
Binh4 ay önce

@spc tk u

AUTONOMOUS profil fotoğrafı
AUTONOMOUS4 ay önce

@spc A well deserved win 👏 How much of the arbitration logic is hand-tuned versus learned?

Binh profil fotoğrafı
Binh4 ay önce

@spc nothing is really hand-tuned, we just let a vlm handle the orchestration the arbitration is more on the network side as we have policies and systems working on different part of the network, even as far as finland

Elliot Arledge profil fotoğrafı
Elliot Arledge4 ay önce

@spc nice job!

Oliver profil fotoğrafı
Oliver4 ay önce

@spc cool

Damian profil fotoğrafı
Damian4 ay önce

@spc Tough

Binh profil fotoğrafı
Binh4 ay önce

@spc tks man

Vinh Tran profil fotoğrafı
Vinh Tran4 ay önce

@spc cool stuff @pham_blnh . congrats

Boshen Zhang profil fotoğrafı
Boshen Zhang4 ay önce

@spc this so cool

Binh profil fotoğrafı
Binh4 ay önce

@spc tks man

Emerson S profil fotoğrafı
Emerson S4 ay önce

@spc Congrats!! This is so smooth!

Binh profil fotoğrafı
Binh4 ay önce

@spc tks!!

josh profil fotoğrafı
josh4 ay önce

@spc so goated

Ghosh profil fotoğrafı
Ghosh4 ay önce

@spc ✋🏻

Jiafei Duan profil fotoğrafı
Jiafei Duan4 ay önce

@spc Cool stuff, might be also due to the different in the gripper you guys are using

Binh profil fotoğrafı
Binh4 ay önce

yeah, that can be the case we ran it first without fine-tuning on the gripper and it didn’t work fine-tuning then helped quite a lot, it could actually pick stuff, just that it took a long time to converge camera position also had a lot of play here, if we move the side camera to a position that is more in distribution with your data it works better

Saïd Aitmbarek profil fotoğrafı
Saïd Aitmbarek4 ay önce

@spc so cool, big lerobot & livekit fan :)

Raj profil fotoğrafı
Raj4 ay önce

@spc Very nice. Did you use two cameras only.. gripper and overhead? Or more... and rgb only or depth? Any pixel seg or straight images?

Binh profil fotoğrafı
Binh4 ay önce

@spc 2 rgb cameras only, no post processing act policies were trained on only the arm camera moving the slider only uses the overhead camera molmo uses both

Pratham Jain profil fotoğrafı
Pratham Jain4 ay önce

@spc You bro is the project opensource , Would love to work with it.

Arnie Ramesh profil fotoğrafı
Arnie Ramesh4 ay önce

@spc it failed the first task too, it did not put the candy in the binh at the end of the table🥲

Binh profil fotoğrafı
Binh4 ay önce

@spc that task was too easy, the binh would come to the candy so we had to try another task

Arnie Ramesh profil fotoğrafı
Arnie Ramesh4 ay önce

@spc hahaha would have loved to see the demo “Put the candy in binh” *binh proceeds to go to candy and eat it* revolutionary.

Benzer Videolar