Video yükleniyor...
Video Yüklenemedi
we won the embodied ai hackathon at last week this is how we did it 🧵
10,669 görüntüleme • 4 ay önce •via X (Twitter)
32 Yorum

In just 2 days, we built a system of robots for assisting you on your desktop from scratch. You can do tasks such as desk cleaning or organization just by talking to a voice agent. For example, you can tell the robot to "put the screwdriver away" and it would pick up the screwdriver on one end of the table and put it to the bin at the other end of the table.

Behind the scene, this is a group of multiple systems working together. First, the robots are orchestrated by a voice agent running on LiveKit. This agent has access to the overview of the table and can use such to make plans based on user-given commands. If the user gives a task like clean the table, the agent would first check what objects are on the table and utilize the tools it have to execute the goal. The agent has access to 3 tools.

The first tool is move_to, which is a P loop for controlling the robot slider to any object on the table. The robot is localized using an April tag while the object is localized using a VLM.

The second tool is run_policy, which allows us to run any of our trained ACT policy on demand. We collected data and trained 2 policy at SPC: > Pick up is an ACT policy trained on 200 episodes of picking random stuff up at SPC. It is used when the agent wants to pick stuff on the table. We also trained a sparse reward model so the agent knows when the pick up policy has finished and can stop it. > Put down is another policy trained on 50 episodes of putting down stuff. However, it didn't work well so we just sample trajectories from the dataset randomly. Just dropping stuff doesn’t require much intelligence anyway.

The third tool is run_molmo, which utilizes MolmoACT2, a VLA that has proven great generalization capabilities on various robot embodiments. We fine-tuned this on all the dataset we gathered at SPC, so the model can better adapt to our embodiment. Results show that the model can reliably localize objects that were not in training dataset, showing its capabilities to generalize. It can occasionally pick up objects too, but it normally takes a long time to converge and is jittery. We suspect this is due to our naive method of remote inference and state sampling on the SO101 arm.

Behind the scene, the entire system is powered by our own arbitration and network infrastructure. The robot is not controlled from a single computer. The voice agent lives on one laptop. The ACT policies live on one laptop. The slider control lives on one laptop. MolmoACT2 lives on a H200 instance in Finland. This is the power of our orchestration library: The datasets, models, and source code of this project can be found here:

🙌

tk u for hosting us 🙌

@spc oh wow congrats!

@spc tk u

@spc A well deserved win 👏 How much of the arbitration logic is hand-tuned versus learned?

@spc nothing is really hand-tuned, we just let a vlm handle the orchestration the arbitration is more on the network side as we have policies and systems working on different part of the network, even as far as finland

@spc nice job!

@spc cool

@spc Tough

@spc tks man

@spc cool stuff @pham_blnh . congrats

@spc this so cool

@spc tks man

@spc Congrats!! This is so smooth!

@spc tks!!

@spc so goated

@spc ✋🏻

@spc Cool stuff, might be also due to the different in the gripper you guys are using

yeah, that can be the case we ran it first without fine-tuning on the gripper and it didn’t work fine-tuning then helped quite a lot, it could actually pick stuff, just that it took a long time to converge camera position also had a lot of play here, if we move the side camera to a position that is more in distribution with your data it works better

@spc so cool, big lerobot & livekit fan :)

@spc Very nice. Did you use two cameras only.. gripper and overhead? Or more... and rgb only or depth? Any pixel seg or straight images?

@spc 2 rgb cameras only, no post processing act policies were trained on only the arm camera moving the slider only uses the overhead camera molmo uses both

@spc You bro is the project opensource , Would love to work with it.

@spc it failed the first task too, it did not put the candy in the binh at the end of the table🥲

@spc that task was too easy, the binh would come to the candy so we had to try another task

@spc hahaha would have loved to see the demo “Put the candy in binh” *binh proceeds to go to candy and eat it* revolutionary.
Benzer Videolar
This is what we did this week at the WAR DEPARTMENT:
DOW Rapid Response
46,023 görüntüleme • 6 ay önce
This is what we did this week at the WAR DEPARTMENT:
DOW Rapid Response
63,833 görüntüleme • 6 ay önce
This is what we did this week at the War Department:
DOW Rapid Response
89,537 görüntüleme • 7 ay önce
This is what we did this week at the WAR DEPARTMENT:
DOW Rapid Response
99,010 görüntüleme • 6 ay önce
