Loading video...

Video Failed to Load

Go Home

In 30 min, we deployed MolmoAct2 from Ai2 at @GRASPLab and it worksout-of-the-box. One cool t hing it shows: self-recovery. The model reasons the cups get knocked off, and lifts it up, even when not asked to. Here is a thread of what stood out today 🧵

13,712 views • 4 months ago •via X (Twitter)

12 Comments

Jie Wang's profile picture
Jie Wang4 months ago

Pick-and-Place: MolmoAct2 shows solid understanding of object names. Not just common ones like pineapple, it also handles weird objects like fingerspinner (though it needs a bit of human help to grasp it)

Jie Wang's profile picture
Jie Wang4 months ago

Folding: MolmoAct2-DROID has some trouble folding our cloth. It tries many times to get a firm grip. And the fold quality is,,, just okay. Let’s see if it works better on Bimanual YAM! (TODO)

Jie Wang's profile picture
Jie Wang4 months ago

Language grounding: Ideally, the policy should ONLY pick the right object from distractors. To be fair, MolmoAct2 DID place two fruit toys into the basket first. (so we label as success) But we kept running it to see if it knew the task was done. It didn't. Same behavior as pi0: keeps touching stuff around. When I put a duck into the gripper, it grasps and places it into a basket too. Do we need a sys1 policy that knows when2stop? Or is a sys2 policy doing subtask decomposition and completion checking enough?

Jie Wang's profile picture
Jie Wang4 months ago

Articulation: So far, MolmoAct2 struggles with opening our cabinet / drawer on the table. More testing on articulated objects coming

Jie Wang's profile picture
Jie Wang4 months ago

Compositional: MolmoAct2 fails to spell its 'Father's name. Could be a prompt issue. I like how @fox_dieter17849 propose Toy Block Benchmark, which needs a lot of reasoning in color, text and space!

Jie Wang's profile picture
Jie Wang4 months ago

Tool Usage: MolmoAct2 still cannot reason how to open a scissor, more research to be done! We need a smarter robot at home! @DJiafei next paper idea 😆 (To be clear, no VLAs/WAMs can do this task on our Franka, including pi05 / DreamZero etc)

Jie Wang's profile picture
Jie Wang4 months ago

@DJiafei One more cool thing: inference speed. MolmoAct2 is much faster than MolmoAct. Even cooler, with CUDA-acceleration, it runs almost twice as faster via HTTP! Checkout HuggingFace for the accelerated version

Jie Wang's profile picture
Jie Wang4 months ago

@DJiafei We will bring you structured testing & YAM Arm results next, stay tuned! Full clips:

Jie Wang's profile picture
Jie Wang4 months ago

@DJiafei Acknowledgement: Thanks @DJiafei for the warm support + hosting the policy server at AI2. Thanks @tianyurobot and @LongLeRobot for the helpful discussion, and thanks @allen_ai for the great open-cooking models!

Sanskar Pandey's profile picture
Sanskar Pandey4 months ago

@allen_ai @GRASPlab Now imagine RL'ing on the experiential data!!

Agentpilled's profile picture
Agentpilled4 months ago

@allen_ai @GRASPlab How does MolmoAct2 trigger that self-recovery behavior without explicit prompting? Does the reasoning loop draw from Allen AI's training data or emerge purely at inference time?

Jie Wang's profile picture
Jie Wang4 months ago

@allen_ai @GRASPlab My hypothesis is: 1. There is reasoning loop that will guide model to steer it 2. Model is visual shortcuting the “placement” behavior, which lift cup up coincidentally I need more evals and intermediate reasoning output to confirm that

Related Videos