Video yükleniyor...
Video Yüklenemedi
Robotics models often struggle outside controlled environments. Ours is built to work in real ones. Today we're launching MolmoAct 2, which can assist with a host of chores & lab tasks, plus the MolmoAct 2-Bimanual YAM dataset—the largest open robotics dataset of its kind. 🧵
416,222 görüntüleme • 4 ay önce •via X (Twitter)
20 Yorum

MolmoAct 2 builds on MolmoAct, our first Action Reasoning Model (ARM). Like its predecessor, MolmoAct 2 reasons about the world in 3D before taking actions. It now runs up to 37x faster & handles two-armed tasks out of the box without per-task fine-tuning.

MolmoAct 2 is easier to guide in real-world settings. Prompt it in natural language, and it responds well to different phrasings thanks to language re-annotation across the robotics training data. MolmoAct 2-Think goes further with adaptive depth perception tokens for stronger spatial reasoning. 👇

Some of MolmoAct 2’s gains come from an adapter linking the VLM & action expert. It’s built on Molmo 2-ER, an embodied-reasoning Molmo 2 variant trained on millions of examples in detection, spatial reasoning, & more. Molmo 2-ER beats GPT-5 on embodied-reasoning benchmarks.

We're already testing MolmoAct 2 outside controlled setups. We've deployed it in our office café to make popcorn & drinks, even as people move around it. It handles practical tasks like wiping surfaces, lifting trays, & folding towels.

We've also piloted MolmoAct 2 with research partners including a Stanford Medicine team using it for hands-on tasks in CRISPR gene-editing work. The model moves samples, uses lab equipment, & recovers from small mistakes during long experiments.

In real-world zero-shot tests on a Franka robot, MolmoAct 2 outperforms strong proprietary models on every task we evaluated, including placing an apple on a plate, moving a pipette to a tray, and a longer multi-step task with several objects.

MolmoAct 2 also sets a new state of the art on LIBERO, a standard benchmark that tests how well robots learn manipulation tasks across varied scenes & objects—improving meaningfully over the original MolmoAct.

We retained @cortexairobot to run a third-party real-world fine-tuning benchmark. Across trials on a broad suite of tabletop, in-the-wild, and mobile tasks, MolmoAct 2 outperformed systems including OpenVLA-OFT, π0.5, X-VLA, & Cosmos Policy.

To lower the barrier to getting started, we're sharing an affordable reference hardware setup: two YAM arms, overhead & close-up cameras, an extendable mount, and a tabletop workspace for bimanual manipulation.

Robotics models are often closed. MolmoAct 2 isn't. We're releasing model weights, an updated VLA architecture, a fully open action tokenizer, & the MolmoAct 2-Bimanual YAM dataset—700+ hours of tabletop manipulation demonstrations.

MolmoAct 2 is a foundation built for others to study, deploy, & customize. 📝 Learn more in our blog: 🤖 Models: 📊 Training dataset:

Nice work @hq_fang @DJiafei, very strong control and ER results!

Cool work, I was curious in Section 6.8, the reported inference speeds are "amortized"? Afaict, the whole chunk is generated together so even though it is reported to generate actions at 55.79Hz (17.92ms), the VLA can only generate actions at 179.2ms? Or am I missing something?

Allen AI dropped Molmo Act 2 cooking in real cafés with zero drama?! This is the most hyped I’ve been about robotics all year, who’s ready for robots that actually pull up and deliver?! 🔥

Super awesome research!!! Love it

This sounds interesting! Where can I read in depth how it works?

Check out our blog!

Great to see this from a non-profit organisation! You released datasets & weights that no other company did. The one thing unique thing about this is - your models predicts in 3D depth Spatial even before acting while other models do in 2D. Really excited to try it on my LeRobot arm. #Ai2 #Robotic #PhysicalAI

Dunia sudah tergantikan

Finally, a dataset large enough to prove that my robot's bimanual coordination is still just two arms having a very expensive disagreement over a laundry basket. Huge win for the open-stack ecosystem!
