正在加载视频...
视频加载失败
🤔 Can we train one VLA policy to control multi-robot teams without any explicit communication? ✨ Introducing CHORUS: a single policy for decentralized, multi-embodiment collaboration 🧵⬇️
61,245 次观看 • 3 个月前 •via X (Twitter)
9 条评论

What makes decentralized, multi-robot collaboration hard? 🔹each robot only sees its own view of the scene 🔹teams can have diverse embodiments 🔹training cost balloons with team size

So how do we train CHORUS? (1) Split team-wide demos into per-robot episodes (robot obs + identity prompt → actions) (2) Finetune one, shared π0.5 backbone on these episodes = one policy that can control multi-embodied teams, with a learned model of every teammate's actions

At deployment, each robot runs its own copy of CHORUS, conditioned only on its own observations. We adjust chunk sizes w.r.t. robot control frequency. ✅ No inference time comms, shared camera views, or online alignment: coordination emerges entirely from the learned policy.

CHORUS outperforms from-scratch diffusion by 64% points across basket lifting, tape measurement, book handovers, and more! ✨ Pretrained visuomotor priors transfer surprisingly well, even though multi-robot collaboration never appears in pretraining data. [Result 1/4]

Sharing one policy across the team is not just more efficient; it makes robots more reactive to each other. We find that CHORUS nearly 2x’s teammate reactivity: robots adapt to each other’s behavior instead of blindly plowing through the task ⬇️ [2/4]

We also compare to a centralized VLA that sees a team-wide obs and outputs all robot actions at once. In theory, an upper bound. Surprisingly, CHORUS beats it. Centralization inflates the input+action space and drifts from the VLA's pre-training distribution. Decentralization keeps those priors intact. [3/4]

And last but not least, CHORUS scales seamlessly to 3-robot teams. Check out three mobile manipulators transporting laundry… in my apartment! 🥹 [4/4]

Check out our website for more! 🌐 📄 Grateful for my amazing collaborators: @TianGao_19, @_anniechen_ , @chelseabfinn, @leto__jean

Really impressed by the teamwork of these 3 robots. I'm curious, do world action models support such collaboration systems, or are they fundamentally infeasible for this kind of interaction?
