Loading video...

Video Failed to Load

Go Home

🤔 Can we train one VLA policy to control multi-robot teams without any explicit communication? ✨ Introducing CHORUS: a single policy for decentralized, multi-embodiment collaboration 🧵⬇️

61,245 views • 3 months ago •via X (Twitter)

9 Comments

Ria Doshi's profile picture
Ria Doshi3 months ago

What makes decentralized, multi-robot collaboration hard? 🔹each robot only sees its own view of the scene 🔹teams can have diverse embodiments 🔹training cost balloons with team size

Ria Doshi's profile picture
Ria Doshi3 months ago

So how do we train CHORUS? (1) Split team-wide demos into per-robot episodes (robot obs + identity prompt → actions) (2) Finetune one, shared π0.5 backbone on these episodes = one policy that can control multi-embodied teams, with a learned model of every teammate's actions

Ria Doshi's profile picture
Ria Doshi3 months ago

At deployment, each robot runs its own copy of CHORUS, conditioned only on its own observations. We adjust chunk sizes w.r.t. robot control frequency. ✅ No inference time comms, shared camera views, or online alignment: coordination emerges entirely from the learned policy.

Ria Doshi's profile picture
Ria Doshi3 months ago

CHORUS outperforms from-scratch diffusion by 64% points across basket lifting, tape measurement, book handovers, and more! ✨ Pretrained visuomotor priors transfer surprisingly well, even though multi-robot collaboration never appears in pretraining data. [Result 1/4]

Ria Doshi's profile picture
Ria Doshi3 months ago

Sharing one policy across the team is not just more efficient; it makes robots more reactive to each other. We find that CHORUS nearly 2x’s teammate reactivity: robots adapt to each other’s behavior instead of blindly plowing through the task ⬇️ [2/4]

Ria Doshi's profile picture
Ria Doshi3 months ago

We also compare to a centralized VLA that sees a team-wide obs and outputs all robot actions at once. In theory, an upper bound. Surprisingly, CHORUS beats it. Centralization inflates the input+action space and drifts from the VLA's pre-training distribution. Decentralization keeps those priors intact. [3/4]

Ria Doshi's profile picture
Ria Doshi3 months ago

And last but not least, CHORUS scales seamlessly to 3-robot teams. Check out three mobile manipulators transporting laundry… in my apartment! 🥹 [4/4]

Ria Doshi's profile picture
Ria Doshi3 months ago

Check out our website for more! 🌐 📄 Grateful for my amazing collaborators: @TianGao_19, @_anniechen_ , @chelseabfinn, @leto__jean

clankr's profile picture
clankr3 months ago

Really impressed by the teamwork of these 3 robots. I'm curious, do world action models support such collaboration systems, or are they fundamentally infeasible for this kind of interaction?

Related Videos