Загрузка видео...

Не удалось загрузить видео

На главную

🤔 Can we train one VLA policy to control multi-robot teams without any explicit communication? ✨ Introducing CHORUS: a single policy for decentralized, multi-embodiment collaboration 🧵⬇️

61,245 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 9

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

What makes decentralized, multi-robot collaboration hard? 🔹each robot only sees its own view of the scene 🔹teams can have diverse embodiments 🔹training cost balloons with team size

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

So how do we train CHORUS? (1) Split team-wide demos into per-robot episodes (robot obs + identity prompt → actions) (2) Finetune one, shared π0.5 backbone on these episodes = one policy that can control multi-embodied teams, with a learned model of every teammate's actions

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

At deployment, each robot runs its own copy of CHORUS, conditioned only on its own observations. We adjust chunk sizes w.r.t. robot control frequency. ✅ No inference time comms, shared camera views, or online alignment: coordination emerges entirely from the learned policy.

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

CHORUS outperforms from-scratch diffusion by 64% points across basket lifting, tape measurement, book handovers, and more! ✨ Pretrained visuomotor priors transfer surprisingly well, even though multi-robot collaboration never appears in pretraining data. [Result 1/4]

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

Sharing one policy across the team is not just more efficient; it makes robots more reactive to each other. We find that CHORUS nearly 2x’s teammate reactivity: robots adapt to each other’s behavior instead of blindly plowing through the task ⬇️ [2/4]

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

We also compare to a centralized VLA that sees a team-wide obs and outputs all robot actions at once. In theory, an upper bound. Surprisingly, CHORUS beats it. Centralization inflates the input+action space and drifts from the VLA's pre-training distribution. Decentralization keeps those priors intact. [3/4]

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

And last but not least, CHORUS scales seamlessly to 3-robot teams. Check out three mobile manipulators transporting laundry… in my apartment! 🥹 [4/4]

Фото профиля Ria Doshi
Ria Doshi3 месяцев назад

Check out our website for more! 🌐 📄 Grateful for my amazing collaborators: @TianGao_19, @_anniechen_ , @chelseabfinn, @leto__jean

Фото профиля clankr
clankr3 месяцев назад

Really impressed by the teamwork of these 3 robots. I'm curious, do world action models support such collaboration systems, or are they fundamentally infeasible for this kind of interaction?

Похожие видео