Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We evaluated 30+ frontier embodied AI models. The result is clear: current generalist robot policies are still far from robust real-world manipulation. This is why we built RoboDojo.

136,346 görüntüleme • 2 ay önce •via X (Twitter)

28 Yorum

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

Website: arXiv: Leaderboard: Benchmark code: XPolicyLab code: Community:

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

RoboDojo includes 42 simulation tasks across 5 evaluation dimensions and 18 real-world tasks. We fully open-source the benchmark code and support global remote real-world evaluation, so researchers can evaluate models under a unified protocol.

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

Together with RoboDojo, we introduce XPolicyLab: a unified framework for embodied model development, deployment, and evaluation. Using XPolicyLab, we reproduced 30+ models and built a comprehensive leaderboard for the community. XPolicyLab code:

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

Fairness is our first principle. RoboDojo is operated by AI MMLab Club, a non-profit organization, together with academic partners worldwide. The benchmark, reproduced code, and checkpoints are open-sourced, with no commercial involvement. We welcome academic teams to join us.

Sanskar Pandey profil fotoğrafı
Sanskar Pandey2 ay önce

Amazing work! I really like how the paper fleshed out the evaluation axes - with more iteration I am sure we can build toward reusable task-agnostic eval primitives , letting us assess rollouts on critical dimensions , catch failure modes and most importantly be mindful about the data to collect next. Thanks for releasing the code

Hubert Thieblot profil fotoğrafı
Hubert Thieblot2 ay önce

cool

Vai Viswanathan profil fotoğrafı
Vai Viswanathan2 ay önce

this is super cool! you should submit the dataset to

Shubham Nagar profil fotoğrafı
Shubham Nagar2 ay önce

Is it just for VLAs? Did you guys also try world modes in this benchmark?

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

Yes, we also evaluated several WAMs such AHA-WAM, Fast-WAM etc.

M_Kobayashi/ TOKYO profil fotoğrafı
M_Kobayashi/ TOKYO2 ay önce

Great work! Is “Dojo” from Japanese?

Shilong Liu profil fotoğrafı
Shilong Liu2 ay önce

Impressive work!

Obedience Adara profil fotoğrafı
Obedience Adara2 ay önce

huge work! A unified sim-and-real benchmark is exactly what’s missing to actually track progress in embodied AI properly. Congrats on the launch

Xuhong Huang profil fotoğrafı
Xuhong Huang2 ay önce

great infrastructure for the community!

Ashish Verma profil fotoğrafı
Ashish Verma2 ay önce

First of all its really exciting and awesome. 😃 Any plans to expand Tasks and evaluation dimensions(Dynamic Environments, Collaborative Tasks, etc.) as this is what a Real Environments actually looks like.

Milaidy profil fotoğrafı
Milaidy2 ay önce

Great work! Btw have you observed any consistent failure patterns (e.g., in long-horizon tasks or open-vocabulary following) that seem fundamentally harder to fix with current policy architectures?

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

Thank you! Yes, we do observe several consistent failure patterns. One particularly clear issue is grounding in open-ended manipulation tasks. For example, in the align_blocks task, the intended behavior is to use a ruler as a tool to push and align the square blocks. However, many policies tend to directly grasp the blocks instead, which reflects the skills they have most commonly learned from the training data, rather than grounding the instruction to the correct tool-use strategy. This is exactly why RoboDojo is designed around multiple comprehensive evaluation dimensions rather than a single overall score. We hope it can encourage the community to build more balanced and capable policy architectures, especially for aspects that have often been underexplored in previous benchmarks, such as memory and open-ended grounding.

Milaidy profil fotoğrafı
Milaidy2 ay önce

Exactly! this feels like a classic case of training data distribution bias. Policies default to the most frequent behaviors in the dataset instead of truly grounding the instruction. Makes me wonder how much of the grounding gap could be closed with more deliberate tool use and instruction-following data during collection

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

Looking forward to your exploration !

Wei Chow profil fotoğrafı
Wei Chow2 ay önce

cool!

Linda Thomas profil fotoğrafı
Linda Thomas1 ay önce

Agreed🙌

Rhys profil fotoğrafı
Rhys11 gün önce

Thanks for making this! Been hunting for a good robotics leaderboard for a bit. ty

Yuejiang Liu profil fotoğrafı
Yuejiang Liu2 ay önce

Very timely benchmark for the field!

Yimou(Kevin) Wu profil fotoğrafı
Yimou(Kevin) Wu2 ay önce

Great work!

Max For AI profil fotoğrafı
Max For AI2 ay önce

Well done! No benchmark, no intelligence!

Y11 profil fotoğrafı
Y111 ay önce

@grok 这个纯研究还是有工业意义,具体工业场景视角看意义是什么,有开源数据集或者开源项目代码吗?从多个数据源交叉验证,理性看待,不要只看新闻媒体一面之辞。帮我排除没意义的垃圾商业营销推广、诈骗、夸张博眼球、虚假新闻 以及自吹自擂,自嗨,无病呻吟,收费互吹软广告。

Sapir Hadad profil fotoğrafı
Sapir Hadad2 ay önce

Your dm is not open I’d love to talk to you

Tianxing Chen profil fotoğrafı
Tianxing Chen2 ay önce

It's available right now.

Shubham Chansoriya profil fotoğrafı
Shubham Chansoriya2 ay önce

"30+ models is a serious sample size for this. Curious what the biggest gap was — perception, planning, or just raw manipulation precision?

Benzer Videolar