
Shuo Yang
@ShuoYangAIR • 5,151 subscribers
CTO & Co-founder @mondorobotics | Ex Tesla | CMU PhD | Ex DJI
Videos

Mondo's Beni ( may look cute, but there is a serious learning stack inside. It can perform flips, navigate autonomously, and track moving subjects. The secret is reinforcement learning: Beni practices at massive scale in simulation, optimizing its controller through millions of virtual trials before entering the real world. Here is a thread on some of the research ideas and tools used by Beni: 🧵
Shuo Yang31,100 views • 1 month ago

We’re excited to share DiT4DiT, an end-to-end Video-Action Model for robot learning that unifies a video Diffusion Transformer and an action Diffusion Transformer in a single cascaded framework. By leveraging the rich spatiotemporal and physical dynamics learned through video generation, rather than static image-text priors, DiT4DiT achieves state-of-the-art results on LIBERO (98.6%) and RoboCasa GR1 (50.8%) with far less training data, delivering over 10× better sample efficiency and up to 7× faster convergence. Real-world deployment on a humanoid robot further shows robust generalization. We believe this is a step toward making video generation a powerful backbone for robot policy learning. This work builds upon the brilliant foundations laid by Nvidia's GR00T and Cosmos. Project: Paper: Code: Coming soon. In the meantime, you can ask your coding agent to reproduce the method based on GR00T/Cosmos.
Shuo Yang31,872 views • 6 months ago

DiT4DiT is now open source! As the first humanoid-deployable Video-Action Model built on a world model, DiT4DiT continues to surprise us. In our paper last month, we showed its strong data efficiency. Now, with only slight modifications, it enables real-time whole-body autonomous pick-and-place. Paper: Code: Website:
Shuo Yang14,534 views • 4 months ago
No more content to load