Yunzhu Li's banner
Yunzhu Li's profile picture

Yunzhu Li

@YunzhuLiYZ10,576 subscribers

Co-Founder of SceniX (now part of @theworldlabs) | Assistant Professor @Columbia @ColumbiaCompSci | Former postdoc @Stanford @StanfordSVL | PhD @MIT_CSAIL

Shorts

Super excited about Hydra-0 from Hongyu Li and team! The key idea is to use flow as a shared visual interface across embodiments/objects for controllable video generation, allowing a single generalist world model to learn from human, handheld-gripper, and robot interaction data. My favorite result is the video below: start from a real video of a human doing the task (left), extract the desired object flow, and condition the model on that flow (right). The model then hallucinates a plausible robot motion that could produce the same object motion. Very cool glimpse of how a generalist world model can bridge human demonstrations and robot control.

Super excited about Hydra-0 from Hongyu Li and team! The key idea is to use flow as a shared visual interface across embodiments/objects for controllable video generation, allowing a single generalist world model to learn from human, handheld-gripper, and robot interaction data. My favorite result is the video below: start from a real video of a human doing the task (left), extract the desired object flow, and condition the model on that flow (right). The model then hallucinates a plausible robot motion that could produce the same object motion. Very cool glimpse of how a generalist world model can bridge human demonstrations and robot control.

10,509 次观看

You can actually interact with the world simulator directly in the browser. 🤖 Here is a quick screen recording (8x speed) of me playing with it: real-time action-conditioned video prediction across rigid objects, deformable objects, rope, and object piles. Try it yourself (no install required): Huge kudos to my student Yixuan Wang for making the interactive demo happen!

You can actually interact with the world simulator directly in the browser. 🤖 Here is a quick screen recording (8x speed) of me playing with it: real-time action-conditioned video prediction across rigid objects, deformable objects, rope, and object piles. Try it yourself (no install required): Huge kudos to my student Yixuan Wang for making the interactive demo happen!

41,513 次观看

Excited to share a few presentations, demos, and workshop talks from our group and collaborators at #ICRA2026! We will present recent work on real-to-sim-to-real robot policy evaluation, model-based planning with learned dynamics, and multi-modal manipulation. We will also have a joint live demo between SceniX and Analog Devices, Inc. on real-to-sim-to-real cable manipulation at the ICRA exhibition. This is a small teaser of what we have been building, with more to come soon! If you are at ICRA, please stop by the sessions or the demo booth. Happy to chat about robot learning, simulation, world models, and sim-to-real!

Excited to share a few presentations, demos, and workshop talks from our group and collaborators at #ICRA2026! We will present recent work on real-to-sim-to-real robot policy evaluation, model-based planning with learned dynamics, and multi-modal manipulation. We will also have a joint live demo between SceniX and Analog Devices, Inc. on real-to-sim-to-real cable manipulation at the ICRA exhibition. This is a small teaser of what we have been building, with more to come soon! If you are at ICRA, please stop by the sessions or the demo booth. Happy to chat about robot learning, simulation, world models, and sim-to-real!

11,102 次观看

I want to call out one of our most important references: SIMPLER ( by Xuanlin Li (Simon), Kyle Hsu, Jiayuan Gu, Jiajun Wu, Hao Su, Quan Vuong, Ted Xiao, and colleagues, which laid the foundation for using simulation for policy evaluation through a systematic study of appearance and dynamics alignment and metrics for measuring sim–real correlation. (It took many nights of Kaifeng Zhang and Shuo Sha's grinding for that correlation to appear — and once it did, extending to new tasks worked like a charm!)

I want to call out one of our most important references: SIMPLER ( by Xuanlin Li (Simon), Kyle Hsu, Jiayuan Gu, Jiajun Wu, Hao Su, Quan Vuong, Ted Xiao, and colleagues, which laid the foundation for using simulation for policy evaluation through a systematic study of appearance and dynamics alignment and metrics for measuring sim–real correlation. (It took many nights of Kaifeng Zhang and Shuo Sha's grinding for that correlation to appear — and once it did, extending to new tasks worked like a charm!)

11,870 次观看

🚀 Excited to share our #ICLR2025 work on planning with neural dynamics models! While our lab has developed diverse neural dynamics models for manipulating rigid, deformable, and granular objects, having the model alone doesn’t solve the problem—planning with it remains a challenge. 💡 Enter BaB-ND, led by Keyi and Jiangwei! We propose a scalable, GPU-accelerated branch-and-bound algorithm, inspired by neural network verification, to enable effective planning for diverse objects modeled with neural dynamics. 🔗 Project page (open-source + detailed docs!): 🎥 Watch the video to see T being pushed around obstacles, and check out Keyi’s thread for more details!

🚀 Excited to share our #ICLR2025 work on planning with neural dynamics models! While our lab has developed diverse neural dynamics models for manipulating rigid, deformable, and granular objects, having the model alone doesn’t solve the problem—planning with it remains a challenge. 💡 Enter BaB-ND, led by Keyi and Jiangwei! We propose a scalable, GPU-accelerated branch-and-bound algorithm, inspired by neural network verification, to enable effective planning for diverse objects modeled with neural dynamics. 🔗 Project page (open-source + detailed docs!): 🎥 Watch the video to see T being pushed around obstacles, and check out Keyi’s thread for more details!

10,561 次观看

Videos

YunzhuLiYZ's profile picture

I was really impressed by the UMI gripper (Cheng Chi et al.), but a key limitation is that **force-related data wasn’t captured**: humans feel haptic feedback through the mechanical springs, but the robot couldn’t leverage that info, limiting the data’s value for fine-grained manipulation tasks. Led by my amazing students Yolanda Zhu and Binghao Huang, we designed a **portable visuo-tactile gripper** by integrating our dense, flexible tactile arrays with the UMI gripper to enable large-scale in-the-wild data collection. 🔗 We demonstrate **cross-modal representation learning** and **downstream policy learning** on tasks requiring in-hand state estimation (e.g., test tube reorientation) and fine-grained force sensing (e.g., pipette fluid transfer). Key takeaways: - Our flexible tactile arrays store the rich haptic information humans perceive as dense tactile signals. - Portability and robustness are key for in-the-wild data collection; our portable gripper is compact, lightweight, and durable. - Touch provides precise, robust measurements of in-hand object pose, invariant to lighting and viewpoint. - Cross-modal pretraining on large-scale in-the-wild data significantly improves policy robustness and sample efficiency (as shown many times before — and verified again here!). Also check out our previous investigations of dense, flexible tactile grids for understanding human-robot-environment interactions: - Dense tactile glove (Nature ’19): - 3D-ViTac (CoRL ’24):

Yunzhu Li

13,188 次观看 • 1 年前

没有更多内容可加载