Video wird geladen...
Video konnte nicht geladen werden
Excited to release τ0-WM: an open-source unified video-action world model for robotic manipulation. It's a 5B-parameter robotic foundation model trained on 27.3K hours of real-robot teleoperation, UMI-style demonstrations, and egocentric interaction videos.
66,164 Aufrufe • vor 4 Monaten •via X (Twitter)
11 Kommentare

Most robot models directly map observations to actions. τ0-WM instead learns a shared predictive representation for acting and imagining: • generate executable robot actions • predict future visual outcomes • evaluate action-conditioned task progress Policy + world model in one framework.

τ0-WM is pretrained on a heterogeneous 27.3K-hour corpus: • 17.8K hrs real-robot teleoperation • 6.5K hrs UMI-style demonstrations • 3.0K hrs egocentric human interaction videos The resulting policy can solve a range of complex robotic manipulation tasks.

At the core of τ0-WM is a Video Action Model that jointly predicts future visual latents and action chunks from multi-view observations, language, and robot state. This enables action-conditioned simulation and test-time computation: imagine futures, evaluate progress, refine actions, then act.

To me, this is also part of a larger direction: Action Foundation Model → World Model → Learning While Deploying → Fleet Data Flywheel Pretraining gives robots broad priors. Deployment gives real feedback. The system should keep improving from both.

Project page: Code and model: Paper:

Did you consider benchmarking against Mimic video?

Which data type (teleop, umi, egocentric) would you prioritize scaling the most for future training runs?

If the ACVS can determine the quality of proposed actions from the VAM, why do we not use it to update the policy itself?

Interesting how the boundary between policy and world model keeps disappearing. Once a model can both act and predict consequences, test time reasoning becomes a lot more natural. Do you think this unified approach eventually replaces standalone VLAs?

Hi Dr. Luo — I'm a UPenn CS graduate following your work in robot learning at BAIR and Google X. I have a collaboration opportunity that could be a great fit with your expertise. Would you be open to a quick 10-minute chat? I'd love to share more and hear your thoughts. Thanks!

Congrats on releasing tO-WM! We’re building the high-quality narrated egocentric data these models need at Tradeye: We capture first-person POV footage from licensed HVAC, plumbing, and electrical crews on real occupied job sites under our California C20 contractor license. This creates messy, long-horizon, decision-heavy data that complements teleop and UMI-style datasets. Happy to explore how our dataset could support future models like tO-WM.
