Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Excited to release τ0-WM: an open-source unified video-action world model for robotic manipulation. It's a 5B-parameter robotic foundation model trained on 27.3K hours of real-robot teleoperation, UMI-style demonstrations, and egocentric interaction videos.

66,164 Aufrufe • vor 4 Monaten •via X (Twitter)

11 Kommentare

Profilbild von Jianlan Luo
Jianlan Luovor 4 Monaten

Most robot models directly map observations to actions. τ0-WM instead learns a shared predictive representation for acting and imagining: • generate executable robot actions • predict future visual outcomes • evaluate action-conditioned task progress Policy + world model in one framework.

Profilbild von Jianlan Luo
Jianlan Luovor 4 Monaten

τ0-WM is pretrained on a heterogeneous 27.3K-hour corpus: • 17.8K hrs real-robot teleoperation • 6.5K hrs UMI-style demonstrations • 3.0K hrs egocentric human interaction videos The resulting policy can solve a range of complex robotic manipulation tasks.

Profilbild von Jianlan Luo
Jianlan Luovor 4 Monaten

At the core of τ0-WM is a Video Action Model that jointly predicts future visual latents and action chunks from multi-view observations, language, and robot state. This enables action-conditioned simulation and test-time computation: imagine futures, evaluate progress, refine actions, then act.

Profilbild von Jianlan Luo
Jianlan Luovor 4 Monaten

To me, this is also part of a larger direction: Action Foundation Model → World Model → Learning While Deploying → Fleet Data Flywheel Pretraining gives robots broad priors. Deployment gives real feedback. The system should keep improving from both.

Profilbild von Jianlan Luo
Jianlan Luovor 4 Monaten

Project page: Code and model: Paper:

Profilbild von Dominique Paul
Dominique Paulvor 4 Monaten

Did you consider benchmarking against Mimic video?

Profilbild von Magnus
Magnusvor 4 Monaten

Which data type (teleop, umi, egocentric) would you prioritize scaling the most for future training runs?

Profilbild von 提前
提前vor 4 Monaten

If the ACVS can determine the quality of proposed actions from the VAM, why do we not use it to update the policy itself?

Profilbild von Nurvai - The Data Layer for Physical AI
Nurvai - The Data Layer for Physical AIvor 4 Monaten

Interesting how the boundary between policy and world model keeps disappearing. Once a model can both act and predict consequences, test time reasoning becomes a lot more natural. Do you think this unified approach eventually replaces standalone VLAs?

Profilbild von Olivia Liu
Olivia Liuvor 4 Monaten

Hi Dr. Luo — I'm a UPenn CS graduate following your work in robot learning at BAIR and Google X. I have a collaboration opportunity that could be a great fit with your expertise. Would you be open to a quick 10-minute chat? I'd love to share more and hear your thoughts. Thanks!

Profilbild von Tradeye
Tradeyevor 4 Monaten

Congrats on releasing tO-WM! We’re building the high-quality narrated egocentric data these models need at Tradeye: We capture first-person POV footage from licensed HVAC, plumbing, and electrical crews on real occupied job sites under our California C20 contractor license. This creates messy, long-horizon, decision-heavy data that complements teleop and UMI-style datasets. Happy to explore how our dataset could support future models like tO-WM.

Ähnliche Videos