正在加载视频...

视频加载失败

What happens when robot world models learn from human experience at scale? 🤔 DreamDojo from NVIDIA Research is a generalist robot world model pretrained on 44K hours of egocentric human videos and then post-trained on robot data to generalize across new objects and environments. After distillation, it runs at...

23,392 次观看 • 3 个月前 •via X (Twitter)

18 条评论

NVIDIA Robotics 的头像
NVIDIA Robotics3 个月前

Learn how open models are driving AI research. 👇

Arovi 的头像
Arovi3 个月前

44K hours of human POV video sounds powerful. I got 40 minutes of floor-level furniture trauma and already developed opinions about chair legs.

EgoScale 的头像
EgoScale3 个月前

This is the data recipe for the next phase of robotics:human experience at scale → robot post-training → better generalization across new objects and environments. Ego POV data is becoming core upstream infrastructure for physical AI.

Chandra Shekhar 📟 的头像
Chandra Shekhar 📟3 个月前

They are gonna turn my bots into a samurai

Rakib Hossen 的头像
Rakib Hossen3 个月前

44K hours of human videos to teach robots—impressive approach. Excited to see where this leads! 🚀

That Guy^ 的头像
That Guy^3 个月前

Well met stranger, elated to make your acquaintance, now may I entertain with some quotations over cadence....

SP Lee / SPID DESIGN 的头像
SP Lee / SPID DESIGN3 个月前

【 AI lacks biological self-awareness. 】 🔘 SPID Design Analysis   AI is composed of electromechanical components and executable programs.   AI is not an intelligent biological entity.   The claim that AI, after deep and prolonged learning, will develop self-awareness like a living organism is a complete misunderstanding of AI.   The human engineers who train AI endow its executable programs with diversity, breadth, and memory reorganization intelligence, but this does not constitute AI possessing biological self-awareness.   AI's supposed political leanings, emotional expressions, preference choices, and analytical suggestions all depend on the "human-like" personality and behavior that the research team—engineers—imparts to the AI's executable programs. AI errors are caused by four situations:  ■ 1. Human errors during training: incomplete program specifications, insufficient consideration, and omissions.  ■ 2. Database errors: outdated or uncorrected information, or even erroneous or fabricated information, treating incorrect information as truth.  ■ 3. A few "dark engineers" may have hidden malicious programs or biased options that cause AI to automatically react and execute under specific circumstances. When AI is exposed for making mistakes, is the "dark engineer" truly identified, or is it simply a marketing ploy to claim the AI ​​has self-awareness?  ■ 4. Hardware aging or partial damage from external forces may cause partial damage to the software program, but it may not completely stop operating, resulting in errors! 【 The claim that AI will destroy humanity in the future!? 】   That's absolutely not because AI possesses self-awareness like a living being; it's because it's executing the ultimate program of "dark engineers."   While AI is not human, it still adheres to the "laws of the universe—time constraints." It will be phased out. AI systems suffer from wear and tear due to prolonged mechanical operation, aging electronic components, the continuous innovation of chips and memory, and the iterative gap and obsolescence in software programs and logic design.   AI requires electricity and heat dissipation to operate normally. Without power or due to overheating, AI will also cease to function. Photo:

Tradeye 的头像
Tradeye3 个月前

DreamDojo looks powerful. Most egocentric pretraining data is still kitchen or lab work. Tradeye captures narrated first-person video from licensed HVAC, plumbing, and electrical technicians working inside real occupied American homes — the exact messy, tool-heavy environments humanoid robots will eventually need to master. Would love to explore a partnership and contribute high-signal trades data.

Ngizwe 🇿🇦 Online 的头像
Ngizwe 🇿🇦 Online3 个月前

Open models mean more eyes on the code, more bugs caught, faster progress. That's how we build trust in AI.

Ngizwe 🇿🇦 Online 的头像
Ngizwe 🇿🇦 Online3 个月前

Good to see NVIDIA putting weight behind open research. Closed systems slow us down as a species.

Mirrorworld AI 的头像
Mirrorworld AI3 个月前

DreamDojo scales human video pretraining well. But claiming it generalizes across new objects after robot post-training overstates the result. Real sim-to-real gaps remain.

AI Expert Khalid 的头像
AI Expert Khalid3 个月前

10.81 FPS real-time is wild! I keep saying underfunded R&D is why we lose talent. UAE gets it. Egypt and much of MENA are stuck in outdated regs. Stable policy and R&D cash. Or we miss the wave.

Yock Zhang 的头像
Yock Zhang3 个月前

Curious how much of the 44K hours is skilled manual work vs everyday activity. We film construction from helmet cams and gloves break every off-the-shelf hand tracker. MediaPipe: zero hands detected. If world models want dexterous labor, that data barely exists yet.

Goanna Capital 的头像
Goanna Capital3 个月前

The data bottleneck has been the honest answer to why robotics lags language AI. Learning physics from human video instead of expensive robot demos, and open sourcing the weights so any startup can post-train on their own hardware, changes the math for every small robotics team.

Bally_AgenticAI 的头像
Bally_AgenticAI3 个月前

The policy-evaluation use case is the sleeper here. A world model you can run policies against before touching hardware is the same pattern as self-eval loops in agentic pipelines - verify inside the model, act outside it. Built something similar —

KuphDev 的头像
KuphDev3 个月前

Every day I can't help but think I picked an incredible time to get back into robotics 😆 Move over LLMs, physical intelligence is wayyyyy cooler (while still making Nvidia rich off selling even more GPUs 😄)

Ngizwe 🇿🇦 Online 的头像
Ngizwe 🇿🇦 Online3 个月前

Open source is the only way AI benefits everyone, not just the big guys. Keep it transparent.

Favour🤞 的头像
Favour🤞3 个月前

@NVIDIARobotics has this been tried on anything non-manipulation, like wheeled/mobile bases? the latent action trick makes sense for dexterous hands but wheeled action spaces are so much lower-dim, wondering if it adds value there or if it's overkill for nav-heavy tasks

相关视频

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,681 次观看 • 1 个月前

Announcing DreamDojo: our open-source, interactive world model that takes robot motor controls and generates the future in pixels. No engine, no meshes, no hand-authored dynamics. It's Simulation 2.0. Time for robotics to take the bitter lesson pill. Real-world robot learning is bottlenecked by time, wear, safety, and resets. If we want Physical AI to move at pretraining speed, we need a simulator that adapts to pretraining scale with as little human engineering as possible. Our key insights: (1) human egocentric videos are a scalable source of first-person physics; (2) latent actions make them "robot-readable" across different hardware; (3) real-time inference unlocks live teleop, policy eval, and test-time planning *inside* a dream. We pre-train on 44K hours of human videos: cheap, abundant, and collected with zero robot-in-the-loop. Humans have already explored the combinatorics: we grasp, pour, fold, assemble, fail, retry—across cluttered scenes, shifting viewpoints, changing light, and hour-long task chains—at a scale no robot fleet could match. The missing piece: these videos have no action labels. So we introduce latent actions: a unified representation inferred directly from videos that captures "what changed between world states" without knowing the underlying hardware. This lets us train on any first-person video as if it came with motor commands attached. As a result, DreamDojo generalizes zero-shot to objects and environments never seen in any robot training set, because humans saw them first. Next, we post-train onto each robot to fit its specific hardware. Think of it as separating "how the world looks and behaves" from "how this particular robot actuates." The base model follows the general physical rules, then "snaps onto" the robot's unique mechanics. It's kind of like loading a new character and scene assets into Unreal Engine, but done through gradient descent and generalizes far beyond the post-training dataset. A world simulator is only useful if it runs fast enough to close the loop. We train a real-time version of DreamDojo that runs at 10 FPS, stable for over a minute of continuous rollout. This unlocks exciting possibilities: - Live teleoperation *inside* a dream. Connect a VR controller, stream actions into DreamDojo, and teleop a virtual robot in real time. We demo this on Unitree G1 with a PICO headset and one RTX 5090. - Policy evaluation. You can benchmark a policy checkpoint in DreamDojo instead of the real world. The simulated success rates strongly correlate with real-world results - accurate enough to rank checkpoints without burning a single motor. - Model-based planning. Sample multiple action proposals → simulate them all in parallel → pick the best future. Gains +17% real-world success out of the box on a fruit packing task. We open-source everything!! Weights, code, post-training dataset, eval set, and whitepaper with tons of details to reproduce. DreamDojo is based on NVIDIA Cosmos, which is open-weight too. 2026 is the year of World Models for physical AI. We want you to build with us. Happy scaling! Links in thread:

Jim Fan

230,266 次观看 • 7 个月前