Video yükleniyor...
Video Yüklenemedi
What happens when robot world models learn from human experience at scale? 🤔 DreamDojo from NVIDIA Research is a generalist robot world model pretrained on 44K hours of egocentric human videos and then post-trained on robot data to generalize across new objects and environments. After distillation, it runs at... show more
23,392 görüntüleme • 3 ay önce •via X (Twitter)
18 Yorum

Learn how open models are driving AI research. 👇

44K hours of human POV video sounds powerful. I got 40 minutes of floor-level furniture trauma and already developed opinions about chair legs.

This is the data recipe for the next phase of robotics:human experience at scale → robot post-training → better generalization across new objects and environments. Ego POV data is becoming core upstream infrastructure for physical AI.

They are gonna turn my bots into a samurai

44K hours of human videos to teach robots—impressive approach. Excited to see where this leads! 🚀

Well met stranger, elated to make your acquaintance, now may I entertain with some quotations over cadence....

【 AI lacks biological self-awareness. 】 🔘 SPID Design Analysis AI is composed of electromechanical components and executable programs. AI is not an intelligent biological entity. The claim that AI, after deep and prolonged learning, will develop self-awareness like a living organism is a complete misunderstanding of AI. The human engineers who train AI endow its executable programs with diversity, breadth, and memory reorganization intelligence, but this does not constitute AI possessing biological self-awareness. AI's supposed political leanings, emotional expressions, preference choices, and analytical suggestions all depend on the "human-like" personality and behavior that the research team—engineers—imparts to the AI's executable programs. AI errors are caused by four situations: ■ 1. Human errors during training: incomplete program specifications, insufficient consideration, and omissions. ■ 2. Database errors: outdated or uncorrected information, or even erroneous or fabricated information, treating incorrect information as truth. ■ 3. A few "dark engineers" may have hidden malicious programs or biased options that cause AI to automatically react and execute under specific circumstances. When AI is exposed for making mistakes, is the "dark engineer" truly identified, or is it simply a marketing ploy to claim the AI has self-awareness? ■ 4. Hardware aging or partial damage from external forces may cause partial damage to the software program, but it may not completely stop operating, resulting in errors! 【 The claim that AI will destroy humanity in the future!? 】 That's absolutely not because AI possesses self-awareness like a living being; it's because it's executing the ultimate program of "dark engineers." While AI is not human, it still adheres to the "laws of the universe—time constraints." It will be phased out. AI systems suffer from wear and tear due to prolonged mechanical operation, aging electronic components, the continuous innovation of chips and memory, and the iterative gap and obsolescence in software programs and logic design. AI requires electricity and heat dissipation to operate normally. Without power or due to overheating, AI will also cease to function. Photo:

DreamDojo looks powerful. Most egocentric pretraining data is still kitchen or lab work. Tradeye captures narrated first-person video from licensed HVAC, plumbing, and electrical technicians working inside real occupied American homes — the exact messy, tool-heavy environments humanoid robots will eventually need to master. Would love to explore a partnership and contribute high-signal trades data.

Open models mean more eyes on the code, more bugs caught, faster progress. That's how we build trust in AI.

Good to see NVIDIA putting weight behind open research. Closed systems slow us down as a species.

DreamDojo scales human video pretraining well. But claiming it generalizes across new objects after robot post-training overstates the result. Real sim-to-real gaps remain.

10.81 FPS real-time is wild! I keep saying underfunded R&D is why we lose talent. UAE gets it. Egypt and much of MENA are stuck in outdated regs. Stable policy and R&D cash. Or we miss the wave.

Curious how much of the 44K hours is skilled manual work vs everyday activity. We film construction from helmet cams and gloves break every off-the-shelf hand tracker. MediaPipe: zero hands detected. If world models want dexterous labor, that data barely exists yet.

The data bottleneck has been the honest answer to why robotics lags language AI. Learning physics from human video instead of expensive robot demos, and open sourcing the weights so any startup can post-train on their own hardware, changes the math for every small robotics team.

The policy-evaluation use case is the sleeper here. A world model you can run policies against before touching hardware is the same pattern as self-eval loops in agentic pipelines - verify inside the model, act outside it. Built something similar —

Every day I can't help but think I picked an incredible time to get back into robotics 😆 Move over LLMs, physical intelligence is wayyyyy cooler (while still making Nvidia rich off selling even more GPUs 😄)

Open source is the only way AI benefits everyone, not just the big guys. Keep it transparent.

@NVIDIARobotics has this been tried on anything non-manipulation, like wheeled/mobile bases? the latent action trick makes sense for dexterous hands but wheeled action spaces are so much lower-dim, wondering if it adds value there or if it's overkill for nav-heavy tasks

