ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

Introducing ๐ƒ๐ซ๐ž๐š๐ฆ๐†๐ž๐ง! We got humanoid robots to perform totally new ๐‘ฃ๐‘’๐‘Ÿ๐‘๐‘  in new environments through video world models. We believe video world models will solve the data problem in robotics. Bringing the paradigm of scaling human hours to GPU hours. Quick ๐Ÿงต

117,574 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰ โ€ขvia X (Twitter)

11 ๆก่ฏ„่ฎบ

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

Currently, robot data scaling is done through human labor. Recent work showed some potential signs of robots doing useful things in unseen homes (i.e., open-world generalization), but this required taking the physical robots and collecting data in 100+ homes.

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

How about new ๐‘ฃ๐‘’๐‘Ÿ๐‘๐‘ ? When prompted, current robots stand still or perform tasks they were trained on (e.g. pick-and-place). There is currently no work in literature that can enable robots to perform new verbs outside of the teleoperation data for visuomotor robot policies.

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

We introduce ๐ƒ๐ซ๐ž๐š๐ฆ๐†๐ž๐ง, an embarrassingly simple 4-step pipeline that generates synthetic robot training data using video world models. (1) Fine-tune, (2) Prompt, (3) Extract actions, (4) Train. Simple as that.

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

Through DreamGen, we generate โ€œDreamsโ€ or ๐‘๐‘’๐‘ข๐‘Ÿ๐‘Ž๐‘™ ๐‘‡๐‘Ÿ๐‘Ž๐‘—๐‘’๐‘๐‘ก๐‘œ๐‘Ÿ๐‘–๐‘’๐‘  of 22 new verbs in 10 unseen environments, and train robots to perform these tasks "zero-shot".

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

DreamGen can also augment seen, contact-rich tasks that are hard to simulate (e.g. folding, scooping m&ms) for different robot systems (Franka & SO-100) and different robot policies (Diffusion Policy, ฯ€โ‚€, GR00T N1), all resulting in non-trivial gains.

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

We also introduce ๐˜‹๐‘Ÿ๐˜ฆ๐‘Ž๐˜ฎ๐บ๐˜ฆ๐‘› ๐ต๐˜ฆ๐‘›๐˜คโ„Ž, a video generative benchmark for robotics that shows a positive correlation with downstream robot policies, so that video model researchers can help enable robotics without actually having to set up their own physical robot systems.

Joel Jang ็š„ๅคดๅƒ
Joel Jang1 ๅนดๅ‰

This was an 8-month research effort, jointly led by @SeonghyeonYe @zy27962986 @szxiangjn, and advised by @scott_e_reed @yukez @DrJimFan , and with amazing collaborators from Nvidia GEAR Lab, Nvidia Cosmos Team, and UW. Check out more videos in the blog post & details in the paper! ๐ŸŒ Blog post: ๐Ÿ“ Paper: We will also be releasing the code & an API to try out DreamGen (prompt -> lerobot dataset) in the upcoming days! Do reach out if you are interested in joining our GR00T Dreams Team. We are scaling up the data, GPUs, and the team!

VistaShares ็š„ๅคดๅƒ
VistaShares1 ๅนดๅ‰

Discover the future of AI investing. AIS delivers exposure to the companies driving the next wave of innovationโ€”semiconductors, data centers, and AI applications. Explore the supercycle today.

Thomas Farfeleder ็š„ๅคดๅƒ
Thomas Farfeleder1 ๅนดๅ‰

@threadreaderapp unroll

JulianSaks ็š„ๅคดๅƒ
JulianSaks1 ๅนดๅ‰

So cool to see this! Congrats ๐Ÿ”ฅ

Raymond Yu ็š„ๅคดๅƒ
Raymond Yu1 ๅนดๅ‰

Congrats Joel! Doing big things๐Ÿ‘€

็›ธๅ…ณ่ง†้ข‘

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's ๐Ÿ”ฅ Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. โ†’ 1,000 hours pre-training โ†’ 20% normalised task performance โ†’ 10,000 hours โ†’ 28% โ†’ 100,000 hours โ†’ 45% โ†’ 1,000,000 hours โ†’ 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ โ™ป๏ธ Join the weekly robotics newsletter, and never miss any news โ†’

Lukas Ziegler

23,576 ๆฌก่ง‚็œ‹ โ€ข 23 ๅคฉๅ‰