Loading video...
Video Failed to Load
For roboticists, one challenge towers above the others: there isn’t enough data. To accelerate the deployment of intelligent robots in the real world, MIT CSAIL’s "LucidSim" uses genAI & physics engines to create diverse & realistic virtual training grounds for robots. W/o any real-world data, their robot achieved expert-level... show more
82,915 views • 1 year ago •via X (Twitter)
9 Comments

ChatGPT comes up w/diverse descriptions of the environment, which are then transformed into realistic pictures by a ControlNet. Guidance from the physics engine ensures that images reflect correct collision geometry and real-world physics.

To make short, 140 millisecond videos that serve as visual "experiences" for the robot, the scientists hacked together a trick called "Dreams In Motion (DIM)" using a mix of image magic. This trick made LucidSim 7x times faster by moving pixels of a single generated image according to the head movement of the robot, & piece together a short, multi-frame video.

LucidSim used AI-generated images to train a robot dog to do parkour — w/o real-world data.

Having the dog learn from its own actions is essential. The team put LucidSim against the alternative, where an expert teacher provides substantially more training data. Robots learn from expert data only succeeded 15% of the time — and even quadrupling the amount of expert training data barely moved the needle. When robots collected their own training data through LucidSim, the story changed dramatically: Just doubling the dataset size catapulted success rates to 88%.

LucidSim outperforms domain randomization, a go-to method from 2017 that produces diverse data, but lacks realism. LucidSim addresses both diversity & realism problems, helping a robot recognize and navigate obstacles in real environments.

LucidSim is paving the way toward a new generation of intelligent machines that elevate & amplify human effort. The researchers envision these machines could safely and autonomously learn how to navigate our complex world w/o ever setting foot in it. The team’s immediate plan is to scale up to a human-sized robot and teach it to manipulate objects on the move, using purely synthesized images.

Authors: Alan Yu (@alany1_), Ge Yang (@EpisodeYang), Ran Choi, Yajvan Ravan, John Leonard (@jleonardmit), & Phillip Isola (@phillip_isola) Paper: Full video: Website:

There's no enough data ? You mean current algorithms are dumb as a brick, children don't need to look at 1 million pictures of cats to learn what a cat is.

The idea of walking your beta robot...seems like there is a song in that....
