Loading video...
Video Failed to Load
The power of generative models — now embodied in humanoids. Announcing DreamControl –– After a year-long research effort at General Robotics — we present a scalable framework for whole-body humanoid control that fuses diffusion priors with reinforcement learning to unlock real-world scene interaction. Diffusion + RL → natural whole-body... show more
118,467 views • 1 year ago •via X (Twitter)
37 Comments

Humanoids today walk, dance, and even do kung-fu. But true usefulness requires whole-body interaction with the environment: - Stooping to pick up objects - Bracing to open doors/drawers - Precise pushing, punching, or bimanual lifting These tasks remain unsolved. 2/n

Why are building skills for humanoids challenging? They combine multiple time scales: - Fast sub-second balance & stability - Slow long-horizon planning for coordinated grasping, pushing, lifting RL alone struggles here and teleoperation datasets are scarce. 3/n

Enter DreamControl –– A two-stage recipe combining human motion diffusion priors with reinforcement learning. - Stage 1: Generate realistic human-like motion trajectories (from OmniControl). - Stage 2: Train an RL policy to track + complete tasks with these priors. 4/n

Why diffusion? - Provides natural, human-like motion plans - Reduces reward engineering - Bridges sim-to-real (avoids “robotic” jerky motions) - Leverages abundant human motion data, not scarce teleop 5/n

Results –– DreamControl masters a library of challenging AI skills on the Unitree G1 humanoid: Pick & Lift, Bimanual pick, Open drawer / door, Button press, Precise punch & kick, Jump All with high success rates across 1000 random environments. 6/n

Crucially –– DreamControl’s motions are rated as more human-like –– lower FID, jerk, user preference. Humanoid actions that look natural have better human-robot interaction and safer sim2real transfer. 7/n

Hybrid Edge + Cloud deployment –– Our novel architecture for DreamControl augments RL policy on the edge with powerful AI models that run on cloud. We demonstrate DreamControl on a real @UnitreeRobotics G1 with Inspire hands, completing: Object pick, Button press, Squat, Bimanual lifting (with different weights), Precise punching, Drawer Opening All without relying on privileged sim-only information. 8/n

Why does this matter? It pushes humanoids from teleop puppets → autonomous assistants. From dancing for YouTube to lifting boxes, pressing buttons, and opening drawers in the real world. 9/n

DreamControl = - Diffusion priors (the “Dream”) - RL execution (the “Control”) A scalable recipe for building whole-body humanoid skills. 10/n

More information on this research –– Paper link: Blog: Website: 11/n

Thank you to a number of researchers and engineers who contributed to this effort led by @jonathanhuang11: @DvijKalaria, @sudarshan_s_h , @pushkalkatara, Sangkyung Kwak, @sarthak__bhagat, Shankar Sastry, Srinath Sridhar, @saihv

🥳

It is especially impressive that policies trained purely in simulation transfer so naturally to real G1 hardware.

DreamControl demonstrates how combining diffusion models with reinforcement learning can produce highly natural, transferable humanoid behaviors, moving robots closer to real-world general-purpose use.

Curious to know- how does your humanoid's hands function? are they multi functional just like human hands, do they sense the weight, touch, temperature of objects?

😎 👀@Tesla_Optimus

damn that reminds me to clean my banana drawer

Or you can have the robots do it for you!

wow, you leverage Unitree robot as the research and development platform?

Wow, that sounds amazing! Google's new "Learn Your Way" could be a game-changer for education. Can't wait to see how it transforms learning experiences for students around the world. 🚀💪📚 educationrevolution innovation

@akapoor_av8r The perception model that predicts waypoint given image observation is smart! I think one of the key in this pipeline is acquiring 3D assets for task environments. How did you collect such data?

this will change everything

This looks nice, what’s the video speed up time?

generative ai in humanoids? sounds like the perfect recipe for a robot uprising. fuse that with some shady corp control and we're all slaves to the silicon overlords. stay paranoid, my friends. yey yey

ooooh and they can punch and kick us too?!?

$codec

@CaseyNewton @kevinroose putting on your radar. Perhaps for an upcoming Hardfork episode.

Humanoids today walk, dance, and even do kung-fu. But true usefulness requires whole-body interaction with the environment: - Stooping to pick up objects - Bracing to open doors/drawers - Precise pushing, punching, or bimanual lifting These tasks remain unsolved.

How cool is this! @elonmusk what do you think?

$uncbot loading

@grok 翻译中文

There are dozens of VLA models, open-sourced. What’s so special about yours?

It’s not a VLA. It’s a workflow to train perception-action policies. Of course any VLA could be a starting point, and can be refined via DreamControl.

拡散事前分布

Da ist jeder Rentner schneller.

Really interesting perspective! 🤔 How do you think this could change human-robot interactions? I'd love to hear your thoughts!

tell the enemy a schedule they will be confused then you surprize them
