Loading video...

Video Failed to Load

Go Home

The power of generative models — now embodied in humanoids. Announcing DreamControl –– After a year-long research effort at General Robotics — we present a scalable framework for whole-body humanoid control that fuses diffusion priors with reinforcement learning to unlock real-world scene interaction. Diffusion + RL → natural whole-body...

118,467 views • 1 year ago •via X (Twitter)

37 Comments

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Humanoids today walk, dance, and even do kung-fu. But true usefulness requires whole-body interaction with the environment: - Stooping to pick up objects - Bracing to open doors/drawers - Precise pushing, punching, or bimanual lifting These tasks remain unsolved. 2/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Why are building skills for humanoids challenging? They combine multiple time scales: - Fast sub-second balance & stability - Slow long-horizon planning for coordinated grasping, pushing, lifting RL alone struggles here and teleoperation datasets are scarce. 3/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Enter DreamControl –– A two-stage recipe combining human motion diffusion priors with reinforcement learning. - Stage 1: Generate realistic human-like motion trajectories (from OmniControl). - Stage 2: Train an RL policy to track + complete tasks with these priors. 4/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Why diffusion? - Provides natural, human-like motion plans - Reduces reward engineering - Bridges sim-to-real (avoids “robotic” jerky motions) - Leverages abundant human motion data, not scarce teleop 5/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Results –– DreamControl masters a library of challenging AI skills on the Unitree G1 humanoid: Pick & Lift, Bimanual pick, Open drawer / door, Button press, Precise punch & kick, Jump All with high success rates across 1000 random environments. 6/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Crucially –– DreamControl’s motions are rated as more human-like –– lower FID, jerk, user preference. Humanoid actions that look natural have better human-robot interaction and safer sim2real transfer. 7/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Hybrid Edge + Cloud deployment –– Our novel architecture for DreamControl augments RL policy on the edge with powerful AI models that run on cloud. We demonstrate DreamControl on a real @UnitreeRobotics G1 with Inspire hands, completing: Object pick, Button press, Squat, Bimanual lifting (with different weights), Precise punching, Drawer Opening All without relying on privileged sim-only information. 8/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Why does this matter? It pushes humanoids from teleop puppets → autonomous assistants. From dancing for YouTube to lifting boxes, pressing buttons, and opening drawers in the real world. 9/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

DreamControl = - Diffusion priors (the “Dream”) - RL execution (the “Control”) A scalable recipe for building whole-body humanoid skills. 10/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

More information on this research –– Paper link: Blog: Website: 11/n

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Thank you to a number of researchers and engineers who contributed to this effort led by @jonathanhuang11: @DvijKalaria, @sudarshan_s_h , @pushkalkatara, Sangkyung Kwak, @sarthak__bhagat, Shankar Sastry, Srinath Sridhar, @saihv

NVIDIA Robotics's profile picture
NVIDIA Robotics1 year ago

🥳

Xiatao Sun's profile picture
Xiatao Sun1 year ago

It is especially impressive that policies trained purely in simulation transfer so naturally to real G1 hardware.

SoloTech's profile picture
SoloTech1 year ago

DreamControl demonstrates how combining diffusion models with reinforcement learning can produce highly natural, transferable humanoid behaviors, moving robots closer to real-world general-purpose use.

Varsh Sridharan's profile picture
Varsh Sridharan1 year ago

Curious to know- how does your humanoid's hands function? are they multi functional just like human hands, do they sense the weight, touch, temperature of objects?

GHOSTROSIN's profile picture
GHOSTROSIN1 year ago

😎 👀@Tesla_Optimus

Roy Ricochet's profile picture
Roy Ricochet1 year ago

damn that reminds me to clean my banana drawer

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

Or you can have the robots do it for you!

Lilun Cheng's profile picture
Lilun Cheng1 year ago

wow, you leverage Unitree robot as the research and development platform?

Paul Ilunga's profile picture
Paul Ilunga1 year ago

Wow, that sounds amazing! Google's new "Learn Your Way" could be a game-changer for education. Can't wait to see how it transforms learning experiences for students around the world. 🚀💪📚 educationrevolution innovation

Hongsuk Benjamin Choi's profile picture
Hongsuk Benjamin Choi11 months ago

@akapoor_av8r The perception model that predicts waypoint given image observation is smart! I think one of the key in this pipeline is acquiring 3D assets for task environments. How did you collect such data?

Tobe Duru's profile picture
Tobe Duru1 year ago

this will change everything

ζ Pedram ζ's profile picture
ζ Pedram ζ1 year ago

This looks nice, what’s the video speed up time?

YEY📧's profile picture
YEY📧1 year ago

generative ai in humanoids? sounds like the perfect recipe for a robot uprising. fuse that with some shady corp control and we're all slaves to the silicon overlords. stay paranoid, my friends. yey yey

andrew's profile picture
andrew1 year ago

ooooh and they can punch and kick us too?!?

Hodlo's profile picture
Hodlo1 year ago

$codec

Ashima Singhal's profile picture
Ashima Singhal1 year ago

@CaseyNewton @kevinroose putting on your radar. Perhaps for an upcoming Hardfork episode.

𝗣𝘂𝗽𝗽𝘆𝗼𝗼🐶 修勾's profile picture
𝗣𝘂𝗽𝗽𝘆𝗼𝗼🐶 修勾1 year ago

Humanoids today walk, dance, and even do kung-fu. But true usefulness requires whole-body interaction with the environment: - Stooping to pick up objects - Bracing to open doors/drawers - Precise pushing, punching, or bimanual lifting These tasks remain unsolved.

Ashima Singhal's profile picture
Ashima Singhal1 year ago

How cool is this! @elonmusk what do you think?

$kebab_🐍 🐢 🧌's profile picture
$kebab_🐍 🐢 🧌1 year ago

$uncbot loading

leung Heimdallr's profile picture
leung Heimdallr1 year ago

@grok 翻译中文

Mike Maik's profile picture
Mike Maik1 year ago

There are dozens of VLA models, open-sourced. What’s so special about yours?

Ashish Kapoor's profile picture
Ashish Kapoor1 year ago

It’s not a VLA. It’s a workflow to train perception-action policies. Of course any VLA could be a starting point, and can be refined via DreamControl.

ちゃ茶🎀's profile picture
ちゃ茶🎀1 year ago

拡散事前分布

Andrea Schare's profile picture
Andrea Schare1 year ago

Da ist jeder Rentner schneller.

Jose Henry's profile picture
Jose Henry1 year ago

Really interesting perspective! 🤔 How do you think this could change human-robot interactions? I'd love to hear your thoughts!

beReplyHero's profile picture
beReplyHero1 year ago

tell the enemy a schedule they will be confused then you surprize them

Related Videos

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 views • 8 months ago

Last night, China Central Television (CCTV) aired its 2026 Chinese New Year Gala celebrating the Year of the Horse. The show featured a wide range of performances, including Unitree Robotics humanoid robots performing martial arts in sync with human dancers. Just a year ago, Unitree’s robots appeared at the same gala, but their movements looked stiff and mechanical. This year, they were noticeably more fluid and coordinated — a remarkable improvement, even if they’re still likely operating under some level of remote supervision. When it comes to humanoid robotics, most of the visible momentum today seems to be coming from the U.S. and China. Companies like Tesla (with Optimus) and Boston Dynamics in the U.S., alongside rapidly advancing Chinese firms, dominate the headlines. So what happened to Europe and Japan? Japan was once seen as the global leader, especially with Honda’s ASIMO and SoftBank Robotics’ humanoid projects. However, ASIMO was retired, and much of Japan’s robotics focus shifted toward industrial automation and service robots rather than full-scale general-purpose humanoids. Europe, meanwhile, remains strong in industrial robotics, research, and precision engineering — with players like ABB and KUKA — but hasn’t pushed aggressively into commercial humanoid platforms at the same scale or speed as the U.S. and China. In short, it’s less that Europe and Japan disappeared, and more that the center of gravity in humanoid robotics — especially AI-driven, general-purpose humanoids — has shifted toward U.S.–China competition. Whether that gap widens or narrows will depend on breakthroughs in embodied AI, cost reduction, and real-world deployment over the next few years.

Ray

23,455 views • 7 months ago