Loading video...

Video Failed to Load

Go Home

I've been working on deformable object manipulation since my PhD. It was totally a nightmare years ago and my PhD advisor was telling me not to work on it for my own good. Today, at ByteDance Seed, we are dropping GR-RL, a new VLA+RL system that manages long-horizon precise...

110,653 views • 10 months ago •via X (Twitter)

45 Comments

Xiao Ma's profile picture
Xiao Ma10 months ago

Why “shoelace threading” matters 🤔 This task is probably one of the most challenging household robotics tasks in terms of precision: 💥 Soft-body chaos – laces deform every frame 💥 Millimeter precision – 1–2 mm slip = total failure 💥 Long-horizon manipulation – hundreds of steps where slightest drift compounds. 💥 Requires recovery – humans retry; IL models freeze It's the perfect torture test for VLA precision. Even our previous model, GR-3 — already trained on massive robot trajectories + human teleop demos + public image-text corpora — failed to get reliable policies here. So we found the bottleneck.

Xiao Ma's profile picture
Xiao Ma10 months ago

Two killers of imitation learning (IL): (1) Human demos are NOT optimal Humans hesitate, retry, fix mistakes mid-trajectory. IL blindly copies ALL of it — including the bad parts. (2) Training vs Deployment Misalignment VLA models output actions. To prevent jitter, robots execute post-processed versions (RHC, smoothing, ensembling). Your “predicted action” ≠ your “executed action.” On millimeter tasks, this mismatch is fatal. So IL alone was a dead end.

Xiao Ma's profile picture
Xiao Ma10 months ago

The Idea: If imitation is broken, then: Let the robot learn from its own experience. GR-RL = ⭐️ Offline RL (data filtering) ⭐️ Symmetry augmentation ⭐️ Online closed-loop Real-World Reinforcement Learning All on top of a single VLA foundation model. Just RGB, proprioception, and language instructions.

Xiao Ma's profile picture
Xiao Ma10 months ago

Offline Stage — Filter the human flaws We train a Critic Transformer via distributional RL: ⭐️ Detects “value drops” when the operator hesitates or messes up ⭐️ Slices every trajectory into high-value vs low-value segments ⭐️ Retains only the cleanest expert behavior Effect: GR-3 45.7% → 61.6% success just from removing “bad micro-behaviors.”

Xiao Ma's profile picture
Xiao Ma10 months ago

Morphological Symmetry Augmentation Our bi-manual robot is left–right symmetric. So we mirror EVERYTHING: 🚀 RGB 🚀 Proprioception 🚀 Actions 🚀 Language Instructions Data size doubles. Spatial reasoning robustness skyrockets. → 72.7% success.

Xiao Ma's profile picture
Xiao Ma10 months ago

Online Stage — Real-World Steering RL Now the robot learns ON THE PHYSICAL PLATFORM. But direct exploration in joint space causes dangerous jitter and is inefficient — you need millimeter accuracy. So GR-RL explores in the latent noise space: A tiny 51.5M-param noise-predictor nudges the flow model’s latent action distribution toward high-value regions. We also introduce a dual-buffer mechanism: 🚀 Off-policy buffer (old rollouts) for the critic warm-up 🚀 On-policy buffer (last 2 checkpoints) for stable improvement 🚀 1:1 sampling stabilizes RL and prevents catastrophic drift. Just 150 online rollouts → +10% performance boost.

Xiao Ma's profile picture
Xiao Ma10 months ago

The Result: Final performance: 83.3% success over continuous shoelace threading. The surprising part? GR-RL learns to: 🔥 retry when the lace slips 🔥 reposition the lace when the initial pose is bad 🔥 "self-correct" mid-task instead of freezing This is the behavior you EXPECT from embodied intelligence — not just mimicry, but competence.

Xiao Ma's profile picture
Xiao Ma10 months ago

GR-RL proves something: IL is inherently limited, and we can do things previously thought impossible by purely visuo-motor control simply making it RL. The future direction is clear: distill RL-enhanced behavior back into the foundation VLA, forming a self-improving, ever-growing loop. Kudos to the wonderful team! @yunfeili98, @simanbulk, @yu_cui123, Zhongren Cui, Zhigang Han, @litchiyoyo, @KT892790499, @YuxiaoLiu4, Hao Niu, Wanli Peng, Jingchao Qiao, Zeyu Ren, Haixin Shi, @ZhiSu22, Jiawen Tian, Yuyang Xiao, Shenyu Zhang, Liwei Zheng

Perry Jia's profile picture
Perry Jia10 months ago

Wow! I love how it naturally pulls the shoe closer so that it can complete the task easier. Awesome work!

Xiao Ma's profile picture
Xiao Ma10 months ago

haha, thanks Perry! btw congrats on your recent release on ACT-1 and Memo! Great work! I really like the design and it's amazing how much it has achieved with pure UMI data.

🄱🄻🄸🅃🅉 🔸's profile picture
🄱🄻🄸🅃🅉 🔸10 months ago

This is brilliant Ma

Adam's profile picture
Adam10 months ago

incredible work, can't wait to see more in the future.

Daniel Domingos's profile picture
Daniel Domingos10 months ago

Congrats!

MAX ONBOARDER ⭕'s profile picture
MAX ONBOARDER ⭕10 months ago

@Scobleizer This is amazing

Bingyi Kang's profile picture
Bingyi Kang10 months ago

Great team! Great Work! Can not wait to see your future project! Congrats @yunfeili98 @yusufma555

Salman's profile picture
Salman10 months ago

This is incredible work. Kudos for seeing this work through for so many years, seems like it’s paying off!

Robert's profile picture
Robert10 months ago

Can you implement an aspect of leaning against the object that you're working on? This is what humans do when we tie our shoe, we don't float in space. This means that you don't have to have accurate control, you know that if you rest your finger against the shoe, then the string will be in the right place.

Xiao Ma's profile picture
Xiao Ma10 months ago

Hey Robert, this is a perfect point. Our model actually sometimes does it. During the offline data collection stage, most of the human demonstrators love to "float" because it is easier for them to control the robot when they became very good at it. However, this causes the issue you mentioned. But there were indeed some samples that the demonstrators love to lean on the shoes. As a result, our VLA model will demonstrate multi behavior modes, including the one you mentioned. But it is just less often, due to the initial data distribution was dominated by the other behavior mode.

Robert's profile picture
Robert10 months ago

Awesome, will this methodology ultimately require the robot to be trained on every task that it carries out or do you have some requirement for a "general" intelligence that can carry our arbitrary tasks? I know this is easier said than done, but interested in your opinion as I have no experience with robotics.

Yu Xiang's profile picture
Yu Xiang10 months ago

This task seems very difficult. Can the policy deal with different shoes?

Xiao Ma's profile picture
Xiao Ma10 months ago

thanks! Currently we tested on three different colored shoes. However, some other shoe designs are extremely difficult to thread, even for human demonstrators. There is much more we can do in the future :)

sam's profile picture
sam10 months ago

very impressive :)

Porters Reserve's profile picture
Porters Reserve10 months ago

😁 this why a professor is just person who memorized a book to tell you to read the same book. Only real world grit and experimentation is the path to success. Great Job now level up send your gear to the reserve to see what really breaks.

Rxct Ftvy's profile picture
Rxct Ftvy10 months ago

Really really cool

Rxct Ftvy's profile picture
Rxct Ftvy10 months ago

Very cool

Offend_no_one's profile picture
Offend_no_one10 months ago

Wow, Yusuf, that’s incredible progress from ByteDance!

Younggyo Seo's profile picture
Younggyo Seo10 months ago

Cool, congrats on the release!!

Xiao Ma's profile picture
Xiao Ma10 months ago

Thank you Younggyo!

Griffin's profile picture
Griffin10 months ago

Incredibly impressive, well done!

Fanf666's profile picture
Fanf66610 months ago

Can you clarify 1) the number of examples to reach your quoted accuracy 2) how many distinct tasks can the model remember and how does it select them

Fanf666's profile picture
Fanf66610 months ago

Here is prior art shoelace tyeing What matter is to advance few shot, max task number and length.

Roméo's profile picture
Roméo10 months ago

@grok analyse et explique ce qui se passe dans cette vidéo

czr's profile picture
czr10 months ago

Cool !

尚月's profile picture
尚月10 months ago

还有哪些常见的日常活动对机器人而言与系鞋带平级或者更复杂?

Dan Wu's profile picture
Dan Wu10 months ago

So sick! Curious what cameras are mounted onto the end effectors?

Xiao Ma's profile picture
Xiao Ma10 months ago

thanks! We have discussed our hardware setups in the section 4 of our paper. It is a realsense D405 camera.

Muhammad Umar 🇺🇸's profile picture
Muhammad Umar 🇺🇸10 months ago

It's just a matter of time we'll solve dexterity. This is just unreal.

Commentary Yi He's profile picture
Commentary Yi He9 months ago

Amazing project! 👏 Message me directly

Commentary Yi He's profile picture
Commentary Yi He9 months ago

Awesome 👏 Message me directly

Reza Sayar's profile picture
Reza Sayar10 months ago

this is awesome! 🔥 congrats! 👏🏼 i noticed even though you had RGB-D cameras, you only kept the RGB for training and inference. did you find that depth data won't help here much?

Laughless's profile picture
Laughless10 months ago

This is more impressive than dancing or kung fu. Do you think it could tie a neck tie without choking a person? That's one of my benchmarks for when robots are ready for home use.

Raul Verdusco's profile picture
Raul Verdusco10 months ago

Incredible progress — deformable object manipulation has always been one of the toughest challenges in robotics, and seeing real-world RL push success rates this high is remarkable. Huge step forward for dexterous autonomy. 🤖🧠 #Robotics #ReinforcementLearning

Golding's profile picture
Golding10 months ago

This is pretty cool man. I'd love to learn more about your research and ByteDance Seed. Can you open your DM's? let's chat!

Farhan Azad Shuvra's profile picture
Farhan Azad Shuvra10 months ago

Impressive Stuff!!

X Æ A-12's profile picture
X Æ A-1210 months ago

you are the best 😍🧡💙💚 !

Related Videos

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 views • 2 months ago