正在加载视频...

视频加载失败

I've been working on deformable object manipulation since my PhD. It was totally a nightmare years ago and my PhD advisor was telling me not to work on it for my own good. Today, at ByteDance Seed, we are dropping GR-RL, a new VLA+RL system that manages long-horizon precise...

110,653 次观看 • 10 个月前 •via X (Twitter)

45 条评论

Xiao Ma 的头像
Xiao Ma10 个月前

Why “shoelace threading” matters 🤔 This task is probably one of the most challenging household robotics tasks in terms of precision: 💥 Soft-body chaos – laces deform every frame 💥 Millimeter precision – 1–2 mm slip = total failure 💥 Long-horizon manipulation – hundreds of steps where slightest drift compounds. 💥 Requires recovery – humans retry; IL models freeze It's the perfect torture test for VLA precision. Even our previous model, GR-3 — already trained on massive robot trajectories + human teleop demos + public image-text corpora — failed to get reliable policies here. So we found the bottleneck.

Xiao Ma 的头像
Xiao Ma10 个月前

Two killers of imitation learning (IL): (1) Human demos are NOT optimal Humans hesitate, retry, fix mistakes mid-trajectory. IL blindly copies ALL of it — including the bad parts. (2) Training vs Deployment Misalignment VLA models output actions. To prevent jitter, robots execute post-processed versions (RHC, smoothing, ensembling). Your “predicted action” ≠ your “executed action.” On millimeter tasks, this mismatch is fatal. So IL alone was a dead end.

Xiao Ma 的头像
Xiao Ma10 个月前

The Idea: If imitation is broken, then: Let the robot learn from its own experience. GR-RL = ⭐️ Offline RL (data filtering) ⭐️ Symmetry augmentation ⭐️ Online closed-loop Real-World Reinforcement Learning All on top of a single VLA foundation model. Just RGB, proprioception, and language instructions.

Xiao Ma 的头像
Xiao Ma10 个月前

Offline Stage — Filter the human flaws We train a Critic Transformer via distributional RL: ⭐️ Detects “value drops” when the operator hesitates or messes up ⭐️ Slices every trajectory into high-value vs low-value segments ⭐️ Retains only the cleanest expert behavior Effect: GR-3 45.7% → 61.6% success just from removing “bad micro-behaviors.”

Xiao Ma 的头像
Xiao Ma10 个月前

Morphological Symmetry Augmentation Our bi-manual robot is left–right symmetric. So we mirror EVERYTHING: 🚀 RGB 🚀 Proprioception 🚀 Actions 🚀 Language Instructions Data size doubles. Spatial reasoning robustness skyrockets. → 72.7% success.

Xiao Ma 的头像
Xiao Ma10 个月前

Online Stage — Real-World Steering RL Now the robot learns ON THE PHYSICAL PLATFORM. But direct exploration in joint space causes dangerous jitter and is inefficient — you need millimeter accuracy. So GR-RL explores in the latent noise space: A tiny 51.5M-param noise-predictor nudges the flow model’s latent action distribution toward high-value regions. We also introduce a dual-buffer mechanism: 🚀 Off-policy buffer (old rollouts) for the critic warm-up 🚀 On-policy buffer (last 2 checkpoints) for stable improvement 🚀 1:1 sampling stabilizes RL and prevents catastrophic drift. Just 150 online rollouts → +10% performance boost.

Xiao Ma 的头像
Xiao Ma10 个月前

The Result: Final performance: 83.3% success over continuous shoelace threading. The surprising part? GR-RL learns to: 🔥 retry when the lace slips 🔥 reposition the lace when the initial pose is bad 🔥 "self-correct" mid-task instead of freezing This is the behavior you EXPECT from embodied intelligence — not just mimicry, but competence.

Xiao Ma 的头像
Xiao Ma10 个月前

GR-RL proves something: IL is inherently limited, and we can do things previously thought impossible by purely visuo-motor control simply making it RL. The future direction is clear: distill RL-enhanced behavior back into the foundation VLA, forming a self-improving, ever-growing loop. Kudos to the wonderful team! @yunfeili98, @simanbulk, @yu_cui123, Zhongren Cui, Zhigang Han, @litchiyoyo, @KT892790499, @YuxiaoLiu4, Hao Niu, Wanli Peng, Jingchao Qiao, Zeyu Ren, Haixin Shi, @ZhiSu22, Jiawen Tian, Yuyang Xiao, Shenyu Zhang, Liwei Zheng

Perry Jia 的头像
Perry Jia10 个月前

Wow! I love how it naturally pulls the shoe closer so that it can complete the task easier. Awesome work!

Xiao Ma 的头像
Xiao Ma10 个月前

haha, thanks Perry! btw congrats on your recent release on ACT-1 and Memo! Great work! I really like the design and it's amazing how much it has achieved with pure UMI data.

🄱🄻🄸🅃🅉 🔸 的头像
🄱🄻🄸🅃🅉 🔸10 个月前

This is brilliant Ma

Adam 的头像
Adam10 个月前

incredible work, can't wait to see more in the future.

Daniel Domingos 的头像
Daniel Domingos10 个月前

Congrats!

MAX ONBOARDER ⭕ 的头像
MAX ONBOARDER ⭕10 个月前

@Scobleizer This is amazing

Bingyi Kang 的头像
Bingyi Kang10 个月前

Great team! Great Work! Can not wait to see your future project! Congrats @yunfeili98 @yusufma555

Salman 的头像
Salman10 个月前

This is incredible work. Kudos for seeing this work through for so many years, seems like it’s paying off!

Robert 的头像
Robert10 个月前

Can you implement an aspect of leaning against the object that you're working on? This is what humans do when we tie our shoe, we don't float in space. This means that you don't have to have accurate control, you know that if you rest your finger against the shoe, then the string will be in the right place.

Xiao Ma 的头像
Xiao Ma10 个月前

Hey Robert, this is a perfect point. Our model actually sometimes does it. During the offline data collection stage, most of the human demonstrators love to "float" because it is easier for them to control the robot when they became very good at it. However, this causes the issue you mentioned. But there were indeed some samples that the demonstrators love to lean on the shoes. As a result, our VLA model will demonstrate multi behavior modes, including the one you mentioned. But it is just less often, due to the initial data distribution was dominated by the other behavior mode.

Robert 的头像
Robert10 个月前

Awesome, will this methodology ultimately require the robot to be trained on every task that it carries out or do you have some requirement for a "general" intelligence that can carry our arbitrary tasks? I know this is easier said than done, but interested in your opinion as I have no experience with robotics.

Yu Xiang 的头像
Yu Xiang10 个月前

This task seems very difficult. Can the policy deal with different shoes?

Xiao Ma 的头像
Xiao Ma10 个月前

thanks! Currently we tested on three different colored shoes. However, some other shoe designs are extremely difficult to thread, even for human demonstrators. There is much more we can do in the future :)

sam 的头像
sam10 个月前

very impressive :)

Porters Reserve 的头像
Porters Reserve10 个月前

😁 this why a professor is just person who memorized a book to tell you to read the same book. Only real world grit and experimentation is the path to success. Great Job now level up send your gear to the reserve to see what really breaks.

Rxct Ftvy 的头像
Rxct Ftvy10 个月前

Really really cool

Rxct Ftvy 的头像
Rxct Ftvy10 个月前

Very cool

Offend_no_one 的头像
Offend_no_one10 个月前

Wow, Yusuf, that’s incredible progress from ByteDance!

Younggyo Seo 的头像
Younggyo Seo10 个月前

Cool, congrats on the release!!

Xiao Ma 的头像
Xiao Ma10 个月前

Thank you Younggyo!

Griffin 的头像
Griffin10 个月前

Incredibly impressive, well done!

Fanf666 的头像
Fanf66610 个月前

Can you clarify 1) the number of examples to reach your quoted accuracy 2) how many distinct tasks can the model remember and how does it select them

Fanf666 的头像
Fanf66610 个月前

Here is prior art shoelace tyeing What matter is to advance few shot, max task number and length.

Roméo 的头像
Roméo10 个月前

@grok analyse et explique ce qui se passe dans cette vidéo

czr 的头像
czr10 个月前

Cool !

尚月 的头像
尚月10 个月前

还有哪些常见的日常活动对机器人而言与系鞋带平级或者更复杂?

Dan Wu 的头像
Dan Wu10 个月前

So sick! Curious what cameras are mounted onto the end effectors?

Xiao Ma 的头像
Xiao Ma10 个月前

thanks! We have discussed our hardware setups in the section 4 of our paper. It is a realsense D405 camera.

Muhammad Umar 🇺🇸 的头像
Muhammad Umar 🇺🇸10 个月前

It's just a matter of time we'll solve dexterity. This is just unreal.

Commentary Yi He 的头像
Commentary Yi He9 个月前

Amazing project! 👏 Message me directly

Commentary Yi He 的头像
Commentary Yi He9 个月前

Awesome 👏 Message me directly

Reza Sayar 的头像
Reza Sayar10 个月前

this is awesome! 🔥 congrats! 👏🏼 i noticed even though you had RGB-D cameras, you only kept the RGB for training and inference. did you find that depth data won't help here much?

Laughless 的头像
Laughless10 个月前

This is more impressive than dancing or kung fu. Do you think it could tie a neck tie without choking a person? That's one of my benchmarks for when robots are ready for home use.

Raul Verdusco 的头像
Raul Verdusco10 个月前

Incredible progress — deformable object manipulation has always been one of the toughest challenges in robotics, and seeing real-world RL push success rates this high is remarkable. Huge step forward for dexterous autonomy. 🤖🧠 #Robotics #ReinforcementLearning

Golding 的头像
Golding10 个月前

This is pretty cool man. I'd love to learn more about your research and ByteDance Seed. Can you open your DM's? let's chat!

Farhan Azad Shuvra 的头像
Farhan Azad Shuvra10 个月前

Impressive Stuff!!

X Æ A-12 的头像
X Æ A-1210 个月前

you are the best 😍🧡💙💚 !

相关视频

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 次观看 • 2 个月前