Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I've been working on deformable object manipulation since my PhD. It was totally a nightmare years ago and my PhD advisor was telling me not to work on it for my own good. Today, at ByteDance Seed, we are dropping GR-RL, a new VLA+RL system that manages long-horizon precise...

110,653 görüntüleme • 10 ay önce •via X (Twitter)

45 Yorum

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Why “shoelace threading” matters 🤔 This task is probably one of the most challenging household robotics tasks in terms of precision: 💥 Soft-body chaos – laces deform every frame 💥 Millimeter precision – 1–2 mm slip = total failure 💥 Long-horizon manipulation – hundreds of steps where slightest drift compounds. 💥 Requires recovery – humans retry; IL models freeze It's the perfect torture test for VLA precision. Even our previous model, GR-3 — already trained on massive robot trajectories + human teleop demos + public image-text corpora — failed to get reliable policies here. So we found the bottleneck.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Two killers of imitation learning (IL): (1) Human demos are NOT optimal Humans hesitate, retry, fix mistakes mid-trajectory. IL blindly copies ALL of it — including the bad parts. (2) Training vs Deployment Misalignment VLA models output actions. To prevent jitter, robots execute post-processed versions (RHC, smoothing, ensembling). Your “predicted action” ≠ your “executed action.” On millimeter tasks, this mismatch is fatal. So IL alone was a dead end.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

The Idea: If imitation is broken, then: Let the robot learn from its own experience. GR-RL = ⭐️ Offline RL (data filtering) ⭐️ Symmetry augmentation ⭐️ Online closed-loop Real-World Reinforcement Learning All on top of a single VLA foundation model. Just RGB, proprioception, and language instructions.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Offline Stage — Filter the human flaws We train a Critic Transformer via distributional RL: ⭐️ Detects “value drops” when the operator hesitates or messes up ⭐️ Slices every trajectory into high-value vs low-value segments ⭐️ Retains only the cleanest expert behavior Effect: GR-3 45.7% → 61.6% success just from removing “bad micro-behaviors.”

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Morphological Symmetry Augmentation Our bi-manual robot is left–right symmetric. So we mirror EVERYTHING: 🚀 RGB 🚀 Proprioception 🚀 Actions 🚀 Language Instructions Data size doubles. Spatial reasoning robustness skyrockets. → 72.7% success.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Online Stage — Real-World Steering RL Now the robot learns ON THE PHYSICAL PLATFORM. But direct exploration in joint space causes dangerous jitter and is inefficient — you need millimeter accuracy. So GR-RL explores in the latent noise space: A tiny 51.5M-param noise-predictor nudges the flow model’s latent action distribution toward high-value regions. We also introduce a dual-buffer mechanism: 🚀 Off-policy buffer (old rollouts) for the critic warm-up 🚀 On-policy buffer (last 2 checkpoints) for stable improvement 🚀 1:1 sampling stabilizes RL and prevents catastrophic drift. Just 150 online rollouts → +10% performance boost.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

The Result: Final performance: 83.3% success over continuous shoelace threading. The surprising part? GR-RL learns to: 🔥 retry when the lace slips 🔥 reposition the lace when the initial pose is bad 🔥 "self-correct" mid-task instead of freezing This is the behavior you EXPECT from embodied intelligence — not just mimicry, but competence.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

GR-RL proves something: IL is inherently limited, and we can do things previously thought impossible by purely visuo-motor control simply making it RL. The future direction is clear: distill RL-enhanced behavior back into the foundation VLA, forming a self-improving, ever-growing loop. Kudos to the wonderful team! @yunfeili98, @simanbulk, @yu_cui123, Zhongren Cui, Zhigang Han, @litchiyoyo, @KT892790499, @YuxiaoLiu4, Hao Niu, Wanli Peng, Jingchao Qiao, Zeyu Ren, Haixin Shi, @ZhiSu22, Jiawen Tian, Yuyang Xiao, Shenyu Zhang, Liwei Zheng

Perry Jia profil fotoğrafı
Perry Jia10 ay önce

Wow! I love how it naturally pulls the shoe closer so that it can complete the task easier. Awesome work!

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

haha, thanks Perry! btw congrats on your recent release on ACT-1 and Memo! Great work! I really like the design and it's amazing how much it has achieved with pure UMI data.

🄱🄻🄸🅃🅉 🔸 profil fotoğrafı
🄱🄻🄸🅃🅉 🔸10 ay önce

This is brilliant Ma

Adam profil fotoğrafı
Adam10 ay önce

incredible work, can't wait to see more in the future.

Daniel Domingos profil fotoğrafı
Daniel Domingos10 ay önce

Congrats!

MAX ONBOARDER ⭕ profil fotoğrafı
MAX ONBOARDER ⭕10 ay önce

@Scobleizer This is amazing

Bingyi Kang profil fotoğrafı
Bingyi Kang10 ay önce

Great team! Great Work! Can not wait to see your future project! Congrats @yunfeili98 @yusufma555

Salman profil fotoğrafı
Salman10 ay önce

This is incredible work. Kudos for seeing this work through for so many years, seems like it’s paying off!

Robert profil fotoğrafı
Robert10 ay önce

Can you implement an aspect of leaning against the object that you're working on? This is what humans do when we tie our shoe, we don't float in space. This means that you don't have to have accurate control, you know that if you rest your finger against the shoe, then the string will be in the right place.

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Hey Robert, this is a perfect point. Our model actually sometimes does it. During the offline data collection stage, most of the human demonstrators love to "float" because it is easier for them to control the robot when they became very good at it. However, this causes the issue you mentioned. But there were indeed some samples that the demonstrators love to lean on the shoes. As a result, our VLA model will demonstrate multi behavior modes, including the one you mentioned. But it is just less often, due to the initial data distribution was dominated by the other behavior mode.

Robert profil fotoğrafı
Robert10 ay önce

Awesome, will this methodology ultimately require the robot to be trained on every task that it carries out or do you have some requirement for a "general" intelligence that can carry our arbitrary tasks? I know this is easier said than done, but interested in your opinion as I have no experience with robotics.

Yu Xiang profil fotoğrafı
Yu Xiang10 ay önce

This task seems very difficult. Can the policy deal with different shoes?

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

thanks! Currently we tested on three different colored shoes. However, some other shoe designs are extremely difficult to thread, even for human demonstrators. There is much more we can do in the future :)

sam profil fotoğrafı
sam10 ay önce

very impressive :)

Porters Reserve profil fotoğrafı
Porters Reserve10 ay önce

😁 this why a professor is just person who memorized a book to tell you to read the same book. Only real world grit and experimentation is the path to success. Great Job now level up send your gear to the reserve to see what really breaks.

Rxct Ftvy profil fotoğrafı
Rxct Ftvy10 ay önce

Really really cool

Rxct Ftvy profil fotoğrafı
Rxct Ftvy10 ay önce

Very cool

Offend_no_one profil fotoğrafı
Offend_no_one10 ay önce

Wow, Yusuf, that’s incredible progress from ByteDance!

Younggyo Seo profil fotoğrafı
Younggyo Seo10 ay önce

Cool, congrats on the release!!

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

Thank you Younggyo!

Griffin profil fotoğrafı
Griffin10 ay önce

Incredibly impressive, well done!

Fanf666 profil fotoğrafı
Fanf66610 ay önce

Can you clarify 1) the number of examples to reach your quoted accuracy 2) how many distinct tasks can the model remember and how does it select them

Fanf666 profil fotoğrafı
Fanf66610 ay önce

Here is prior art shoelace tyeing What matter is to advance few shot, max task number and length.

Roméo profil fotoğrafı
Roméo10 ay önce

@grok analyse et explique ce qui se passe dans cette vidéo

czr profil fotoğrafı
czr10 ay önce

Cool !

尚月 profil fotoğrafı
尚月10 ay önce

还有哪些常见的日常活动对机器人而言与系鞋带平级或者更复杂?

Dan Wu profil fotoğrafı
Dan Wu10 ay önce

So sick! Curious what cameras are mounted onto the end effectors?

Xiao Ma profil fotoğrafı
Xiao Ma10 ay önce

thanks! We have discussed our hardware setups in the section 4 of our paper. It is a realsense D405 camera.

Muhammad Umar 🇺🇸 profil fotoğrafı
Muhammad Umar 🇺🇸10 ay önce

It's just a matter of time we'll solve dexterity. This is just unreal.

Commentary Yi He profil fotoğrafı
Commentary Yi He9 ay önce

Amazing project! 👏 Message me directly

Commentary Yi He profil fotoğrafı
Commentary Yi He9 ay önce

Awesome 👏 Message me directly

Reza Sayar profil fotoğrafı
Reza Sayar10 ay önce

this is awesome! 🔥 congrats! 👏🏼 i noticed even though you had RGB-D cameras, you only kept the RGB for training and inference. did you find that depth data won't help here much?

Laughless profil fotoğrafı
Laughless10 ay önce

This is more impressive than dancing or kung fu. Do you think it could tie a neck tie without choking a person? That's one of my benchmarks for when robots are ready for home use.

Raul Verdusco profil fotoğrafı
Raul Verdusco10 ay önce

Incredible progress — deformable object manipulation has always been one of the toughest challenges in robotics, and seeing real-world RL push success rates this high is remarkable. Huge step forward for dexterous autonomy. 🤖🧠 #Robotics #ReinforcementLearning

Golding profil fotoğrafı
Golding10 ay önce

This is pretty cool man. I'd love to learn more about your research and ByteDance Seed. Can you open your DM's? let's chat!

Farhan Azad Shuvra profil fotoğrafı
Farhan Azad Shuvra10 ay önce

Impressive Stuff!!

X Æ A-12 profil fotoğrafı
X Æ A-1210 ay önce

you are the best 😍🧡💙💚 !

Benzer Videolar

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 görüntüleme • 2 ay önce