Загрузка видео...

Не удалось загрузить видео

На главную

I've been working on deformable object manipulation since my PhD. It was totally a nightmare years ago and my PhD advisor was telling me not to work on it for my own good. Today, at ByteDance Seed, we are dropping GR-RL, a new VLA+RL system that manages long-horizon precise...

110,653 просмотров • 10 месяцев назад •via X (Twitter)

Комментарии: 45

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Why “shoelace threading” matters 🤔 This task is probably one of the most challenging household robotics tasks in terms of precision: 💥 Soft-body chaos – laces deform every frame 💥 Millimeter precision – 1–2 mm slip = total failure 💥 Long-horizon manipulation – hundreds of steps where slightest drift compounds. 💥 Requires recovery – humans retry; IL models freeze It's the perfect torture test for VLA precision. Even our previous model, GR-3 — already trained on massive robot trajectories + human teleop demos + public image-text corpora — failed to get reliable policies here. So we found the bottleneck.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Two killers of imitation learning (IL): (1) Human demos are NOT optimal Humans hesitate, retry, fix mistakes mid-trajectory. IL blindly copies ALL of it — including the bad parts. (2) Training vs Deployment Misalignment VLA models output actions. To prevent jitter, robots execute post-processed versions (RHC, smoothing, ensembling). Your “predicted action” ≠ your “executed action.” On millimeter tasks, this mismatch is fatal. So IL alone was a dead end.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

The Idea: If imitation is broken, then: Let the robot learn from its own experience. GR-RL = ⭐️ Offline RL (data filtering) ⭐️ Symmetry augmentation ⭐️ Online closed-loop Real-World Reinforcement Learning All on top of a single VLA foundation model. Just RGB, proprioception, and language instructions.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Offline Stage — Filter the human flaws We train a Critic Transformer via distributional RL: ⭐️ Detects “value drops” when the operator hesitates or messes up ⭐️ Slices every trajectory into high-value vs low-value segments ⭐️ Retains only the cleanest expert behavior Effect: GR-3 45.7% → 61.6% success just from removing “bad micro-behaviors.”

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Morphological Symmetry Augmentation Our bi-manual robot is left–right symmetric. So we mirror EVERYTHING: 🚀 RGB 🚀 Proprioception 🚀 Actions 🚀 Language Instructions Data size doubles. Spatial reasoning robustness skyrockets. → 72.7% success.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Online Stage — Real-World Steering RL Now the robot learns ON THE PHYSICAL PLATFORM. But direct exploration in joint space causes dangerous jitter and is inefficient — you need millimeter accuracy. So GR-RL explores in the latent noise space: A tiny 51.5M-param noise-predictor nudges the flow model’s latent action distribution toward high-value regions. We also introduce a dual-buffer mechanism: 🚀 Off-policy buffer (old rollouts) for the critic warm-up 🚀 On-policy buffer (last 2 checkpoints) for stable improvement 🚀 1:1 sampling stabilizes RL and prevents catastrophic drift. Just 150 online rollouts → +10% performance boost.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

The Result: Final performance: 83.3% success over continuous shoelace threading. The surprising part? GR-RL learns to: 🔥 retry when the lace slips 🔥 reposition the lace when the initial pose is bad 🔥 "self-correct" mid-task instead of freezing This is the behavior you EXPECT from embodied intelligence — not just mimicry, but competence.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

GR-RL proves something: IL is inherently limited, and we can do things previously thought impossible by purely visuo-motor control simply making it RL. The future direction is clear: distill RL-enhanced behavior back into the foundation VLA, forming a self-improving, ever-growing loop. Kudos to the wonderful team! @yunfeili98, @simanbulk, @yu_cui123, Zhongren Cui, Zhigang Han, @litchiyoyo, @KT892790499, @YuxiaoLiu4, Hao Niu, Wanli Peng, Jingchao Qiao, Zeyu Ren, Haixin Shi, @ZhiSu22, Jiawen Tian, Yuyang Xiao, Shenyu Zhang, Liwei Zheng

Фото профиля Perry Jia
Perry Jia10 месяцев назад

Wow! I love how it naturally pulls the shoe closer so that it can complete the task easier. Awesome work!

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

haha, thanks Perry! btw congrats on your recent release on ACT-1 and Memo! Great work! I really like the design and it's amazing how much it has achieved with pure UMI data.

Фото профиля 🄱🄻🄸🅃🅉 🔸
🄱🄻🄸🅃🅉 🔸10 месяцев назад

This is brilliant Ma

Фото профиля Adam
Adam10 месяцев назад

incredible work, can't wait to see more in the future.

Фото профиля Daniel Domingos
Daniel Domingos10 месяцев назад

Congrats!

Фото профиля MAX ONBOARDER ⭕
MAX ONBOARDER ⭕10 месяцев назад

@Scobleizer This is amazing

Фото профиля Bingyi Kang
Bingyi Kang10 месяцев назад

Great team! Great Work! Can not wait to see your future project! Congrats @yunfeili98 @yusufma555

Фото профиля Salman
Salman10 месяцев назад

This is incredible work. Kudos for seeing this work through for so many years, seems like it’s paying off!

Фото профиля Robert
Robert10 месяцев назад

Can you implement an aspect of leaning against the object that you're working on? This is what humans do when we tie our shoe, we don't float in space. This means that you don't have to have accurate control, you know that if you rest your finger against the shoe, then the string will be in the right place.

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Hey Robert, this is a perfect point. Our model actually sometimes does it. During the offline data collection stage, most of the human demonstrators love to "float" because it is easier for them to control the robot when they became very good at it. However, this causes the issue you mentioned. But there were indeed some samples that the demonstrators love to lean on the shoes. As a result, our VLA model will demonstrate multi behavior modes, including the one you mentioned. But it is just less often, due to the initial data distribution was dominated by the other behavior mode.

Фото профиля Robert
Robert10 месяцев назад

Awesome, will this methodology ultimately require the robot to be trained on every task that it carries out or do you have some requirement for a "general" intelligence that can carry our arbitrary tasks? I know this is easier said than done, but interested in your opinion as I have no experience with robotics.

Фото профиля Yu Xiang
Yu Xiang10 месяцев назад

This task seems very difficult. Can the policy deal with different shoes?

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

thanks! Currently we tested on three different colored shoes. However, some other shoe designs are extremely difficult to thread, even for human demonstrators. There is much more we can do in the future :)

Фото профиля sam
sam10 месяцев назад

very impressive :)

Фото профиля Porters Reserve
Porters Reserve10 месяцев назад

😁 this why a professor is just person who memorized a book to tell you to read the same book. Only real world grit and experimentation is the path to success. Great Job now level up send your gear to the reserve to see what really breaks.

Фото профиля Rxct Ftvy
Rxct Ftvy10 месяцев назад

Really really cool

Фото профиля Rxct Ftvy
Rxct Ftvy10 месяцев назад

Very cool

Фото профиля Offend_no_one
Offend_no_one10 месяцев назад

Wow, Yusuf, that’s incredible progress from ByteDance!

Фото профиля Younggyo Seo
Younggyo Seo10 месяцев назад

Cool, congrats on the release!!

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

Thank you Younggyo!

Фото профиля Griffin
Griffin10 месяцев назад

Incredibly impressive, well done!

Фото профиля Fanf666
Fanf66610 месяцев назад

Can you clarify 1) the number of examples to reach your quoted accuracy 2) how many distinct tasks can the model remember and how does it select them

Фото профиля Fanf666
Fanf66610 месяцев назад

Here is prior art shoelace tyeing What matter is to advance few shot, max task number and length.

Фото профиля Roméo
Roméo10 месяцев назад

@grok analyse et explique ce qui se passe dans cette vidéo

Фото профиля czr
czr10 месяцев назад

Cool !

Фото профиля 尚月
尚月10 месяцев назад

还有哪些常见的日常活动对机器人而言与系鞋带平级或者更复杂?

Фото профиля Dan Wu
Dan Wu10 месяцев назад

So sick! Curious what cameras are mounted onto the end effectors?

Фото профиля Xiao Ma
Xiao Ma10 месяцев назад

thanks! We have discussed our hardware setups in the section 4 of our paper. It is a realsense D405 camera.

Фото профиля Muhammad Umar 🇺🇸
Muhammad Umar 🇺🇸10 месяцев назад

It's just a matter of time we'll solve dexterity. This is just unreal.

Фото профиля Commentary Yi He
Commentary Yi He9 месяцев назад

Amazing project! 👏 Message me directly

Фото профиля Commentary Yi He
Commentary Yi He9 месяцев назад

Awesome 👏 Message me directly

Фото профиля Reza Sayar
Reza Sayar10 месяцев назад

this is awesome! 🔥 congrats! 👏🏼 i noticed even though you had RGB-D cameras, you only kept the RGB for training and inference. did you find that depth data won't help here much?

Фото профиля Laughless
Laughless10 месяцев назад

This is more impressive than dancing or kung fu. Do you think it could tie a neck tie without choking a person? That's one of my benchmarks for when robots are ready for home use.

Фото профиля Raul Verdusco
Raul Verdusco10 месяцев назад

Incredible progress — deformable object manipulation has always been one of the toughest challenges in robotics, and seeing real-world RL push success rates this high is remarkable. Huge step forward for dexterous autonomy. 🤖🧠 #Robotics #ReinforcementLearning

Фото профиля Golding
Golding10 месяцев назад

This is pretty cool man. I'd love to learn more about your research and ByteDance Seed. Can you open your DM's? let's chat!

Фото профиля Farhan Azad Shuvra
Farhan Azad Shuvra10 месяцев назад

Impressive Stuff!!

Фото профиля X Æ A-12
X Æ A-1210 месяцев назад

you are the best 😍🧡💙💚 !

Похожие видео

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 просмотров • 2 месяцев назад