Загрузка видео...

Не удалось загрузить видео

На главную

Lighting differences can make a huge difference in robotics. Today, I found a quirk in my model exemplifying this. > I collected 10h of training data. > 3h in, I notice that the left arm following the right arm for the final movement could be good for the final...

97,968 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 24

Фото профиля KuphDev
KuphDev3 месяцев назад

Yep. I had to re-train one of my ACT models last night cuz when I recorded the original data I had a light placed in a weird position which gave me all sorts of reliability issues. Haha I hope the new model works cuz the stream is going live tomorrow 😅

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

Cool!

Фото профиля KuphDev
KuphDev3 месяцев назад

I've actually still got a few hours left til the training wraps up and I can test again 🤞

Фото профиля Binh
Binh3 месяцев назад

we should have better vision encoders for this, like having a loss func that doesn’t penalize difference in lighting condition in reconstruction so latents for 2 images with diff lighting is the same

Фото профиля pfung
pfung3 месяцев назад

perhaps should encase in a self-lighted box like some other startups do

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

Parts are already on their way!

Фото профиля Nathaniel Nifong
Nathaniel Nifong3 месяцев назад

There are so many little details to learn, but I am starting to feel confident in a few things and one is that you really should turn on image augmentations every time you train something. Specifically, this means the slight brightness color and contrast jitter settings. And without you having to do any extra work in data collection, this helps with things like lighting sensitivity.

Фото профиля Luděk Čižinský
Luděk Čižinský3 месяцев назад

Is that VLA based policy?

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

Yes, π0.5 finetune on 10h of data

Фото профиля Mi.lu.
Mi.lu.3 месяцев назад

Now imagine this in real production: every company, every site, different lighting, different conditions. The operator is definitely not going to lower the blinds just to make it work. Do you already have a solution for this?

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

Yeah, collect consistent data the next time 😂

Фото профиля Mi.lu.
Mi.lu.3 месяцев назад

Fair 😂 if it were only that simple, robotics would be half as fun....

Фото профиля atharva ☆
atharva ☆3 месяцев назад

haha so cool

Фото профиля Daniel Friis
Daniel Friis3 месяцев назад

Can't this be fixed by creating new training data based on original but with exposure changed? At least partly

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

Yes totally. But i need to collect that data now haha

Фото профиля Daniel Friis
Daniel Friis3 месяцев назад

I mean by creating variations of your training data programmatically with changed exposure, brightness, color etc? :)

Фото профиля Angkul
Angkul3 месяцев назад

I never train a real robot but I don’t want to spend any time optimising these micro learnings. ideally models should generalise to different lighting conditions. > I change behaviour can u explain your 3rd point. I didn’t get it. How you changed the behaviour?

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

I collected data in a different way. I originally did it the way you see in the first trial and then changed to the behaviour of the robot when its dark. If you don't want to spend time optimising these things I'd just wait another two years before getting into robotics haha

Фото профиля Angkul
Angkul3 месяцев назад

Got it. Makes sense to me now. haha I will surely enjoy it but we’ll see once I get my hands on some hardware. rn i m playing sim-sim only

Фото профиля Antoni
Antoni3 месяцев назад

Check it out - - they focused exactly on the problem of maintaining robustness under sensor noise, and every lab has implemented some notion of that

Фото профиля Dominique Paul
Dominique Paul3 месяцев назад

Cool thanks!

Фото профиля Antoni
Antoni3 месяцев назад

Not a problem! I can't wait to shake your robotic hand at the local manufacturer

Фото профиля Diego P. Jaccottet
Diego P. Jaccottet3 месяцев назад

JEPA is meant to solve this.

Фото профиля Nurvai - The Data Layer for Physical AI
Nurvai - The Data Layer for Physical AI3 месяцев назад

This is a great robotics example of a classic ML issue. It reminds us of the CNN that learned to classify wolves by detecting snow in the background instead of the animal itself. With black box models, dataset diversity is key to ensuring the policy learns the task, not accidental shortcuts.

Похожие видео

One question that's been on my mind for years now is: could we use regular multimodal LLMs not necessarily trained for robotics to do the high level robotics intelligence part that VLAs and WAMs attempt to do? The latest explosion of powerful opensource multi-modal LLMs has, IMO, begun to make this possible due both to intelligence and speed. This is GLM 5.3 Flash, which has vision understanding, but isn't meant to be a VLA/VLM/WAM/robotics model at all, controlling an XGO mini wheeled robot quadruped with an arm & gripper. GLM 5.3F simply has access to the robot's high level SDK for controlling movement, arm joints, open/close gripper...etc. It analyzes the frames from the camera and makes adjustments all on its own to solve the task. Nothing was trained here, nothing fine-tuned for this task. Z AI did not make this model for robots and tbh I think they're surprised this works when I talk to them about it! This also works quite well with DSV4F + a vision capable model like Qwen 3.8 27B. I havent tried JUST Qwen 3.8 27B, but I'm sure it works too. I like the "logic" to be a model that's as fast as possible (but still intelligent). There's also an experimental vision version of DSV4F, I'm confident that'll work too and might even be better bc the full loop might be the fastest of all with this model. An obvious question you might wonder is: well why not use VLA or VLM? The hard part about robotics isn't object detection, that's long solved. This also isn't a solution for gait/locomotion...yet, but I actually don't think this is far away either and I've done some experimentation with LLMs in this space in the past and it does show promise. It might actually already be here for quadrupeds, since you dont need super fast IMU readings to maintain balance. I've also tried many of the larger, more generalist, VLAs that you should be able to use with popular robots and tbh there are just so many edge cases that make things hard and not work. You gotta get the camera, lighting, task, everything *just right* or the demo fails. This is for the actual hard part in robotics right now: intelligence, logic, and planning for all the ways the real world just simply isn't perfect. I've trained VLAs. They're super finicky and you're always running into sim2real issues, especially around the camera. You also have to build the whole training pipeline in a simulator, and, if everything does work, you still just have a robot that does this 1 single thing after weeks of work. If you use teleop, this overcomes the "2real" problem, but now you need to painstakingly collect teleop data, and it's only good at that specific task and that particular robot. There is a growing set of egocentric training data for "general purpose" VLAs and world action models (for humanoid form factors), but I'm really starting to wonder: Why? I think we might just sidestep this whole area of research entirely. I didn't need any training data or special environment to work with this quadruped and arm to do the task I was after. This particular quadruped and arm doesn't even exist in the wild yet really, it's a demo build from a company launching it on kickstarter, so it's not like this robot's data exists in the LLM to any real extent. I think this is cool as heck that this works and I am interested to see just how far I can push it. Also this marks the first time that I've finally got a generalist solution to a task I've been trying to solve ever since I became a dad of twins: pick up toys off the ground. This is a big day!

Harrison Kinsley

53,755 просмотров • 27 дней назад