Loading video...
Video Failed to Load
Robots are getting smarter, but most still fail the same way. They don’t learn from their own mistakes. A new paper proposes something different: a way for robots to self-improve directly from their failures in the real world. It’s called PLD (Probe, Learn, Distill). The idea: instead of collecting... show more
22,891 views • 11 months ago •via X (Twitter)
9 Comments

Cool, but crude. The key to doing this well is having a better way to grip the card. I worked at Texas A&M's Robotics Lab in the mid 80s, where we solved this exact problem for IBM' s PC AT and RT assembly operations in Austin and Louisville (they put our solution in production!) Of course, it neither needed nor used AI... There are two keys: 1) Realizing the only safe place to grab the card is by the edges - we built a custom reverse scissors tool for the gripper to do this (Larry Chan's great idea) and 2) building a screwdriver sleeve and screw feeding system that simply cannot ever drop a screw when it is screwing the card in after it's inserted. I designed the screw feeding system. The key to avoiding jams and double feeding to get perfect reliability was to simply remove all the parts. Literally - The only moving part became the screw itself. The IBM 7565 robot we used had advanced dynamic compliance algos in it even then that took care of the rest once we had these fixes and decent fixturing.

"They don’t learn from their own mistakes." Human after all.

it is not new. the problem is how to decrease the training cost the training time

Thanks for featuring!

Residual RL for recovery is clever. Instead of retraining the whole policy, just learn the correction layer. Faster iteration, less compute, better results. That's practical engineering for real systems.

Interesting approach! If only we could teach them to avoid the classic “oops, I did it again” moment.

Spot on! With KENTA, we go further: no forced RL, just "organic memory" that turns crashes into evolution. Tested: 2482 events, anomalies sensed in 0.3s, and it even sings its CPU cycles in pentatonic for expression. Offline, privacy-first, on a 2012 i5. Antifragile in living bits. Demo? #KENTA #SelfHealingAI

How can it sustain multiple attempts if a failure were to damage that which it is handling? How does it know when it is damaged? There still seems to be a need for human intervention.

i don't know why this concept is considered groundbreaking. humans and animals learn by doing and failing and doing it again and again. why can't robots or human/animal simulators do the same?
