Video wird geladen...
Video konnte nicht geladen werden
🤖Adding new RL algorithms to LeRobot just got much easier. Demo: HIL-SERL training with a SAC-based RL algorithm on an SO-100 for a hole-in-hand peg-in-hole task. Sparse reward, only 30 offline demos mixed with live robot experience, and ~1 hour of online training with human interventions only when the... show more
31,237 Aufrufe • vor 5 Monaten •via X (Twitter)
10 Kommentare

Check out the docs:

The biggest thing here is making RL algorithms feel plug and play instead of tightly coupled to the training stack. That’s probably more important long term than the SAC result itself

finally a framework that lets me swap algorithms without rewriting the whole pipeline.

Interesting the use of sparse reward with only 30 offline demos. Most curious the speed of convergence.

Clean refactor on LeRobot.

Interesting to see RL infrastructure becoming modular enough that algorithms are almost plug and play now. The intervention curve is especially compelling for real world training. How much do you think reusable infrastructure will accelerate experimentation compared to new algorithms themselves?

Just fine-tuned SmolVLA on a bimanual folding task, any insights appreciated:

so adding new algorithms is just copy paste now? guess the robot will learn to hate manual configs

the rl side is exciting but what happens after training is where most teams stall. going from a lerobot checkpoint to something running on real hardware at inference speed is still a whole separate engineering problem

nice demo. I’m mostly curious about transfer here: same policy idea on a different arm or fixture, or does that become a whole new weekend?

