Video wird geladen...
Video konnte nicht geladen werden
Specifying reward functions for robots is one of the hardest things about reinforcement learning. Robot rewards often need to be very detailed; metrics like progress can be ill-defined and hard to estimate. This leads to most robot learning defaulting to sparse rewards or simple preference learning. But naive preference... show more
13,768 Aufrufe • vor 4 Tagen •via X (Twitter)
2 Kommentare

Nghẹo | DOPEvor 3 Tagen
Sparse rewards often hide the behaviors we actually want

Muhammad Ahmedvor 3 Tagen
A preference like 'better trajectory' hides the reason for choosing it. Being able to say 'less force' or 'fewer unnecessary motions' gives robot learning feedback that's much closer to what we actually care about.
