Video wird geladen...
Video konnte nicht geladen werden
Introducing Scalable Option Learning (SOL☀️), a blazingly fast hierarchical RL algorithm that makes progress on long-horizon tasks and demonstrates positive scaling trends on the largely unsolved NetHack benchmark, when trained for 30 billion samples. Details, paper and code in >
21,956 Aufrufe • vor 1 Jahr •via X (Twitter)
5 Kommentare

Hierarchy is a natural way to tackle long horizons, but until now has remained at relatively small scale. With SOL, we identify and solve bottlenecks in scaling hierarchical RL, resulting in a ~35-580x speed increase over prior hierarchical methods.

We demonstrate SOL's performance and scalability by training hierarchical agents for 30B steps on the complex game of NetHack, significantly outperforming flat agents and demonstrating promising scaling trends. Our agents still seem to be improving, even after 30B steps.

SOL can be run on any RL problem for which we can define a few reasonable intrinsic rewards. We include some simple PointMaze and MiniHack environments to show this generality - this may also be useful for others working in HRL since they are faster to iterate on than NetHack.

All details are in our paper and code release: It was lots of fun working with Scott Fujimoto, @mitrma and Mike Rabbat on this!

This is so cool! Thanks: ) been looking for good HRL repos!
