Loading video...
Video Failed to Load
I hand-wrote a 500-LoC RL stack to make hacking on RL research much easier. Most RL stacks are either massive and unhackable, or duct-taped research scripts. I am open-sourcing Mithrl, a modular RLVR stack. Next items on my checklist: adding more complex environment examples, supporting multi-gpu + async RL,... show more
17,218 views • 3 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
