正在加载视频...
视频加载失败
I hand-wrote a 500-LoC RL stack to make hacking on RL research much easier. Most RL stacks are either massive and unhackable, or duct-taped research scripts. I am open-sourcing Mithrl, a modular RLVR stack. Next items on my checklist: adding more complex environment examples, supporting multi-gpu + async RL,... show more
17,676 次观看 • 6 个月前 •via X (Twitter)
36 条评论

cc for feedback: @willccb @QGallouedec @_rajanagarwal @yacinelearning @danielhanchen (learnt vllm sleep/wake from here) @karpathy (easy to port to autoresearch which is what I am doing next)

github:

hell yeah

ship squad

keep cooking omkizzy

me and you forever

@hamostaf04 this is sick

@hamostaf04 yessir thank you

lets gooo this is fire

appreciate you brother

yoooo that’s tuff

thank you brother

man this is awesome

i appreciate it man, lmk if you have ppl experimenting w LLM RL I could talk to

@novasarc01 potentially

@novasarc01 hi @novasarc01, would love to chat if this is interesting

Hand wrote, nice

this is fire

appreciate it brother, lmk if you run any experiments on top

this is dope

thank you brother, you first put me on actual RLVR training

🔥

LFG!

yesssir

this is incredible

thank you brother

@k7agar Nice work!

@k7agar thank you!

Nice work, will try this out if i get time to.

so based

yessir down to collab on an interesting env

sick!!!

@therealkmodi thank you brother!

Let's connect

holy shit

would love to collab if andera is looking into internal rl envs
