Loading video...

Video Failed to Load

Go Home

I hand-wrote a 500-LoC RL stack to make hacking on RL research much easier. Most RL stacks are either massive and unhackable, or duct-taped research scripts. I am open-sourcing Mithrl, a modular RLVR stack. Next items on my checklist: adding more complex environment examples, supporting multi-gpu + async RL,...

17,676 views • 6 months ago •via X (Twitter)

36 Comments

omkaar's profile picture
omkaar6 months ago

cc for feedback: @willccb @QGallouedec @_rajanagarwal @yacinelearning @danielhanchen (learnt vllm sleep/wake from here) @karpathy (easy to port to autoresearch which is what I am doing next)

omkaar's profile picture
omkaar6 months ago

github:

parth's profile picture
parth6 months ago

hell yeah

omkaar's profile picture
omkaar6 months ago

ship squad

surya's profile picture
surya6 months ago

keep cooking omkizzy

omkaar's profile picture
omkaar6 months ago

me and you forever

Jibran's profile picture
Jibran6 months ago

@hamostaf04 this is sick

omkaar's profile picture
omkaar6 months ago

@hamostaf04 yessir thank you

mikael haji's profile picture
mikael haji6 months ago

lets gooo this is fire

omkaar's profile picture
omkaar6 months ago

appreciate you brother

saksham's profile picture
saksham6 months ago

yoooo that’s tuff

omkaar's profile picture
omkaar6 months ago

thank you brother

ari dutilh's profile picture
ari dutilh6 months ago

man this is awesome

omkaar's profile picture
omkaar6 months ago

i appreciate it man, lmk if you have ppl experimenting w LLM RL I could talk to

ari dutilh's profile picture
ari dutilh6 months ago

@novasarc01 potentially

omkaar's profile picture
omkaar6 months ago

@novasarc01 hi @novasarc01, would love to chat if this is interesting

Damon Crockett's profile picture
Damon Crockett6 months ago

Hand wrote, nice

toki's profile picture
toki6 months ago

this is fire

omkaar's profile picture
omkaar6 months ago

appreciate it brother, lmk if you run any experiments on top

Madhav Singhal's profile picture
Madhav Singhal6 months ago

this is dope

omkaar's profile picture
omkaar6 months ago

thank you brother, you first put me on actual RLVR training

Chris Samra's profile picture
Chris Samra6 months ago

🔥

Fahim Ahmed's profile picture
Fahim Ahmed6 months ago

LFG!

omkaar's profile picture
omkaar6 months ago

yesssir

atharva ☆'s profile picture
atharva ☆6 months ago

this is incredible

omkaar's profile picture
omkaar6 months ago

thank you brother

Ishan's profile picture
Ishan6 months ago

@k7agar Nice work!

omkaar's profile picture
omkaar6 months ago

@k7agar thank you!

Jash Thakkar's profile picture
Jash Thakkar6 months ago

Nice work, will try this out if i get time to.

aarush's profile picture
aarush6 months ago

so based

omkaar's profile picture
omkaar6 months ago

yessir down to collab on an interesting env

Krish Modi's profile picture
Krish Modi6 months ago

sick!!!

omkaar's profile picture
omkaar6 months ago

@therealkmodi thank you brother!

shiv's profile picture
shiv6 months ago

Let's connect

adi's profile picture
adi6 months ago

holy shit

omkaar's profile picture
omkaar6 months ago

would love to collab if andera is looking into internal rl envs

Related Videos

🚨 RL for LLMs is finally accessible. Introducing OpenTinker: The first community-driven, open-source framework designed to democratize Reinforcement Learning for LLMs. Inspired by Thinking Machines's amazing Tinker, we realize the biggest bottleneck in agentic LLM research isn’t the math—it’s the setup. Current RL pipelines are messy. Configuring VeRL for every single experiment is a productivity killer. OpenTinker fixed it. 🛠 How OpenTinker Works: Decoupled Design of Server and Client - Setup Once, Run Forever: Configure the OpenTinker backend on your GPU cluster once. - Develop Locally: Define your RL environments directly on your laptop. - Train on the Cloud: Simply point your local client to the backend. The cluster handles the compute; you handle the science. 📉 The 10x Development Efficiency Thanks to our elegant architectural decomposition, OpenTinker reduces the time to develop a new RL training pipeline by at least an order of magnitude. ⚡ Turn Idle GPU Compute into Gold Small labs often have underutilized hardware. OpenTinker turns your idle GPUs into an internal/external API service for - RL Training - SFT - Inference 🎯 Who needs OpenTinker? - Researchers tired of infrastructure hell. - Labs needing to standardize workflows. - Teams wanting to maximize hardware ROI. Thanks my amazing PhD student Siqi Zhu for leading the project. We are building the future of open RL infra. Be the first to build with us. 👇 Start Building with OpenTinker Now 🚀 Repo: 🌐 Blog: If you believe RL should be accessible to everyone, give us a star, repost this 🔄 post, and let us know what agents you plan to build!

Jiaxuan You

58,326 views • 9 months ago

OpenClaw meets RL! OpenClaw Agents adapt through memory files and skills, but the base model weights never actually change. OpenClaw-RL solves this! It wraps a self-hosted model as an OpenAI-compatible API, intercepts live conversations from OpenClaw, and trains the policy in the background using RL. The architecture is fully async. This means serving, reward scoring, and training all run in parallel. Once done, weights get hot-swapped after every batch while the agent keeps responding. Currently, it has two training modes: - Binary RL (GRPO): A process reward model scores each turn as good, bad, or neutral. That scalar reward drives policy updates via a PPO-style clipped objective. - On-Policy Distillation: When concrete corrections come in like "you should have checked that file first," it uses that feedback as a richer, directional training signal at the token level. When to use OpenClaw-RL? To be fair, a lot of agent behavior can already be improved through better memory and skill design. OpenClaw's existing skill ecosystem and community-built self-improvement skills handle a wide range of use cases without touching model weights at all. If the agent keeps forgetting preferences, that's a memory problem. And if it doesn't know how to handle a specific workflow, that's a skill problem. Both are solvable at the prompt and context layer. Where RL becomes interesting is when the failure pattern lives deeper in the model's reasoning itself. Things like consistently poor tool selection order, weak multi-step planning, or failing to interpret ambiguous instructions the way a specific user intends. Research on agentic RL (like ARTIST and Agent-R1) has shown that these behavioral patterns hit a ceiling with prompt-based approaches alone, especially in complex multi-turn tasks where the model needs to recover from tool failures or adapt its strategy mid-execution. That's the layer OpenClaw-RL targets, and it's a meaningful distinction from what OpenClaw offers. I have shared the repo in the replies!

Avi Chawla

138,769 views • 6 months ago

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts

Sonya Huang 🐥

81,332 views • 4 months ago

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a higher-dimensional feedback channel. Every major AI lab trains models with RL today (OpenAI, Anthropic, DeepSeek). And their key bottleneck has always been the reward functions. GRPO by DeepSeek worked well for math and code because the environment gave a binary signal. But for real agent tasks, someone still has to hand-code the scoring function. That takes days and breaks every time the pipeline changes. RULER (implemented in OpenPipe ART, 10k stars) addresses the exact problem Karpathy identified. The reward criteria are defined in plain English, and an LLM evaluates each trajectory against that description to provide feedback for training. I trained a Qwen3 1.4B agent that plays 2048 using GRPO with this exact workflow. In this case, the agent saw the board, picked a direction, and RULER evaluated the outcome, all from this natural language definition. You can see the full implementation on GitHub and try it yourself. Here's the ART Repo: (don't forget to star it ⭐ ) Just like RLHF replaced manual rankings and GRPO replaced the critic model, natural language rewards are replacing hand-coded scoring functions. RL reward engineering is now prompt engineering. I wrote a full walkthrough covering RL for LLM agents, from RLHF to GRPO to RULER, in the article below.

Avi Chawla

350,798 views • 4 months ago