Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I hand-wrote a 500-LoC RL stack to make hacking on RL research much easier. Most RL stacks are either massive and unhackable, or duct-taped research scripts. I am open-sourcing Mithrl, a modular RLVR stack. Next items on my checklist: adding more complex environment examples, supporting multi-gpu + async RL,...

17,676 Aufrufe • vor 6 Monaten •via X (Twitter)

36 Kommentare

Profilbild von omkaar
omkaarvor 6 Monaten

cc for feedback: @willccb @QGallouedec @_rajanagarwal @yacinelearning @danielhanchen (learnt vllm sleep/wake from here) @karpathy (easy to port to autoresearch which is what I am doing next)

Profilbild von omkaar
omkaarvor 6 Monaten

github:

Profilbild von parth
parthvor 6 Monaten

hell yeah

Profilbild von omkaar
omkaarvor 6 Monaten

ship squad

Profilbild von surya
suryavor 6 Monaten

keep cooking omkizzy

Profilbild von omkaar
omkaarvor 6 Monaten

me and you forever

Profilbild von Jibran
Jibranvor 6 Monaten

@hamostaf04 this is sick

Profilbild von omkaar
omkaarvor 6 Monaten

@hamostaf04 yessir thank you

Profilbild von mikael haji
mikael hajivor 6 Monaten

lets gooo this is fire

Profilbild von omkaar
omkaarvor 6 Monaten

appreciate you brother

Profilbild von saksham
sakshamvor 6 Monaten

yoooo that’s tuff

Profilbild von omkaar
omkaarvor 6 Monaten

thank you brother

Profilbild von ari dutilh
ari dutilhvor 6 Monaten

man this is awesome

Profilbild von omkaar
omkaarvor 6 Monaten

i appreciate it man, lmk if you have ppl experimenting w LLM RL I could talk to

Profilbild von ari dutilh
ari dutilhvor 6 Monaten

@novasarc01 potentially

Profilbild von omkaar
omkaarvor 6 Monaten

@novasarc01 hi @novasarc01, would love to chat if this is interesting

Profilbild von Damon Crockett
Damon Crockettvor 6 Monaten

Hand wrote, nice

Profilbild von toki
tokivor 6 Monaten

this is fire

Profilbild von omkaar
omkaarvor 6 Monaten

appreciate it brother, lmk if you run any experiments on top

Profilbild von Madhav Singhal
Madhav Singhalvor 6 Monaten

this is dope

Profilbild von omkaar
omkaarvor 6 Monaten

thank you brother, you first put me on actual RLVR training

Profilbild von Chris Samra
Chris Samravor 6 Monaten

🔥

Profilbild von Fahim Ahmed
Fahim Ahmedvor 6 Monaten

LFG!

Profilbild von omkaar
omkaarvor 6 Monaten

yesssir

Profilbild von atharva ☆
atharva ☆vor 6 Monaten

this is incredible

Profilbild von omkaar
omkaarvor 6 Monaten

thank you brother

Profilbild von Ishan
Ishanvor 6 Monaten

@k7agar Nice work!

Profilbild von omkaar
omkaarvor 6 Monaten

@k7agar thank you!

Profilbild von Jash Thakkar
Jash Thakkarvor 6 Monaten

Nice work, will try this out if i get time to.

Profilbild von aarush
aarushvor 6 Monaten

so based

Profilbild von omkaar
omkaarvor 6 Monaten

yessir down to collab on an interesting env

Profilbild von Krish Modi
Krish Modivor 6 Monaten

sick!!!

Profilbild von omkaar
omkaarvor 6 Monaten

@therealkmodi thank you brother!

Profilbild von shiv
shivvor 6 Monaten

Let's connect

Profilbild von adi
adivor 6 Monaten

holy shit

Profilbild von omkaar
omkaarvor 6 Monaten

would love to collab if andera is looking into internal rl envs

Ähnliche Videos

🚨 RL for LLMs is finally accessible. Introducing OpenTinker: The first community-driven, open-source framework designed to democratize Reinforcement Learning for LLMs. Inspired by Thinking Machines's amazing Tinker, we realize the biggest bottleneck in agentic LLM research isn’t the math—it’s the setup. Current RL pipelines are messy. Configuring VeRL for every single experiment is a productivity killer. OpenTinker fixed it. 🛠 How OpenTinker Works: Decoupled Design of Server and Client - Setup Once, Run Forever: Configure the OpenTinker backend on your GPU cluster once. - Develop Locally: Define your RL environments directly on your laptop. - Train on the Cloud: Simply point your local client to the backend. The cluster handles the compute; you handle the science. 📉 The 10x Development Efficiency Thanks to our elegant architectural decomposition, OpenTinker reduces the time to develop a new RL training pipeline by at least an order of magnitude. ⚡ Turn Idle GPU Compute into Gold Small labs often have underutilized hardware. OpenTinker turns your idle GPUs into an internal/external API service for - RL Training - SFT - Inference 🎯 Who needs OpenTinker? - Researchers tired of infrastructure hell. - Labs needing to standardize workflows. - Teams wanting to maximize hardware ROI. Thanks my amazing PhD student Siqi Zhu for leading the project. We are building the future of open RL infra. Be the first to build with us. 👇 Start Building with OpenTinker Now 🚀 Repo: 🌐 Blog: If you believe RL should be accessible to everyone, give us a star, repost this 🔄 post, and let us know what agents you plan to build!

Jiaxuan You

58,326 Aufrufe • vor 9 Monaten

OpenClaw meets RL! OpenClaw Agents adapt through memory files and skills, but the base model weights never actually change. OpenClaw-RL solves this! It wraps a self-hosted model as an OpenAI-compatible API, intercepts live conversations from OpenClaw, and trains the policy in the background using RL. The architecture is fully async. This means serving, reward scoring, and training all run in parallel. Once done, weights get hot-swapped after every batch while the agent keeps responding. Currently, it has two training modes: - Binary RL (GRPO): A process reward model scores each turn as good, bad, or neutral. That scalar reward drives policy updates via a PPO-style clipped objective. - On-Policy Distillation: When concrete corrections come in like "you should have checked that file first," it uses that feedback as a richer, directional training signal at the token level. When to use OpenClaw-RL? To be fair, a lot of agent behavior can already be improved through better memory and skill design. OpenClaw's existing skill ecosystem and community-built self-improvement skills handle a wide range of use cases without touching model weights at all. If the agent keeps forgetting preferences, that's a memory problem. And if it doesn't know how to handle a specific workflow, that's a skill problem. Both are solvable at the prompt and context layer. Where RL becomes interesting is when the failure pattern lives deeper in the model's reasoning itself. Things like consistently poor tool selection order, weak multi-step planning, or failing to interpret ambiguous instructions the way a specific user intends. Research on agentic RL (like ARTIST and Agent-R1) has shown that these behavioral patterns hit a ceiling with prompt-based approaches alone, especially in complex multi-turn tasks where the model needs to recover from tool failures or adapt its strategy mid-execution. That's the layer OpenClaw-RL targets, and it's a meaningful distinction from what OpenClaw offers. I have shared the repo in the replies!

Avi Chawla

138,769 Aufrufe • vor 6 Monaten

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts

Sonya Huang 🐥

81,332 Aufrufe • vor 4 Monaten

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a higher-dimensional feedback channel. Every major AI lab trains models with RL today (OpenAI, Anthropic, DeepSeek). And their key bottleneck has always been the reward functions. GRPO by DeepSeek worked well for math and code because the environment gave a binary signal. But for real agent tasks, someone still has to hand-code the scoring function. That takes days and breaks every time the pipeline changes. RULER (implemented in OpenPipe ART, 10k stars) addresses the exact problem Karpathy identified. The reward criteria are defined in plain English, and an LLM evaluates each trajectory against that description to provide feedback for training. I trained a Qwen3 1.4B agent that plays 2048 using GRPO with this exact workflow. In this case, the agent saw the board, picked a direction, and RULER evaluated the outcome, all from this natural language definition. You can see the full implementation on GitHub and try it yourself. Here's the ART Repo: (don't forget to star it ⭐ ) Just like RLHF replaced manual rankings and GRPO replaced the critic model, natural language rewards are replacing hand-coded scoring functions. RL reward engineering is now prompt engineering. I wrote a full walkthrough covering RL for LLM agents, from RLHF to GRPO to RULER, in the article below.

Avi Chawla

350,798 Aufrufe • vor 4 Monaten