Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I hand-wrote a 500-LoC RL stack to make hacking on RL research much easier. Most RL stacks are either massive and unhackable, or duct-taped research scripts. I am open-sourcing Mithrl, a modular RLVR stack. Next items on my checklist: adding more complex environment examples, supporting multi-gpu + async RL,...

17,676 görüntüleme • 6 ay önce •via X (Twitter)

36 Yorum

omkaar profil fotoğrafı
omkaar6 ay önce

cc for feedback: @willccb @QGallouedec @_rajanagarwal @yacinelearning @danielhanchen (learnt vllm sleep/wake from here) @karpathy (easy to port to autoresearch which is what I am doing next)

omkaar profil fotoğrafı
omkaar6 ay önce

github:

parth profil fotoğrafı
parth6 ay önce

hell yeah

omkaar profil fotoğrafı
omkaar6 ay önce

ship squad

surya profil fotoğrafı
surya6 ay önce

keep cooking omkizzy

omkaar profil fotoğrafı
omkaar6 ay önce

me and you forever

Jibran profil fotoğrafı
Jibran6 ay önce

@hamostaf04 this is sick

omkaar profil fotoğrafı
omkaar6 ay önce

@hamostaf04 yessir thank you

mikael haji profil fotoğrafı
mikael haji6 ay önce

lets gooo this is fire

omkaar profil fotoğrafı
omkaar6 ay önce

appreciate you brother

saksham profil fotoğrafı
saksham6 ay önce

yoooo that’s tuff

omkaar profil fotoğrafı
omkaar6 ay önce

thank you brother

ari dutilh profil fotoğrafı
ari dutilh6 ay önce

man this is awesome

omkaar profil fotoğrafı
omkaar6 ay önce

i appreciate it man, lmk if you have ppl experimenting w LLM RL I could talk to

ari dutilh profil fotoğrafı
ari dutilh6 ay önce

@novasarc01 potentially

omkaar profil fotoğrafı
omkaar6 ay önce

@novasarc01 hi @novasarc01, would love to chat if this is interesting

Damon Crockett profil fotoğrafı
Damon Crockett6 ay önce

Hand wrote, nice

toki profil fotoğrafı
toki6 ay önce

this is fire

omkaar profil fotoğrafı
omkaar6 ay önce

appreciate it brother, lmk if you run any experiments on top

Madhav Singhal profil fotoğrafı
Madhav Singhal6 ay önce

this is dope

omkaar profil fotoğrafı
omkaar6 ay önce

thank you brother, you first put me on actual RLVR training

Chris Samra profil fotoğrafı
Chris Samra6 ay önce

🔥

Fahim Ahmed profil fotoğrafı
Fahim Ahmed6 ay önce

LFG!

omkaar profil fotoğrafı
omkaar6 ay önce

yesssir

atharva ☆ profil fotoğrafı
atharva ☆6 ay önce

this is incredible

omkaar profil fotoğrafı
omkaar6 ay önce

thank you brother

Ishan profil fotoğrafı
Ishan6 ay önce

@k7agar Nice work!

omkaar profil fotoğrafı
omkaar6 ay önce

@k7agar thank you!

Jash Thakkar profil fotoğrafı
Jash Thakkar6 ay önce

Nice work, will try this out if i get time to.

aarush profil fotoğrafı
aarush6 ay önce

so based

omkaar profil fotoğrafı
omkaar6 ay önce

yessir down to collab on an interesting env

Krish Modi profil fotoğrafı
Krish Modi6 ay önce

sick!!!

omkaar profil fotoğrafı
omkaar6 ay önce

@therealkmodi thank you brother!

shiv profil fotoğrafı
shiv6 ay önce

Let's connect

adi profil fotoğrafı
adi6 ay önce

holy shit

omkaar profil fotoğrafı
omkaar6 ay önce

would love to collab if andera is looking into internal rl envs

Benzer Videolar

🚨 RL for LLMs is finally accessible. Introducing OpenTinker: The first community-driven, open-source framework designed to democratize Reinforcement Learning for LLMs. Inspired by Thinking Machines's amazing Tinker, we realize the biggest bottleneck in agentic LLM research isn’t the math—it’s the setup. Current RL pipelines are messy. Configuring VeRL for every single experiment is a productivity killer. OpenTinker fixed it. 🛠 How OpenTinker Works: Decoupled Design of Server and Client - Setup Once, Run Forever: Configure the OpenTinker backend on your GPU cluster once. - Develop Locally: Define your RL environments directly on your laptop. - Train on the Cloud: Simply point your local client to the backend. The cluster handles the compute; you handle the science. 📉 The 10x Development Efficiency Thanks to our elegant architectural decomposition, OpenTinker reduces the time to develop a new RL training pipeline by at least an order of magnitude. ⚡ Turn Idle GPU Compute into Gold Small labs often have underutilized hardware. OpenTinker turns your idle GPUs into an internal/external API service for - RL Training - SFT - Inference 🎯 Who needs OpenTinker? - Researchers tired of infrastructure hell. - Labs needing to standardize workflows. - Teams wanting to maximize hardware ROI. Thanks my amazing PhD student Siqi Zhu for leading the project. We are building the future of open RL infra. Be the first to build with us. 👇 Start Building with OpenTinker Now 🚀 Repo: 🌐 Blog: If you believe RL should be accessible to everyone, give us a star, repost this 🔄 post, and let us know what agents you plan to build!

Jiaxuan You

58,326 görüntüleme • 9 ay önce

OpenClaw meets RL! OpenClaw Agents adapt through memory files and skills, but the base model weights never actually change. OpenClaw-RL solves this! It wraps a self-hosted model as an OpenAI-compatible API, intercepts live conversations from OpenClaw, and trains the policy in the background using RL. The architecture is fully async. This means serving, reward scoring, and training all run in parallel. Once done, weights get hot-swapped after every batch while the agent keeps responding. Currently, it has two training modes: - Binary RL (GRPO): A process reward model scores each turn as good, bad, or neutral. That scalar reward drives policy updates via a PPO-style clipped objective. - On-Policy Distillation: When concrete corrections come in like "you should have checked that file first," it uses that feedback as a richer, directional training signal at the token level. When to use OpenClaw-RL? To be fair, a lot of agent behavior can already be improved through better memory and skill design. OpenClaw's existing skill ecosystem and community-built self-improvement skills handle a wide range of use cases without touching model weights at all. If the agent keeps forgetting preferences, that's a memory problem. And if it doesn't know how to handle a specific workflow, that's a skill problem. Both are solvable at the prompt and context layer. Where RL becomes interesting is when the failure pattern lives deeper in the model's reasoning itself. Things like consistently poor tool selection order, weak multi-step planning, or failing to interpret ambiguous instructions the way a specific user intends. Research on agentic RL (like ARTIST and Agent-R1) has shown that these behavioral patterns hit a ceiling with prompt-based approaches alone, especially in complex multi-turn tasks where the model needs to recover from tool failures or adapt its strategy mid-execution. That's the layer OpenClaw-RL targets, and it's a meaningful distinction from what OpenClaw offers. I have shared the repo in the replies!

Avi Chawla

138,769 görüntüleme • 6 ay önce

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts

Sonya Huang 🐥

81,332 görüntüleme • 4 ay önce

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a higher-dimensional feedback channel. Every major AI lab trains models with RL today (OpenAI, Anthropic, DeepSeek). And their key bottleneck has always been the reward functions. GRPO by DeepSeek worked well for math and code because the environment gave a binary signal. But for real agent tasks, someone still has to hand-code the scoring function. That takes days and breaks every time the pipeline changes. RULER (implemented in OpenPipe ART, 10k stars) addresses the exact problem Karpathy identified. The reward criteria are defined in plain English, and an LLM evaluates each trajectory against that description to provide feedback for training. I trained a Qwen3 1.4B agent that plays 2048 using GRPO with this exact workflow. In this case, the agent saw the board, picked a direction, and RULER evaluated the outcome, all from this natural language definition. You can see the full implementation on GitHub and try it yourself. Here's the ART Repo: (don't forget to star it ⭐ ) Just like RLHF replaced manual rankings and GRPO replaced the critic model, natural language rewards are replacing hand-coded scoring functions. RL reward engineering is now prompt engineering. I wrote a full walkthrough covering RL for LLM agents, from RLHF to GRPO to RULER, in the article below.

Avi Chawla

350,798 görüntüleme • 4 ay önce