正在加载视频...

视频加载失败

I hand-wrote a 500-LoC RL stack to make hacking on RL research much easier. Most RL stacks are either massive and unhackable, or duct-taped research scripts. I am open-sourcing Mithrl, a modular RLVR stack. Next items on my checklist: adding more complex environment examples, supporting multi-gpu + async RL,...

17,676 次观看 • 6 个月前 •via X (Twitter)

36 条评论

omkaar 的头像
omkaar6 个月前

cc for feedback: @willccb @QGallouedec @_rajanagarwal @yacinelearning @danielhanchen (learnt vllm sleep/wake from here) @karpathy (easy to port to autoresearch which is what I am doing next)

omkaar 的头像
omkaar6 个月前

github:

parth 的头像
parth6 个月前

hell yeah

omkaar 的头像
omkaar6 个月前

ship squad

surya 的头像
surya6 个月前

keep cooking omkizzy

omkaar 的头像
omkaar6 个月前

me and you forever

Jibran 的头像
Jibran6 个月前

@hamostaf04 this is sick

omkaar 的头像
omkaar6 个月前

@hamostaf04 yessir thank you

mikael haji 的头像
mikael haji6 个月前

lets gooo this is fire

omkaar 的头像
omkaar6 个月前

appreciate you brother

saksham 的头像
saksham6 个月前

yoooo that’s tuff

omkaar 的头像
omkaar6 个月前

thank you brother

ari dutilh 的头像
ari dutilh6 个月前

man this is awesome

omkaar 的头像
omkaar6 个月前

i appreciate it man, lmk if you have ppl experimenting w LLM RL I could talk to

ari dutilh 的头像
ari dutilh6 个月前

@novasarc01 potentially

omkaar 的头像
omkaar6 个月前

@novasarc01 hi @novasarc01, would love to chat if this is interesting

Damon Crockett 的头像
Damon Crockett6 个月前

Hand wrote, nice

toki 的头像
toki6 个月前

this is fire

omkaar 的头像
omkaar6 个月前

appreciate it brother, lmk if you run any experiments on top

Madhav Singhal 的头像
Madhav Singhal6 个月前

this is dope

omkaar 的头像
omkaar6 个月前

thank you brother, you first put me on actual RLVR training

Chris Samra 的头像
Chris Samra6 个月前

🔥

Fahim Ahmed 的头像
Fahim Ahmed6 个月前

LFG!

omkaar 的头像
omkaar6 个月前

yesssir

atharva ☆ 的头像
atharva ☆6 个月前

this is incredible

omkaar 的头像
omkaar6 个月前

thank you brother

Ishan 的头像
Ishan6 个月前

@k7agar Nice work!

omkaar 的头像
omkaar6 个月前

@k7agar thank you!

Jash Thakkar 的头像
Jash Thakkar6 个月前

Nice work, will try this out if i get time to.

aarush 的头像
aarush6 个月前

so based

omkaar 的头像
omkaar6 个月前

yessir down to collab on an interesting env

Krish Modi 的头像
Krish Modi6 个月前

sick!!!

omkaar 的头像
omkaar6 个月前

@therealkmodi thank you brother!

shiv 的头像
shiv6 个月前

Let's connect

adi 的头像
adi6 个月前

holy shit

omkaar 的头像
omkaar6 个月前

would love to collab if andera is looking into internal rl envs

相关视频

🚨 RL for LLMs is finally accessible. Introducing OpenTinker: The first community-driven, open-source framework designed to democratize Reinforcement Learning for LLMs. Inspired by Thinking Machines's amazing Tinker, we realize the biggest bottleneck in agentic LLM research isn’t the math—it’s the setup. Current RL pipelines are messy. Configuring VeRL for every single experiment is a productivity killer. OpenTinker fixed it. 🛠 How OpenTinker Works: Decoupled Design of Server and Client - Setup Once, Run Forever: Configure the OpenTinker backend on your GPU cluster once. - Develop Locally: Define your RL environments directly on your laptop. - Train on the Cloud: Simply point your local client to the backend. The cluster handles the compute; you handle the science. 📉 The 10x Development Efficiency Thanks to our elegant architectural decomposition, OpenTinker reduces the time to develop a new RL training pipeline by at least an order of magnitude. ⚡ Turn Idle GPU Compute into Gold Small labs often have underutilized hardware. OpenTinker turns your idle GPUs into an internal/external API service for - RL Training - SFT - Inference 🎯 Who needs OpenTinker? - Researchers tired of infrastructure hell. - Labs needing to standardize workflows. - Teams wanting to maximize hardware ROI. Thanks my amazing PhD student Siqi Zhu for leading the project. We are building the future of open RL infra. Be the first to build with us. 👇 Start Building with OpenTinker Now 🚀 Repo: 🌐 Blog: If you believe RL should be accessible to everyone, give us a star, repost this 🔄 post, and let us know what agents you plan to build!

Jiaxuan You

58,326 次观看 • 9 个月前

OpenClaw meets RL! OpenClaw Agents adapt through memory files and skills, but the base model weights never actually change. OpenClaw-RL solves this! It wraps a self-hosted model as an OpenAI-compatible API, intercepts live conversations from OpenClaw, and trains the policy in the background using RL. The architecture is fully async. This means serving, reward scoring, and training all run in parallel. Once done, weights get hot-swapped after every batch while the agent keeps responding. Currently, it has two training modes: - Binary RL (GRPO): A process reward model scores each turn as good, bad, or neutral. That scalar reward drives policy updates via a PPO-style clipped objective. - On-Policy Distillation: When concrete corrections come in like "you should have checked that file first," it uses that feedback as a richer, directional training signal at the token level. When to use OpenClaw-RL? To be fair, a lot of agent behavior can already be improved through better memory and skill design. OpenClaw's existing skill ecosystem and community-built self-improvement skills handle a wide range of use cases without touching model weights at all. If the agent keeps forgetting preferences, that's a memory problem. And if it doesn't know how to handle a specific workflow, that's a skill problem. Both are solvable at the prompt and context layer. Where RL becomes interesting is when the failure pattern lives deeper in the model's reasoning itself. Things like consistently poor tool selection order, weak multi-step planning, or failing to interpret ambiguous instructions the way a specific user intends. Research on agentic RL (like ARTIST and Agent-R1) has shown that these behavioral patterns hit a ceiling with prompt-based approaches alone, especially in complex multi-turn tasks where the model needs to recover from tool failures or adapt its strategy mid-execution. That's the layer OpenClaw-RL targets, and it's a meaningful distinction from what OpenClaw offers. I have shared the repo in the replies!

Avi Chawla

138,769 次观看 • 6 个月前

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts

Sonya Huang 🐥

81,332 次观看 • 4 个月前

Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a higher-dimensional feedback channel. Every major AI lab trains models with RL today (OpenAI, Anthropic, DeepSeek). And their key bottleneck has always been the reward functions. GRPO by DeepSeek worked well for math and code because the environment gave a binary signal. But for real agent tasks, someone still has to hand-code the scoring function. That takes days and breaks every time the pipeline changes. RULER (implemented in OpenPipe ART, 10k stars) addresses the exact problem Karpathy identified. The reward criteria are defined in plain English, and an LLM evaluates each trajectory against that description to provide feedback for training. I trained a Qwen3 1.4B agent that plays 2048 using GRPO with this exact workflow. In this case, the agent saw the board, picked a direction, and RULER evaluated the outcome, all from this natural language definition. You can see the full implementation on GitHub and try it yourself. Here's the ART Repo: (don't forget to star it ⭐ ) Just like RLHF replaced manual rankings and GRPO replaced the critic model, natural language rewards are replacing hand-coded scoring functions. RL reward engineering is now prompt engineering. I wrote a full walkthrough covering RL for LLM agents, from RLHF to GRPO to RULER, in the article below.

Avi Chawla

350,798 次观看 • 4 个月前