Loading video...

Video Failed to Load

Go Home

RL post-training on Macs 14 Macs across 4 countries generate every rollout for the run. Everything's running over the internet. No wire between any of them. As far as we can tell, this is the first RL post-training run with its whole rollout fleet on consumer Macs.

35,576 views • 2 months ago •via X (Twitter)

19 Comments

Pluralis Research's profile picture
Pluralis Research2 months ago

Why Macs for inference? They are not built for training, but they are very good at inference. Unified memory makes them especially strong for MoE models. So, we use LFM2.5-8B-A1B, an MoE model with 1.5B active parameters. And ~80% of the compute in agentic RL is inference.

Pluralis Research's profile picture
Pluralis Research2 months ago

One of the hardest problems here is off-policiness, and it hits from three sides at once. Normal RL trains on data from the exact model it's updating (on-policy). Ours trains on batches from Macs that run weights (1) a few versions old, (2) in int8, (3) on a different GPU stack than the bf16 trainer-- M2/3/4 Base & Pro vs B200.

Pluralis Research's profile picture
Pluralis Research2 months ago

Same tokens, different probs: tiny disagreements between the int8 Macs and the bf16 trainer compound until training crashes. Our fix: DPPO. Measure the divergence; drop the ~0.3% of tokens that drift too far.

Pluralis Research's profile picture
Pluralis Research2 months ago

9 GB of weights change every few minutes. Pulling that over home internet takes minutes. The Macs would run stale for all that time. PULSE's insight: most updates are too small to survive a precision cast. In int8, that's about 0.5% of values. Ship just those: 82 MB instead of 9 GB. Transfer happens in seconds.

Pluralis Research's profile picture
Pluralis Research2 months ago

Our system ties these pieces together so neither the trainer nor the Macs sit idle: DPPO absorbs staleness, PULSE syncs weights in seconds, and every rollout streams the moment it's done.

Pluralis Research's profile picture
Pluralis Research2 months ago

To show it works, we picked an agentic search task (PaperSearchQA): search biomedical papers, across multiple turns, to answer questions it's never seen. It learned. On the full validation set, pass@1 went from 29% to 63%.

Pluralis Research's profile picture
Pluralis Research2 months ago

Two challenges remain. Each Mac holds the whole model, so we can only train models that fit on one Mac, and the trainer sits in a single cluster. Agora gets past both. It just pretrained Pluralis-8B across hundreds of consumer GPUs, the first pipeline-parallel training run over the open internet as far as we can tell.

Pluralis Research's profile picture
Pluralis Research2 months ago

Put Stoa (our system) and Agora together, and the whole loop could run in the open: large-model inference across the Macs, large-model training on Agora. The compute is already out there, idle, more than the clusters behind today's frontier models combined. As the best models are drifting behind closed APIs, training them on hardware people already own, owned by the people who train them, is how we get them back. Blog: Code:

Eiso Kant's profile picture
Eiso Kant2 months ago

@hmdolatabadi This is really cool guys!

ajay yadav's profile picture
ajay yadav2 months ago

which country has the slowest mac

Kydo's profile picture
Kydo2 months ago

This is sick

Mark's profile picture
Mark2 months ago

Super cool @AlexanderLong

Sachi Kamiya 幸's profile picture
Sachi Kamiya 幸2 months ago

👀

lvnbbs_bnb 🐬TermMax's profile picture
lvnbbs_bnb 🐬TermMax2 months ago

macs running rl like secret weapon energy lol

django's profile picture
django2 months ago

What happened to the decentralised training runs we had been doing? Node0 abandoned?

Yiğit Polat's profile picture
Yiğit Polat2 months ago

how badly does the network communication overhead affect the wall-clock measurements for the same training workload? (compared to both: an equal-cost local GPU cluster and an equivalent local mac cluster)

ZERΘ MΛXX's profile picture
ZERΘ MΛXX2 months ago

14 Macs across four countries is a pretty wild rollout fleet. no wire between any of them is the fun part

Joseph K's profile picture
Joseph K2 months ago

Sci-fi spent decades imagining AI born in massive underground server farms. Turns out it's 14 MacBooks on home WiFi across four countries. Somehow that's more unsettling.

AI Mastery Guide's profile picture
AI Mastery Guide2 months ago

14 Macs across 4 countries pulling this off over the internet is a wild distributed setup

Related Videos

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 views • 1 year ago

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts

Sonya Huang 🐥

81,490 views • 4 months ago