Video wird geladen...
Video konnte nicht geladen werden
RL post-training on Macs 14 Macs across 4 countries generate every rollout for the run. Everything's running over the internet. No wire between any of them. As far as we can tell, this is the first RL post-training run with its whole rollout fleet on consumer Macs.
35,576 Aufrufe • vor 2 Monaten •via X (Twitter)
19 Kommentare

Why Macs for inference? They are not built for training, but they are very good at inference. Unified memory makes them especially strong for MoE models. So, we use LFM2.5-8B-A1B, an MoE model with 1.5B active parameters. And ~80% of the compute in agentic RL is inference.

One of the hardest problems here is off-policiness, and it hits from three sides at once. Normal RL trains on data from the exact model it's updating (on-policy). Ours trains on batches from Macs that run weights (1) a few versions old, (2) in int8, (3) on a different GPU stack than the bf16 trainer-- M2/3/4 Base & Pro vs B200.

Same tokens, different probs: tiny disagreements between the int8 Macs and the bf16 trainer compound until training crashes. Our fix: DPPO. Measure the divergence; drop the ~0.3% of tokens that drift too far.

9 GB of weights change every few minutes. Pulling that over home internet takes minutes. The Macs would run stale for all that time. PULSE's insight: most updates are too small to survive a precision cast. In int8, that's about 0.5% of values. Ship just those: 82 MB instead of 9 GB. Transfer happens in seconds.

Our system ties these pieces together so neither the trainer nor the Macs sit idle: DPPO absorbs staleness, PULSE syncs weights in seconds, and every rollout streams the moment it's done.

To show it works, we picked an agentic search task (PaperSearchQA): search biomedical papers, across multiple turns, to answer questions it's never seen. It learned. On the full validation set, pass@1 went from 29% to 63%.

Two challenges remain. Each Mac holds the whole model, so we can only train models that fit on one Mac, and the trainer sits in a single cluster. Agora gets past both. It just pretrained Pluralis-8B across hundreds of consumer GPUs, the first pipeline-parallel training run over the open internet as far as we can tell.

Put Stoa (our system) and Agora together, and the whole loop could run in the open: large-model inference across the Macs, large-model training on Agora. The compute is already out there, idle, more than the clusters behind today's frontier models combined. As the best models are drifting behind closed APIs, training them on hardware people already own, owned by the people who train them, is how we get them back. Blog: Code:

@hmdolatabadi This is really cool guys!

which country has the slowest mac

This is sick

Super cool @AlexanderLong

👀

macs running rl like secret weapon energy lol

What happened to the decentralised training runs we had been doing? Node0 abandoned?

how badly does the network communication overhead affect the wall-clock measurements for the same training workload? (compared to both: an equal-cost local GPU cluster and an equivalent local mac cluster)

14 Macs across four countries is a pretty wild rollout fleet. no wire between any of them is the fun part

Sci-fi spent decades imagining AI born in massive underground server farms. Turns out it's 14 MacBooks on home WiFi across four countries. Somehow that's more unsettling.

14 Macs across 4 countries pulling this off over the internet is a wild distributed setup
