RL Post-Training on Macs | Pluralis Research
A decoupled reinforcement learning run of LFM2.5-8B-A1B on 14 Macs across 4 countries generating int8 rollouts, with one B200 computing every gradient in bf16. On PaperSearchQA, held-out pass@1 more than doubled, from 29% to 63%.
We just finished a multi-turn reinforcement learning (RL) run of LFM2.5-8B-A1B, [2] an 8.3B-parameter MoE, on 14 Macs spread across four countries. And one of them was my MacBook. The Macs generated the rollouts; a single B200, an ocean away, did the training. On PaperSearchQA, [14] an agentic task where the model answers biomedical questions by searching a corpus of research papers, held-out pass@1 more than doubled, from 29% to 63%. As far as we can tell, this is the first time anything like this has happened: the largest multi-Mac setups we could find, eight EXO machines [9] and four RDMA-l
Explore this link on the map →related reading
- Composer2.pdfcursor.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- Orbit - Ultra-efficient RL Pipelinespherelab.ai
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai