RL Post-Training on Macs | Pluralis Research
A decoupled reinforcement learning run of LFM2.5-8B-A1B on 14 Macs across 4 countries generating int8 rollouts, with one B200 computing every gradient in bf16. On PaperSearchQA, held-out pass@1 more than doubled, from 29% to 63%.
We just finished a multi-turn reinforcement learning (RL) run of LFM2.5-8B-A1B, [2] an 8.3B-parameter MoE, on 14 Macs spread across four countries. And one of them was my MacBook. The Macs generated the rollouts; a single B200, an ocean away, did the training. On PaperSearchQA, [14] an agentic task where the model answers biomedical questions by searching a corpus of research papers, held-out pass@1 more than doubled, from 29% to 63%. As far as we can tell, this is the first time anything like this has happened: the largest multi-Mac setups we could find, eight EXO machines [9] and four RDMA-l
saved by
related reading
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Composer2.pdfcursor.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Is Frontier Asynchronous RL Solved? — Luke J. Huangluk-huang.github.io
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org