flâneur — a map of the web's best reading

RL Post-Training on Macs | Pluralis Research

pluralis.ai · 7,396 words · saved by 1 readers

A decoupled reinforcement learning run of LFM2.5-8B-A1B on 14 Macs across 4 countries generating int8 rollouts, with one B200 computing every gradient in bf16. On PaperSearchQA, held-out pass@1 more than doubled, from 29% to 63%.

We just finished a multi-turn reinforcement learning (RL) run of LFM2.5-8B-A1B, [2] an 8.3B-parameter MoE, on 14 Macs spread across four countries. And one of them was my MacBook. The Macs generated the rollouts; a single B200, an ocean away, did the training. On PaperSearchQA, [14] an agentic task where the model answers biomedical questions by searching a corpus of research papers, held-out pass@1 more than doubled, from 29% to 63%. As far as we can tell, this is the first time anything like this has happened: the largest multi-Mac setups we could find, eight EXO machines [9] and four RDMA-l

Explore this link on the map →

related reading