RL at 1T Scale: prime-rl Performance Deep Dive
prime-rl 0.6.0 trains trillion-parameter MoE models on heavy agentic workloads at the highest efficiency. A deep dive into the inference and training optimizations behind it — from FP8 and wide expert parallelism to P/D disaggregation, router replay, and 3-D parallelism.
RL at 1T Scale: prime-rl Performance Deep Dive Today we are releasing prime-rl version 0.6.0. This version enables us (and you) to train models of trillion-parameter scale on heavy agentic workloads at the highest efficiency. We have been relentlessly optimizing our RL infrastructure to maximize performance on large MoE models, reducing the cost, time and suffering required to post-train OSS models on agentic workflows. We are able to train GLM-5 on SWE tasks at up to 131k sequence length, with sub-5-minute step times and a batch size of 256 rollouts, on only 28 H200 nodes. In this blog we wil
saved by
related reading
- GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpressprimeintellect.ai
- Journey to 2-second Inter-node RL Weight Transferle.qun.ch
- Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Orglmsys.org
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- How To Scale Your Modeljax-ml.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Composer2.pdfcursor.com
- How We Build Trillion Parameter Reasoning RL with 10% GPUsmacaron.im
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Orbit - Ultra-efficient RL Pipelinespherelab.ai
- Async RL in Pure JAXdivyamakkar0.github.io
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com