RL at 1T Scale: prime-rl Performance Deep Dive
prime-rl 0.6.0 trains trillion-parameter MoE models on heavy agentic workloads at the highest efficiency. A deep dive into the inference and training optimizations behind it — from FP8 and wide expert parallelism to P/D disaggregation, router replay, and 3-D parallelism.
RL at 1T Scale: prime-rl Performance Deep Dive Today we are releasing prime-rl version 0.6.0. This version enables us (and you) to train models of trillion-parameter scale on heavy agentic workloads at the highest efficiency. We have been relentlessly optimizing our RL infrastructure to maximize performance on large MoE models, reducing the cost, time and suffering required to post-train OSS models on agentic workflows. We are able to train GLM-5 on SWE tasks at up to 131k sequence length, with sub-5-minute step times and a batch size of 256 rollouts, on only 28 H200 nodes. In this blog we wil
Explore this link on the map →related reading
- How To Scale Your Modeljax-ml.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Composer2.pdfcursor.com
- How We Build Trillion Parameter Reasoning RL with 10% GPUsmacaron.im
- Orbit - Ultra-efficient RL Pipelinespherelab.ai
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Rethinking RL Infra for Agents | B'Logbillxbf.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- LLM Resourcesforrestbicker.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com