flâneur — a map of the web's best reading

RL at 1T Scale: prime-rl Performance Deep Dive

primeintellect.ai · 2,305 words · saved by 1 readers

prime-rl 0.6.0 trains trillion-parameter MoE models on heavy agentic workloads at the highest efficiency. A deep dive into the inference and training optimizations behind it — from FP8 and wide expert parallelism to P/D disaggregation, router replay, and 3-D parallelism.

RL at 1T Scale: prime-rl Performance Deep Dive Today we are releasing prime-rl version 0.6.0. This version enables us (and you) to train models of trillion-parameter scale on heavy agentic workloads at the highest efficiency. We have been relentlessly optimizing our RL infrastructure to maximize performance on large MoE models, reducing the cost, time and suffering required to post-train OSS models on agentic workflows. We are able to train GLM-5 on SWE tasks at up to 131k sequence length, with sub-5-minute step times and a batch size of 256 rollouts, on only 28 H200 nodes. In this blog we wil

Explore this link on the map →

related reading