INTELLECT-2 Release: The First Globally Trained 32B Parameter Model Reinforcement Learning Training Run
We're excited to release INTELLECT-2, the first 32B parameter model trained via globally distributed reinforcement learning. Unlike traditional centralized training efforts, INTELLECT-2 trains a reasoning language model using fully asynchronous RL across a dynamic, heterogeneous swarm of permissionless compute contributors. To enable a training run with this unique infrastructure, we built various components from scratch: we introduce PRIME-RL, our training framework purpose-built for distributed asynchronous reinforcement learning, based on top of novel components such as TOPLOC, which verifies rollouts from untrusted inference workers, and SHARDCAST, which efficiently broadcasts policy weights from training nodes to inference workers. Beyond infrastructure components, we propose modifications to the standard GRPO training recipe and data filtering techniques that were crucial to achieve training stability and ensure that our model successfully learned its training objective, thus im
INTELLECT-2 Release: The First 32B Parameter Model Trained Through Globally Distributed Reinforcement Learning We're excited to release INTELLECT-2, the first 32B parameter model trained via globally distributed reinforcement learning. Unlike traditional centralized training efforts, INTELLECT-2 trains a reasoning language model using fully asynchronous RL across a dynamic, heterogeneous swarm of compute contributors. To enable a training run with this unique infrastructure, we built various components from scratch: we introduce PRIME-RL, our training framework purpose-built for distributed as
Explore this link on the map →related reading
- Composer2.pdfcursor.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- DeepSeek-R1arxiv.org
- Explore | alphaXivalphaxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- How We Build Trillion Parameter Reasoning RL with 10% GPUsmacaron.im
- LLM Resourcesforrestbicker.com
- Teaching a Language Model Arithmetic with Reinforcement Learning - Sami Khansamikhan.ai
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Unpacking decentralized training - knower's substacktheknower.substack.com