Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Org
We introduced a RDMA-based, Peer to Peer weight update mechanism for RL workloads in SGLang as a supplement to traditional NCCL broadcast methods, compatible with all major open source models. By util...
Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Org Projects Blog About Donations Contact ‹ Back to Blog ‹ Back to Blog Contents Background Challenges with Existing NCCL Broadcast Design Initialization During Each Update Implementation Results Usage Future Plans Engineering Appendix Design Iterations Peer to Peer Transfer Plan Identify SGLang Tensors to Transfer Transfer Flow with Shared Replica Other Quantization and Post Processes Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL Jiadong Guo, Xin Ji, Letian Rua
saved by
related reading
- Journey to 2-second Inter-node RL Weight Transferle.qun.ch
- GLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpressprimeintellect.ai
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- How To Scale Your Modeljax-ml.github.io
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- How to Parallelize a Transformer for Training — an explorable explanationezyang.github.io
- 1910.02054v3arxiv.org