Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Org
We introduced a RDMA-based, Peer to Peer weight update mechanism for RL workloads in SGLang as a supplement to traditional NCCL broadcast methods, compatible with all major open source models. By util...
Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL - LMSYS Org Projects Blog About Donations Contact ‹ Back to Blog ‹ Back to Blog Contents Background Challenges with Existing NCCL Broadcast Design Initialization During Each Update Implementation Results Usage Future Plans Engineering Appendix Design Iterations Peer to Peer Transfer Plan Identify SGLang Tensors to Transfer Transfer Flow with Shared Replica Other Quantization and Post Processes Updating 1T parameters in seconds — P2P weight transfer in Large Scale Distributed RL Jiadong Guo, Xin Ji, Letian Rua
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- 5D parallelism in LLM training - gdymind's Bloggdymind.com
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- Parallelism in Distributed Deep Learning · Better Tomorrow with Computer Scienceinsujang.github.io
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- [2411.19870] DeMo: Decoupled Momentum Optimizationarxiv.org
- Orbit - Ultra-efficient RL Pipelinespherelab.ai